Read GCS and local paths with the same interface, clone of tensorflow.io.gfile

Project description

blobfile

This is a standalone clone of TensorFlow's gfile, supporting local paths, Google Cloud Storage paths (gs://), and Azure Blobs paths (https://<account>.blob.core.windows.net/<container>/).

The main function is BlobFile, a replacement for GFile. There are also a few additional functions, basename, dirname, and join, which mostly do the same thing as their os.path namesakes, only they also support GCS paths and Azure Storage paths.

Installation

pip install blobfile

Usage

import blobfile as bf

with bf.BlobFile("gs://my-bucket-name/cats", "wb") as w:
    w.write(b"meow!")

Here are the functions:

BlobFile - like open() but works with remote paths too, data can be streamed to/from the remote file. It accepts the following arguments:
- streaming:
  - The default for streaming is True when mode is in "r", "rb" and False when mode is in "w", "wb", "a", "ab".
  - streaming=True:
    - Reading is done without downloading the entire remote file.
    - Writing is done to the remote file directly, but only in chunks of a few MB in size. flush() will not cause an early write.
    - Appending is not implemented.
  - streaming=False:
    - Reading is done by downloading the remote file to a local file during the constructor.
    - Writing is done by uploading the file on close() or during destruction.
    - Appending is done by downloading the file during construction and uploading on close().
- buffer_size: number of bytes to buffer, this can potentially make reading more efficient.
- cache_dir: a directory in which to cache files for reading, only valid if streaming=False and mode is in "r", "rb". You are reponsible for cleaning up the cache directory.

Some are inspired by existing os.path and shutil functions:

copy - copy a file from one path to another, this will do a remote copy between two remote paths on the same blob storage service
exists - returns True if the file or directory exists
glob/scanglob - return files matching a glob-style pattern as a generator. Globs can have surprising performance characteristics when used with blob storage. Character ranges are not supported in patterns.
isdir - returns True if the path is a directory
listdir/scandir - list contents of a directory as a generator
makedirs - ensure that a directory and all parent directories exist
remove - remove a file
rmdir - remove an empty directory
rmtree - remove a directory tree
stat - get the size and modification time of a file
walk - walk a directory tree with a generator that yields (dirpath, dirnames, filenames) tuples
basename - get the final component of a path
dirname - get the path except for the final component
join - join 2 or more paths together, inserting directory separators between each component

There are a few bonus functions:

get_url - returns a url for a path (usable by an HTTP client without any authentication) along with the expiration for that url (or None)
md5 - get the md5 hash for a path, for GCS this is often fast, but for other backends this may be slow. On Azure, if the md5 of a file is calculated and is missing from the file, the file will be updated with the calculated md5.
set_mtime - set the modified timestamp for a file
configure - set global configuration options for blobfile
- log_callback=_default_log_fn: a log callback function log(msg: string) to use instead of printing to stdout
- connection_pool_max_size=32: the max size for each per-host connection pool
- max_connection_pool_count=10: the maximum count of per-host connection pools
- azure_write_chunk_size=4 * 2 ** 20: the size of blocks to write to Azure Storage blobs, can be set to a maximum of 100MB. This determines both the unit of request retries as well as the maximum file size, which is 50,000 * azure_write_chunk_size.
- retry_log_threshold=0: set a retry count threshold above which to log failures to the log callback function

Authentication

Google Cloud Storage

The following methods will be tried in order:

Check the environment variable GOOGLE_APPLICATION_CREDENTIALS for a path to service account credentials in JSON format.
Check for "application default credentials". To setup application default credentials, run gcloud auth application-default login.
Check for a GCE metadata server (if running on GCE) and get credentials from that service.

Azure Blobs

The following methods will be tried in order:

Check the environment variable AZURE_STORAGE_KEY for an azure storage account key (these are per-storage account shared keys described in https://docs.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage)
Check the environment variable AZURE_APPLICATION_CREDENTIALS which should point to JSON credentials for a service principal output by the command az ad sp create-for-rbac --name <name>
Check the environment variables AZURE_CLIENT_ID, AZURE_CLIENT_SECRET, AZURE_TENANT_ID corresponding to a service principal described in the previous step but without the JSON file.
Check the environment variable AZURE_STORAGE_CONNECTION_STRING for an Azure Storage connection string
Use credentials from the az command line tool if they can be found.

Paths

For Google Cloud Storage and Azure Blobs directories don't really exist. These storage systems store the files in a single flat list. The "/" separators are just part of the filenames and there is no need to call the equivalent of os.mkdir on one of these systems.

To make local behavior consistent with the remote storage systems, missing local directories will be created automatically when opening a file in write mode.

Local

These are just normal paths for the current machine, e.g. /root/hello.txt

Google Cloud Storage

GCS paths have the format gs://<bucket>/<blob>, you cannot perform any operations on gs:// itself.

Azure Blobs

Azure Blobs URLs have the format https://<account>.blob.core.windows.net/<container>/<blob>. The highest you can go up the hierarchy is https://<account>.blob.core.windows.net/<container>/, blobfile cannot perform any operations on https://<account>.blob.core.windows.net/.

Errors

Error - base class for library-specific exceptions
RequestFailure(Error) - a request has failed permanently, has message:str, request:Request, and response:urllib3.HTTPResponse attributes.
RestartableStreamingWriteFailure(RequestFailure) - a streaming write has failed permanently, which requires restarting from the beginning of the stream.
ConcurrentWriteFailure(RequestFailure) - a write failed because another process was writing to the same file at the same time.
The following generic exceptions are raised from some functions to make the behavior similar to the original versions: FileNotFoundError, FileExistsError, IsADirectoryError, NotADirectoryError, OSError, ValueError, io.UnsupportedOperation

Concurrent Writers

Google Cloud Storage supports multiple writers for the same blob and the last one to finish should win. However, in the event of a large number of simultaneous writers, the service will return 503 errors and most writers will stall. In this case, write to different blobs instead.

Azure Blobs doesn't support multiple writers for the same blob. With the way BlobFile is currently configured, the last writer to start writing will win. Other writers will get a ConcurrentWriteFailure. In addition, all writers could fail if the file size is large. In this case, you can write to a temporary blob (with a random filename), copy it to the final location, and then delete the original. The copy will be within a container so should be fast.

Logging

blobfile will keep retrying transient errors until they succeed or a permanent error is encountered (which will raise an exception). In order to make diagnosing stalls easier, blobfile will log when retrying requests.

To route those log lines, use configure(log_callback=<fn>) to set a callback function which will be called whenever a log line should be printed. The default callback prints to stdout with the prefix blobfile:.

While blobfile does not use the python logging module, it does use other libraries which uses that module. So if you configure the python logging module, you may need to change the settings to adjust logging behavior:

urllib3: logging.getLogger("urllib3").setLevel(logging.ERROR)
filelock: logging.getLogger("filelock").setLevel(logging.ERROR)

Examples

Write and read a file

import blobfile as bf

with bf.BlobFile("gs://my-bucket/file.name", "wb") as f:
    f.write(b"meow")

print("exists:", bf.exists("gs://my-bucket/file.name"))

print("contents:", bf.BlobFile("gs://my-bucket/file.name", "rb").read())

See Parallel Examples

Changes

See CHANGES

Project details

Release history Release notifications | RSS feed

3.0.0

Aug 27, 2024

2.1.1

Nov 2, 2023

2.1.0

Oct 12, 2023

2.0.2

Apr 21, 2023

2.0.1

Jan 10, 2023

2.0.0

Sep 30, 2022

1.3.4

Sep 24, 2022

1.3.3

Aug 25, 2022

1.3.2

Aug 24, 2022

1.3.1

May 7, 2022

1.3.0

May 2, 2022

1.2.9

Apr 11, 2022

1.2.8

Jan 19, 2022

1.2.7

Nov 20, 2021

1.2.6

Nov 11, 2021

1.2.5

Sep 24, 2021

1.2.4

Sep 21, 2021

1.2.3

Jun 5, 2021

1.2.2

Jun 4, 2021

1.2.1

May 20, 2021

1.2.0

Apr 17, 2021

1.1.3

Apr 17, 2021

1.1.2

Mar 27, 2021

1.1.1

Feb 3, 2021

1.1.0

Jan 30, 2021

1.0.7

Dec 20, 2020

1.0.6

Nov 26, 2020

1.0.5

Nov 14, 2020

1.0.4

Nov 12, 2020

1.0.3

Nov 6, 2020

1.0.2

Nov 6, 2020

1.0.1

Nov 6, 2020

1.0.0

Oct 18, 2020

0.17.3

Oct 13, 2020

0.17.2

Sep 18, 2020

0.17.1

Sep 1, 2020

This version

0.17.0

Jul 2, 2020

0.16.11

Jun 17, 2020

0.16.10

Jun 7, 2020

0.16.9

May 8, 2020

0.16.8

May 5, 2020

0.16.7

Apr 28, 2020

0.16.6

Apr 24, 2020

0.16.5

Apr 11, 2020

0.16.4

Apr 7, 2020

0.16.3

Mar 19, 2020

0.16.2

Mar 17, 2020

0.16.1

Mar 9, 2020

0.16.0

Mar 5, 2020

0.15.0

Mar 2, 2020

0.14.0

Feb 12, 2020

0.12.0

Feb 1, 2020

0.11.0

Jan 31, 2020

0.10.2

Jan 30, 2020

0.10.1

Dec 14, 2019

0.10.0

Dec 7, 2019

0.9.0

Nov 26, 2019

0.8.1

Nov 24, 2019

0.8.0

Nov 16, 2019

0.7.0

Nov 14, 2019

0.6.1

Nov 2, 2019

0.5.0

Oct 30, 2019

0.4.5

Oct 30, 2019

0.4.4

Oct 27, 2019

0.4.3

Oct 27, 2019

0.4.2

Oct 26, 2019

0.4.1

Oct 24, 2019

0.4.0

Oct 23, 2019

0.3.3

Oct 20, 2019

0.3.2

Oct 19, 2019

0.3.1

Sep 10, 2019

0.3.0

Sep 8, 2019

0.2.3

Jul 11, 2019

0.2.2

Jul 1, 2019

0.2.1

Jun 26, 2019

0.2.0

Jun 25, 2019

0.1

Jun 24, 2019

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

blobfile-0.17.0-py3-none-any.whl (50.6 kB view hashes)

Uploaded Jul 2, 2020 Python 3

Hashes for blobfile-0.17.0-py3-none-any.whl

Hashes for blobfile-0.17.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`b383b395228659c270ec940481c6c77d2d95fe36dd76ebc09c0a0ece5f234f84`
MD5	`022e4889a47e0e8f3287c26136afc1fd`
BLAKE2b-256	`3bb2202279033cd0285beb713a4066a25d5844f5c5915b14d80c4372dbe4d3e2`