adlfs·PyPI

Access Azure Datalake Gen1 with fsspec and dask

These details have not been verified by PyPI

Project description

Filesystem interface to Azure-Datalake Gen1 and Gen2 Storage

Quickstart

This package can be installed using:

pip install adlfs

conda install -c conda-forge adlfs

The adl:// and abfs:// protocols are included in fsspec's known_implementations registry in fsspec > 0.6.1, otherwise users must explicitly inform fsspec about the supported adlfs protocols.

To use the Gen1 filesystem:

import dask.dataframe as dd

storage_options={'tenant_id': TENANT_ID, 'client_id': CLIENT_ID, 'client_secret': CLIENT_SECRET}

dd.read_csv('adl://{STORE_NAME}/{FOLDER}/*.csv', storage_options=storage_options)

To use the Gen2 filesystem you can use the protocol abfs or az:

import dask.dataframe as dd

storage_options={'account_name': ACCOUNT_NAME, 'account_key': ACCOUNT_KEY}

ddf = dd.read_csv('abfs://{CONTAINER}/{FOLDER}/*.csv', storage_options=storage_options)
ddf = dd.read_parquet('az://{CONTAINER}/folder.parquet', storage_options=storage_options)

Accepted protocol / uri formats include:
'PROTOCOL://container/path-part/file'
'PROTOCOL://container@account.dfs.core.windows.net/path-part/file'

or optionally, if AZURE_STORAGE_ACCOUNT_NAME and an AZURE_STORAGE_<CREDENTIAL> is 
set as an environmental variable, then storage_options will be read from the environmental
variables

To read from a public storage blob you are required to specify the 'account_name'. For example, you can access NYC Taxi & Limousine Commission as:

storage_options = {'account_name': 'azureopendatastorage'}
ddf = dd.read_parquet('az://nyctlc/green/puYear=2019/puMonth=*/*.parquet', storage_options=storage_options)

Details

The package includes pythonic filesystem implementations for both Azure Datalake Gen1 and Azure Datalake Gen2, that facilitate interactions between both Azure Datalake implementations and Dask. This is done leveraging the intake/filesystem_spec base class and Azure Python SDKs.

Operations against both Gen1 Datalake currently only work with an Azure ServicePrincipal with suitable credentials to perform operations on the resources of choice.

Operations against the Gen2 Datalake are implemented by leveraging Azure Blob Storage Python SDK.

Setting credentials

The storage_options can be instantiated with a variety of keyword arguments depending on the filesystem. The most commonly used arguments are:

connection_string
account_name
account_key
sas_token
tenant_id, client_id, and client_secret are combined for an Azure ServicePrincipal e.g. storage_options={'account_name': ACCOUNT_NAME, 'tenant_id': TENANT_ID, 'client_id': CLIENT_ID, 'client_secret': CLIENT_SECRET}
anon: bool, optional. The value to use for whether to attempt anonymous access if no other credential is passed. By default (None), the AZURE_STORAGE_ANON environment variable is checked. False values (false, 0, f) will resolve to False and anonymous access will not be attempted. Otherwise the value for anon resolves to True.
location_mode: valid values are "primary" or "secondary" and apply to RA-GRS accounts

For more argument details see all arguments for AzureBlobFileSystem here and AzureDatalakeFileSystem here.

The following environmental variables can also be set and picked up for authentication:

"AZURE_STORAGE_CONNECTION_STRING"
"AZURE_STORAGE_ACCOUNT_NAME"
"AZURE_STORAGE_ACCOUNT_KEY"
"AZURE_STORAGE_SAS_TOKEN"
"AZURE_STORAGE_TENANT_ID"
"AZURE_STORAGE_CLIENT_ID"
"AZURE_STORAGE_CLIENT_SECRET"

The filesystem can be instantiated for different use cases based on a variety of storage_options combinations. The following list describes some common use cases utilizing AzureBlobFileSystem, i.e. protocols abfsor az. Note that all cases require the account_name argument to be provided:

Anonymous connection to public container: storage_options={'account_name': ACCOUNT_NAME, 'anon': True} will assume the ACCOUNT_NAME points to a public container, and attempt to use an anonymous login. Note, the default value for anon is True.
Auto credential solving using Azure's DefaultAzureCredential() library: storage_options={'account_name': ACCOUNT_NAME, 'anon': False} will use DefaultAzureCredential to get valid credentials to the container ACCOUNT_NAME. DefaultAzureCredential attempts to authenticate via the mechanisms and order visualized here.
Auto credential solving without requiring storage_options: Set AZURE_STORAGE_ANON to false, resulting in automatic credential resolution. Useful for compatibility with fsspec.
Azure ServicePrincipal: tenant_id, client_id, and client_secret are all used as credentials for an Azure ServicePrincipal: e.g. storage_options={'account_name': ACCOUNT_NAME, 'tenant_id': TENANT_ID, 'client_id': CLIENT_ID, 'client_secret': CLIENT_SECRET}.

Append Blob

The AzureBlobFileSystem accepts all of the Async BlobServiceClient arguments.

By default, write operations create BlockBlobs in Azure, which, once written can not be appended. It is possible to create an AppendBlob using mode="ab" when creating and operating on blobs. Currently, AppendBlobs are not available if hierarchical namespaces are enabled.

Project details

These details have not been verified by PyPI

Release history Release notifications | RSS feed

This version

2024.12.0

Dec 15, 2024

2024.7.0

Jul 22, 2024

2024.4.1

Apr 15, 2024

2024.4.0

Apr 13, 2024

2024.2.0

Feb 5, 2024

2024.1.0

Jan 29, 2024

2023.12.0

Dec 23, 2023

2023.10.0

Oct 17, 2023

2023.9.0

Sep 17, 2023

2023.8.0

Aug 8, 2023

2023.4.0

Apr 27, 2023

2023.1.0

Jan 17, 2023

2022.11.2

Nov 24, 2022

2022.11.1

Nov 24, 2022

2022.11.0 yanked

Nov 23, 2022

Reason this release was yanked:

AzureDatalakeFileSystem fails import

2022.10.0

Oct 3, 2022

2022.9.1

Sep 6, 2022

2022.9.0

Sep 6, 2022

2022.7.0

Jul 9, 2022

2022.4.0

Apr 15, 2022

2022.4.0a0 pre-release

Apr 15, 2022

2022.2.0

Feb 5, 2022

2021.10.0

Oct 3, 2021

2021.9.1

Sep 10, 2021

2021.8.2

Aug 18, 2021

2021.8.1

Aug 13, 2021

2021.7.1

Jul 19, 2021

2021.7.0 yanked

Jul 12, 2021

Reason this release was yanked:

Changing behavior with anonymous logins to public repos causes user issues

0.7.7

Jun 14, 2021

0.7.6

Jun 9, 2021

0.7.5

May 11, 2021

0.7.4

Apr 26, 2021

0.7.3

Apr 15, 2021

0.7.2

Apr 12, 2021

0.7.1

Apr 9, 2021

0.7.0

Mar 31, 2021

0.6.3

Feb 16, 2021

0.6.2

Feb 12, 2021

0.6.1

Feb 9, 2021

0.6.0

Jan 15, 2021

0.5.9

Dec 19, 2020

0.5.8

Dec 9, 2020

0.5.7

Nov 19, 2020

0.5.5

Oct 6, 2020

0.5.4

Oct 4, 2020

0.5.3

Sep 15, 2020

0.5.2 yanked

Sep 15, 2020

0.5.1

Sep 10, 2020

0.5.0

Sep 7, 2020

0.4.0

Aug 20, 2020

0.3.3

Aug 13, 2020

0.3.2

Aug 2, 2020

0.3.1

Jun 15, 2020

0.3.0

May 19, 2020

0.2.5

May 19, 2020

0.2.4

Apr 21, 2020

0.2.3

Apr 21, 2020

0.2.2

Apr 20, 2020

0.2.0

Feb 15, 2020

0.1.5

Dec 17, 2019

0.1.4

Dec 16, 2019

0.1.3

Dec 15, 2019

0.1.3a0 pre-release

Dec 16, 2019

0.1.2

Nov 25, 2019

0.1.1

Nov 14, 2019

0.1.0

Oct 20, 2019

0.0.11

Oct 15, 2019

0.0.10.post2

Oct 14, 2019

0.0.10.post1

Oct 14, 2019

0.0.10.post0

Oct 14, 2019

0.0.10

Oct 14, 2019

0.0.9.post0

Oct 9, 2019

0.0.9

Oct 9, 2019

0.0.8.post3

Oct 9, 2019

0.0.8.post2

Oct 9, 2019

0.0.8.post1

Oct 9, 2019

0.0.8.post0

Oct 9, 2019

0.0.8

Oct 9, 2019

0.0.8a0 pre-release

Oct 9, 2019

0.0.7

Sep 23, 2019

0.0.6

Sep 19, 2019

0.0.5

Sep 11, 2019

0.0.5a0 pre-release

Sep 18, 2019

0.0.2

Aug 11, 2019

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

adlfs-2024.12.0.tar.gz (49.2 kB view details)

Uploaded Dec 15, 2024 Source

Built Distribution

adlfs-2024.12.0-py3-none-any.whl (41.8 kB view details)

Uploaded Dec 15, 2024 Python 3

File details

Details for the file adlfs-2024.12.0.tar.gz.

File metadata

Download URL: adlfs-2024.12.0.tar.gz
Upload date: Dec 15, 2024
Size: 49.2 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.0.1 CPython/3.9.21

File hashes

Hashes for adlfs-2024.12.0.tar.gz
Algorithm	Hash digest
SHA256	`04582bf7461a57365766d01a295a0a88b2b8c42c4fea06e2d673f62675cac5c6`
MD5	`b6dbf1e784042037b3b11befff664b53`
BLAKE2b-256	`7682e30891af574fb358449fb9436aac53569814452cb88b0cba4f488171b8dc`

See more details on using hashes here.

File details

Details for the file adlfs-2024.12.0-py3-none-any.whl.

File metadata

Download URL: adlfs-2024.12.0-py3-none-any.whl
Upload date: Dec 15, 2024
Size: 41.8 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.0.1 CPython/3.9.21

File hashes

Hashes for adlfs-2024.12.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`00aab061ddec0413b2039487e656b62e01ece8ef1ca0493f76034a596cf069e3`
MD5	`a5d99660e17ed7627f238602817b8563`
BLAKE2b-256	`cbedd1bf75c089857d38332cf45416e419b47382b345ba5dfc4fae69397830d9`

See more details on using hashes here.

adlfs 2024.12.0

Navigation

Verified details

Maintainers

Unverified details

Meta

Classifiers

Project description

Filesystem interface to Azure-Datalake Gen1 and Gen2 Storage

Quickstart

Details

Setting credentials

Append Blob

Project details

Verified details

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes