Skip to main content

Simple. Streaming. Resilient. MFA-ready. Fetch files from SharePoint via Microsoft Graph.

Project description

🚀 spfetch

spfetch_lg

Simple. Streaming. Resilient. MFA-ready.
List and fetch files from SharePoint via Microsoft Graph with clean APIs and cloud-native downloads.


✨ What is spfetch?

spfetch is an asynchronous Python library built for modern data pipelines:

  • 📂 List SharePoint folders with structured metadata
  • ⬇️ Stream large files directly to Local Disk, S3, GCS, or Azure without memory crashes
  • Smart Buffering – Control chunk and buffer sizes to optimize Cloud I/O (50+ MB/s)
  • 📊 Load small files directly into Pandas DataFrames
  • 🔐 Authenticate via MFA (Device Code) or Silent (Client Secret) flows
  • 🛡️ Auto-Recover from Microsoft API Throttling (HTTP 429) with Exponential Backoff

🚀 Performance Benchmark (v0.1.3)

Zero Intermediate Disk Architecture + Smart Buffering

Benchmark Results

  • Payload: 10.10 GB CSV (SharePoint ➡ Azure Data Lake)
  • Time: 3m 11s (191.96s)
  • Average Speed: 53.87 MB/s
  • Config: chunk_size_mb=1 | buffer_size_mb=100

🏗️ Technical Architecture

The library is designed with a layered approach to ensure high throughput and resilience. By decoupling the reading rate from the writing rate, we maximize the performance of both the Microsoft Graph API and Cloud Providers.

diagram-spfetch

The Data Pipeline Flow:

  1. Source (SharePoint): Chunks are read at a light rate (default 1MB) to avoid API throttling.
  2. Core (Smart Buffer): Data is accumulated in a memory buffer managed by the Orchestrator.
  3. Destination (Cloud): Once the buffer reaches the set size (e.g., 100MB), a single high-speed write is performed via fsspec.
  4. Resilience: The @retry_on_429 shield monitors all requests, applying exponential backoff if the source is overloaded.

🔐 1. Authentication

Instantiate the client using your Microsoft Entra ID (Azure AD) credentials.


Option A: Interactive / Local (Device Code Flow)

Ideal for local scripts. Supports MFA.

from spfetch.auth import DeviceCodeAuth
from spfetch.client import SharePointClient

auth = DeviceCodeAuth(
    tenant_id="<YOUR_TENANT_ID>",
    client_id="<YOUR_CLIENT_ID>"
)

client = SharePointClient(auth=auth)

Option B: Automated / CI/CD (Client Secret Flow)

Ideal for Airflow, Databricks, GitHub Actions.

from spfetch.auth import ClientSecretAuth
from spfetch.client import SharePointClient

auth = ClientSecretAuth(
    tenant_id="<YOUR_TENANT_ID>",
    client_id="<YOUR_CLIENT_ID>",
    client_secret="<YOUR_CLIENT_SECRET>"
)

client = SharePointClient(auth=auth)

📊 2. Telemetry & Dual Progress Bar

By default, spfetch does not override your logging configuration (uses NullHandler).

To enable structured logs and dual progress bars:

import asyncio
from spfetch.auth import ClientSecretAuth
from spfetch.client import SharePointClient
from spfetch.destinations import LocalDestination
from spfetch import enable_console_logs # <-- Add this line

enable_console_logs() # <-- Add this line

async def main():
    auth = ClientSecretAuth(
        tenant_id="YOUR_TENANT_ID",
        client_id="YOUR_CLIENT_ID",
        client_secret="YOUR_CLIENT_SECRET"
    )

    client = SharePointClient(auth=auth)
    destination = LocalDestination()

    await client.download(
        hostname="your_company.sharepoint.com",
        site_path="/sites/YourSite",
        file_path="/Folder/your_file.csv",
        dest_path="./data/your_file.csv",
        destination=destination
    )

if __name__ == "__main__":
    asyncio.run(main())

🖥️ Expected Terminal Output

🚀 [INGESTÃO STREAMING | STREAMING INGESTION] Started at: YYYY-MM-DD HH:MM:SS
📍 Destination: <DestinationClass> -> path/to/file.ext (Chunk: 1MB | Buffer: 16MB)
📂 Source: /path/in/sharepoint/file.ext (X.XX GB)

📥 Reading | Leitura: 100%|███████████████| X.XXG/X.XXG [MM:SS<00:00, XX.XMB/s]
📤 Saving | Salvando: 100%|███████████████| X.XXG/X.XXG [MM:SS<00:00, XX.XMB/s]

-------------------------------------------------------
✅ INGESTION COMPLETED SUCCESSFULLY
⏱ Total Time: XXX.XXs
⚡ Average Speed: XX.XX MB/s
🏁 Finished at: YYYY-MM-DD HH:MM:SS
-------------------------------------------------------

📖 3. Exploration – Listing Folders

📦 Installation:

pip install spfetch
import asyncio

async def list_files():
    items = await client.ls(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        folder_path="/Shared Documents/General"
    )

    for item in items:
        print(item["name"], item["size"], item["is_folder"])

asyncio.run(list_files())

🌊 4. Ingestion Workflows


☁️ A) Azure (ADLS / Blob)

📦 Installation:

pip install "spfetch[azure]"
from spfetch.destinations import AzureDestination
import asyncio

async def download_to_azure():
    destination = AzureDestination(
        account_name="<YOUR_STORAGE_ACCOUNT_NAME>",
        account_key="<YOUR_STORAGE_ACCOUNT_KEY>"
    )

    await client.download(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        file_path="/Shared Documents/Data/file.parquet",
        dest_path="abfs://<container>/bronze/file.parquet",
        destination=destination,
        chunk_size_mb=1,
        buffer_size_mb=100
    )

asyncio.run(download_to_azure())

☁️ B) Amazon S3

📦 Installation:

pip install "spfetch[s3]"
from spfetch.destinations import S3Destination
import asyncio

async def download_to_s3():
    destination = S3Destination(
        key="<AWS_ACCESS_KEY_ID>",
        secret="<AWS_SECRET_ACCESS_KEY>"
    )

    await client.download(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        file_path="/Shared Documents/Data/file.csv",
        dest_path="s3://<bucket>/raw/file.csv",
        destination=destination,
        chunk_size_mb=1,
        buffer_size_mb=16
    )

asyncio.run(download_to_s3())

☁️ C) Google Cloud Storage (GCS)

📦 Installation:

pip install "spfetch[gcs]"
from spfetch.destinations import GCSDestination
import asyncio

async def download_to_gcs():
    destination = GCSDestination(
        project="<my-gcp-project-id>",
        token="google_default"
    )

    await client.download(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        file_path="/Shared Documents/Data/file.csv",
        dest_path="gs://<bucket>/raw/file.csv",
        destination=destination
    )

asyncio.run(download_to_gcs())

💻 D) Local Disk

📦 Installation:

pip install spfetch
from spfetch.destinations import LocalDestination
import asyncio

async def download_local():
    destination = LocalDestination()

    await client.download(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        file_path="/Shared Documents/Data/file.csv",
        dest_path="./local_downloads/file.csv",
        destination=destination
    )

asyncio.run(download_local())

📊 E) Read Directly to Pandas

📦 Installation:

pip install "spfetch[pandas]"
import asyncio

async def read_to_memory():
    df = await client.read_df(
        hostname="<tenant>.sharepoint.com",
        site_path="/sites/<YourSite>",
        file_path="/Shared Documents/Reports/data.xlsx",
        sheet_name="Sheet1",
        skiprows=2,
        usecols="A:D"
    )

    print(df.head())

asyncio.run(read_to_memory())

🛡️ 5. Resilience – Handling HTTP 429

spfetch automatically handles Microsoft Graph throttling.

If HTTP 429 Too Many Requests occurs:

  1. Execution pauses
  2. Retry-After header is read
  3. Exponential Backoff is applied
  4. Retries up to 5 times

Your pipeline will wait and recover gracefully instead of crashing.


🤝 Contributing

Pull Requests are welcome.

Before submitting:

make test

Ensure all tests pass.


📄 License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

spfetch-0.1.4.tar.gz (17.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

spfetch-0.1.4-py3-none-any.whl (13.8 kB view details)

Uploaded Python 3

File details

Details for the file spfetch-0.1.4.tar.gz.

File metadata

  • Download URL: spfetch-0.1.4.tar.gz
  • Upload date:
  • Size: 17.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for spfetch-0.1.4.tar.gz
Algorithm Hash digest
SHA256 90de0ba70dcf0717b6df131249585e312dfee069b40d3787913747e17031445e
MD5 8e12c228be6de647eb88806cdf0afe67
BLAKE2b-256 42fbd9f28b6e515f2af60513931b12dcfe31de51b6e49a6a0f36910b638e3409

See more details on using hashes here.

File details

Details for the file spfetch-0.1.4-py3-none-any.whl.

File metadata

  • Download URL: spfetch-0.1.4-py3-none-any.whl
  • Upload date:
  • Size: 13.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for spfetch-0.1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 572a9b135c9f015572d5771e65c68815696fdd904efa4a4a82df81ee2b8fb1b8
MD5 1f7aa34260c19324469df82602729132
BLAKE2b-256 ea8361b2b21ee35d7796e8a08f36adabefe6cedf208fd26b650c4dac6902baa7

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page