Simple. Streaming. Resilient. MFA-ready. Fetch files from SharePoint via Microsoft Graph.
Project description
🚀 spfetch
Simple. Streaming. Resilient. MFA-ready.
List and fetch files from SharePoint via Microsoft Graph with clean APIs and cloud-native downloads.
✨ What is spfetch?
spfetch is an asynchronous Python library built for modern data pipelines:
- 📂 List SharePoint folders with structured metadata
- ⬇️ Stream large files directly to Local Disk, S3, GCS, or Azure without memory crashes
- ⚡ Smart Buffering – Control chunk and buffer sizes to optimize Cloud I/O (50+ MB/s)
- 📊 Load small files directly into Pandas DataFrames
- 🔐 Authenticate via MFA (Device Code) or Silent (Client Secret) flows
- 🛡️ Auto-Recover from Microsoft API Throttling (HTTP 429) with Exponential Backoff
🚀 Performance Benchmark (v0.1.3)
Zero Intermediate Disk Architecture + Smart Buffering
Benchmark Results
- Payload: 10.10 GB CSV (SharePoint ➡ Azure Data Lake)
- Time: 3m 11s (191.96s)
- Average Speed: 53.87 MB/s
- Config:
chunk_size_mb=1|buffer_size_mb=100
🏗️ Technical Architecture
The library is designed with a layered approach to ensure high throughput and resilience. By decoupling the reading rate from the writing rate, we maximize the performance of both the Microsoft Graph API and Cloud Providers.
The Data Pipeline Flow:
- Source (SharePoint): Chunks are read at a light rate (default 1MB) to avoid API throttling.
- Core (Smart Buffer): Data is accumulated in a memory buffer managed by the Orchestrator.
- Destination (Cloud): Once the buffer reaches the set size (e.g., 100MB), a single high-speed write is performed via
fsspec. - Resilience: The
@retry_on_429shield monitors all requests, applying exponential backoff if the source is overloaded.
🔐 1. Authentication
Instantiate the client using your Microsoft Entra ID (Azure AD) credentials.
Option A: Interactive / Local (Device Code Flow)
Ideal for local scripts. Supports MFA.
from spfetch.auth import DeviceCodeAuth
from spfetch.client import SharePointClient
auth = DeviceCodeAuth(
tenant_id="<YOUR_TENANT_ID>",
client_id="<YOUR_CLIENT_ID>"
)
client = SharePointClient(auth=auth)
Option B: Automated / CI/CD (Client Secret Flow)
Ideal for Airflow, Databricks, GitHub Actions.
from spfetch.auth import ClientSecretAuth
from spfetch.client import SharePointClient
auth = ClientSecretAuth(
tenant_id="<YOUR_TENANT_ID>",
client_id="<YOUR_CLIENT_ID>",
client_secret="<YOUR_CLIENT_SECRET>"
)
client = SharePointClient(auth=auth)
📊 2. Telemetry & Dual Progress Bar
By default, spfetch does not override your logging configuration (uses NullHandler).
To enable structured logs and dual progress bars:
import asyncio
from spfetch.auth import ClientSecretAuth
from spfetch.client import SharePointClient
from spfetch.destinations import LocalDestination
from spfetch import enable_console_logs # <-- Add this line
enable_console_logs() # <-- Add this line
async def main():
auth = ClientSecretAuth(
tenant_id="YOUR_TENANT_ID",
client_id="YOUR_CLIENT_ID",
client_secret="YOUR_CLIENT_SECRET"
)
client = SharePointClient(auth=auth)
destination = LocalDestination()
await client.download(
hostname="your_company.sharepoint.com",
site_path="/sites/YourSite",
file_path="/Folder/your_file.csv",
dest_path="./data/your_file.csv",
destination=destination
)
if __name__ == "__main__":
asyncio.run(main())
🖥️ Expected Terminal Output
🚀 [INGESTÃO STREAMING | STREAMING INGESTION] Started at: YYYY-MM-DD HH:MM:SS
📍 Destination: <DestinationClass> -> path/to/file.ext (Chunk: 1MB | Buffer: 16MB)
📂 Source: /path/in/sharepoint/file.ext (X.XX GB)
📥 Reading | Leitura: 100%|███████████████| X.XXG/X.XXG [MM:SS<00:00, XX.XMB/s]
📤 Saving | Salvando: 100%|███████████████| X.XXG/X.XXG [MM:SS<00:00, XX.XMB/s]
-------------------------------------------------------
✅ INGESTION COMPLETED SUCCESSFULLY
⏱ Total Time: XXX.XXs
⚡ Average Speed: XX.XX MB/s
🏁 Finished at: YYYY-MM-DD HH:MM:SS
-------------------------------------------------------
📖 3. Exploration – Listing Folders
📦 Installation:
pip install spfetch
import asyncio
async def list_files():
items = await client.ls(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
folder_path="/Shared Documents/General"
)
for item in items:
print(item["name"], item["size"], item["is_folder"])
asyncio.run(list_files())
🌊 4. Ingestion Workflows
☁️ A) Azure (ADLS / Blob)
📦 Installation:
pip install "spfetch[azure]"
from spfetch.destinations import AzureDestination
import asyncio
async def download_to_azure():
destination = AzureDestination(
account_name="<YOUR_STORAGE_ACCOUNT_NAME>",
account_key="<YOUR_STORAGE_ACCOUNT_KEY>"
)
await client.download(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
file_path="/Shared Documents/Data/file.parquet",
dest_path="abfs://<container>/bronze/file.parquet",
destination=destination,
chunk_size_mb=1,
buffer_size_mb=100
)
asyncio.run(download_to_azure())
☁️ B) Amazon S3
📦 Installation:
pip install "spfetch[s3]"
from spfetch.destinations import S3Destination
import asyncio
async def download_to_s3():
destination = S3Destination(
key="<AWS_ACCESS_KEY_ID>",
secret="<AWS_SECRET_ACCESS_KEY>"
)
await client.download(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
file_path="/Shared Documents/Data/file.csv",
dest_path="s3://<bucket>/raw/file.csv",
destination=destination,
chunk_size_mb=1,
buffer_size_mb=16
)
asyncio.run(download_to_s3())
☁️ C) Google Cloud Storage (GCS)
📦 Installation:
pip install "spfetch[gcs]"
from spfetch.destinations import GCSDestination
import asyncio
async def download_to_gcs():
destination = GCSDestination(
project="<my-gcp-project-id>",
token="google_default"
)
await client.download(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
file_path="/Shared Documents/Data/file.csv",
dest_path="gs://<bucket>/raw/file.csv",
destination=destination
)
asyncio.run(download_to_gcs())
💻 D) Local Disk
📦 Installation:
pip install spfetch
from spfetch.destinations import LocalDestination
import asyncio
async def download_local():
destination = LocalDestination()
await client.download(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
file_path="/Shared Documents/Data/file.csv",
dest_path="./local_downloads/file.csv",
destination=destination
)
asyncio.run(download_local())
📊 E) Read Directly to Pandas
📦 Installation:
pip install "spfetch[pandas]"
import asyncio
async def read_to_memory():
df = await client.read_df(
hostname="<tenant>.sharepoint.com",
site_path="/sites/<YourSite>",
file_path="/Shared Documents/Reports/data.xlsx",
sheet_name="Sheet1",
skiprows=2,
usecols="A:D"
)
print(df.head())
asyncio.run(read_to_memory())
🛡️ 5. Resilience – Handling HTTP 429
spfetch automatically handles Microsoft Graph throttling.
If HTTP 429 Too Many Requests occurs:
- Execution pauses
Retry-Afterheader is read- Exponential Backoff is applied
- Retries up to 5 times
Your pipeline will wait and recover gracefully instead of crashing.
🤝 Contributing
Pull Requests are welcome.
Before submitting:
make test
Ensure all tests pass.
📄 License
MIT License
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file spfetch-0.1.4.tar.gz.
File metadata
- Download URL: spfetch-0.1.4.tar.gz
- Upload date:
- Size: 17.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
90de0ba70dcf0717b6df131249585e312dfee069b40d3787913747e17031445e
|
|
| MD5 |
8e12c228be6de647eb88806cdf0afe67
|
|
| BLAKE2b-256 |
42fbd9f28b6e515f2af60513931b12dcfe31de51b6e49a6a0f36910b638e3409
|
File details
Details for the file spfetch-0.1.4-py3-none-any.whl.
File metadata
- Download URL: spfetch-0.1.4-py3-none-any.whl
- Upload date:
- Size: 13.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
572a9b135c9f015572d5771e65c68815696fdd904efa4a4a82df81ee2b8fb1b8
|
|
| MD5 |
1f7aa34260c19324469df82602729132
|
|
| BLAKE2b-256 |
ea8361b2b21ee35d7796e8a08f36adabefe6cedf208fd26b650c4dac6902baa7
|