Simple. Streaming. Resilient. MFA-ready. Fetch files from SharePoint via Microsoft Graph.
Project description
🚀 spfetch
Simple. Streaming. MFA-ready.
List and fetch files from SharePoint via Microsoft Graph with clean APIs and cloud-native downloads.
✨ What is spfetch?
spfetch is a Python library to:
- 📂 List SharePoint folders
- ⬇️ Download files in streaming mode
- ☁️ Send files directly to S3 / GCS / ADLS
- 🔐 Authenticate with MFA (Device Code Flow)
- 📊 Optionally read small files into pandas
Built for data engineers and analytics workflows that need reliability and simplicity.
🤔 Why?
Accessing SharePoint programmatically usually involves:
- 🔐 Complex authentication flows (MFA included)
- 🔎 Confusing browser URLs vs API paths
- 🚦 Throttling (HTTP 429)
- 🧠 Memory issues when handling large files
spfetch solves this with:
- ✅ MFA-compatible authentication
- ✅ Streaming-first downloads (no full in-memory load)
- ✅ fsspec integration (cloud-native destinations)
- ✅ Minimal and explicit API
📦 Installation
Core
pip install spfetch
☁️ Cloud destinations (optional)
pip install spfetch[s3] # Amazon S3 (s3fs)
pip install spfetch[gcs] # Google Cloud Storage (gcsfs)
pip install spfetch[azure] # Azure Data Lake (adlfs)
📊 Pandas helpers (optional)
pip install spfetch[pandas]
🚀 Quickstart
Since spfetch is built for modern data engineering, all network operations are asynchronous.
🔐 1️⃣ Authenticate (Device Code / MFA)
For Local Development (Device Code / MFA) Works with MFA-enabled accounts. No secrets required.
from spfetch.auth import DeviceCodeAuth
from spfetch.client import SharePointClient
auth = DeviceCodeAuth(
tenant_id="YOUR_TENANT_ID",
client_id="YOUR_CLIENT_ID"
)
client = SharePointClient(auth=auth)
✔️ Works with MFA-enabled accounts
✔️ No secrets required
✔️ Ideal for local development
For CI/CD & Automated Pipelines (Client Secret)
from spfetch.auth import ClientSecretAuth
auth = ClientSecretAuth(
tenant_id="YOUR_TENANT_ID",
client_id="YOUR_CLIENT_ID",
client_secret="YOUR_CLIENT_SECRET"
)
client = SharePointClient(auth=auth)
📂 2️⃣ List a Folder
import asyncio
async def main():
items = await client.ls(
hostname="tenant.sharepoint.com",
site_path="/sites/MySite",
folder_path="/Shared Documents/Reports"
)
for item in items:
icon = "📁" if item["is_folder"] else "📄"
print(f"{icon} {item['name']} | Size: {item['size']} | ID: {item['id']}")
asyncio.run(main())
Returns structured metadata for files and folders.
⬇️ 3️⃣ Streaming Download (Recommended for Large Files)
🔥 Files are streamed in chunks — no full in-memory loading. Perfect for large CSVs, parquet files, and data pipelines.
🖥️ Download to Local Filesystem
async def main():
await client.download(
hostname="tenant.sharepoint.com",
site_path="/sites/MySite",
file_path="/Shared Documents/Big/base.csv",
dest_path="stage/base.csv"
)
☁️ Download Directly to Cloud Storage
from spfetch.destinations import S3Destination
# from spfetch.destinations import AzureDestination, GCSDestination
async def main():
# Pass your cloud credentials configuration
s3_dest = S3Destination(key="AWS_KEY", secret="AWS_SECRET")
await client.download(
hostname="tenant.sharepoint.com",
site_path="/sites/MySite",
file_path="/Shared Documents/Big/base.csv",
dest_path="s3://my-bucket/stage/base.csv",
destination=s3_dest
)
🔥 Files are streamed in chunks — no full in-memory loading.
Perfect for large CSVs, parquet files, exports, and data pipelines.
📊 4️⃣ Read Small Files into pandas (Optional)
read_df automatically detects .csv, .xlsx, or .xls and loads it into memory without writing to disk.
async def main():
df = await client.read_df(
hostname="tenant.sharepoint.com",
site_path="/sites/MySite",
file_path="/Shared Documents/Reports/sales.xlsx",
sheet_name="Base", # Passed down to pandas
usecols="B:F",
skiprows=8
)
print(df.head())
⚠️ For very large files, prefer
download()and process them using Spark, Dask, or your data engine of choice.
🛡️ Resilience (Handling HTTP 429)
spfetch includes a built-in resilience layer. If Microsoft Graph throws an HTTP 429 Too Many Requests error, the client will automatically:
- Catch the error.
- Read the
Retry-Afterheader provided by Microsoft. - Pause execution asynchronously using Exponential Backoff.
- Retry the request automatically up to
max_retries(Default: 3 for API calls, 5 for downloads).
Your pipelines will simply wait and recover gracefully.
🔐 Microsoft Entra ID (Azure AD) Setup
To use spfetch, create an App Registration in Microsoft Entra ID.
🧭 Setup Steps
- Create an App Registration
- Copy the
tenant_id - Copy the
client_id - Enable Public client flows (Device Code flow)
- Grant required Microsoft Graph permissions
- (If needed) Request Admin Consent
🔑 Required Permissions
Minimum permissions depend on your use case.
📖 Read-only access
Sites.Read.All
✍️ Read & Write access
Sites.ReadWrite.All
Some tenants may require:
- Admin consent
- Site-specific permission configuration
🧱 Design Principles
- 🔐 MFA-first authentication
- 🌊 Streaming over buffering
- ☁️ Cloud-native architecture
- 🧩 Minimal API surface
- 🔎 Explicit behavior over magic
🗺️ Roadmap
Planned features:
- 🔑
connect_client_secret()for pipelines / CI - 🔄
sync_folder()with ETag / Last-Modified support - 🧭 Path and URL normalization helpers
- 📊 Richer metadata filtering
- 🔁 Configurable retry / backoff strategy
🤝 Contributing
PRs are welcome!
Please:
- Check our
CONTRIBUTING.md. - Ensure all tests pass
via make test. - We use the Conventional Commits specification.
📄 License
MIT
Built for modern data workflows 🚀
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file spfetch-0.1.0.tar.gz.
File metadata
- Download URL: spfetch-0.1.0.tar.gz
- Upload date:
- Size: 16.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5d3352ee49b28f49d67c1bca4a94ebdaf46e4f052677a5f67ead307509913e90
|
|
| MD5 |
0765086e15795910c8e287f0f21339a1
|
|
| BLAKE2b-256 |
81f4c300a9a6bb588ad69a1df5d67551329b1bd2aa9638966d9a33ef8fe3b08d
|
File details
Details for the file spfetch-0.1.0-py3-none-any.whl.
File metadata
- Download URL: spfetch-0.1.0-py3-none-any.whl
- Upload date:
- Size: 13.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d730b3259cc2efab3bf7c444fc98f27cb657721c1e9b14d2a2b9665f8939726f
|
|
| MD5 |
a12c612d347a148642be69bf729579ab
|
|
| BLAKE2b-256 |
80752de1c0436181c8bfbb607f4d85c3bee05520e0db447aa82e48e814fdf745
|