Python SDK for Unsiloed Vision API - Parse, Extract, Classify, and Split documents
Project description
Unsiloed Python SDK
The official Python SDK for Unsiloed - a powerful document processing platform that enables you to parse, extract, classify, and split documents with ease.
Features
- Parse: Extract structured content from documents (PDFs, images, etc.)
- Extract: Extract specific data using JSON schemas with citations
- Classify: Classify documents into predefined categories
- Split: Split document pages based on categories
- Both Sync & Async: Choose the client that fits your application
Installation
pip install unsiloed-sdk
Quick Start
Synchronous Usage
Perfect for scripts, notebooks, and traditional Python applications:
from unsiloed_sdk import UnsiloedClient
with UnsiloedClient(api_key="your-api-key-here") as client:
result = client.parse_and_wait(file="document.pdf")
print(f"Total chunks: {result.total_chunks}")
Async Usage
Perfect for FastAPI apps and concurrent processing:
import asyncio
from unsiloed_sdk import AsyncUnsiloedClient
async def main():
async with AsyncUnsiloedClient(api_key="your-api-key-here") as client:
result = await client.parse_and_wait(file="document.pdf")
print(f"Total chunks: {result.total_chunks}")
asyncio.run(main())
Authentication
Get your API key from the Unsiloed Dashboard.
from unsiloed_sdk import UnsiloedClient
client = UnsiloedClient(api_key="your-api-key-here")
Set as environment variable:
export UNSILOED_API_KEY="your-api-key"
import os
from unsiloed_sdk import UnsiloedClient
client = UnsiloedClient(api_key=os.getenv("UNSILOED_API_KEY"))
Usage Examples
Parse Documents
Extract structured content from any document:
from unsiloed_sdk import UnsiloedClient
# Initialize the client
with UnsiloedClient(api_key="your-api-key") as client:
# Parse a document and wait for results
result = client.parse_and_wait(file="document.pdf")
# Access the parsed content
print(f"Total chunks: {result.total_chunks}")
# Get the embed content
for chunk in result.chunks:
print(f"\n--- {chunk['embed'][:100]} ---")
Extract Data with Schema
Define exactly what data you need using JSON schema. The property names are the fields you want to extract, and descriptions specify what type of data to look for:
from unsiloed_sdk import UnsiloedClient
schema = {
"type": "object",
"properties": {
"invoice_number": {
"type": "string",
"description": "Invoice number from the document"
},
"date": {
"type": "string",
"description": "Invoice date"
},
"total_amount": {
"type": "number",
"description": "Total amount"
}
},
"required": ["invoice_number", "date", "total_amount"],
"additionalProperties": False
}
with UnsiloedClient(api_key="your-api-key") as client:
result = client.extract_and_wait(
file="invoice.pdf",
schema=schema
)
# Results include confidence scores
print(f"Invoice #: {result.result['invoice_number']['value']}")
print(f"Confidence: {result.result['invoice_number']['score']}")
print(f"Total: ${result.result['total_amount']['value']}")
Advanced Example - Extracting shareholding data:
from unsiloed_sdk import UnsiloedClient
schema = {
"type": "object",
"properties": {
"Individuals": {
"type": "string",
"description": "Percentage Holding"
},
"LIC of India": {
"type": "string",
"description": "No of Shares Held"
},
"United bank of india": {
"type": "string",
"description": "No of shares held by United bank of india"
}
},
"required": ["Individuals", "LIC of India", "United bank of india"],
"additionalProperties": False
}
with UnsiloedClient(api_key="your-api-key") as client:
result = client.extract_and_wait(
file="shareholding.pdf",
schema=schema
)
for field, data in result.result.items():
print(f"{field}: {data['value']} (confidence: {data['score']:.2%})")
Classify Documents
Automatically categorize your documents:
from unsiloed_sdk import UnsiloedClient
with UnsiloedClient(api_key="your-api-key") as client:
result = client.classify_and_wait(
file="document.pdf",
categories=["Invoice", "Receipt", "Contract", "Letter"]
)
print(f"Type: {result.result['classification']}")
print(f"Confidence: {result.result['confidence']}")
Split Documents
Separate multi-document files by page type:
from unsiloed_sdk import UnsiloedClient, Category
categories = [
Category(name="Cover Page", description="Document cover or title page"),
Category(name="Main Content", description="Primary document content and body text")
]
with UnsiloedClient(api_key="your-api-key") as client:
result = client.split_and_wait(
file="report.pdf",
categories=categories
)
# Check if split was successful
if result.result['success']:
print(f"✓ {result.result['message']}")
# Access the generated split files
for file_info in result.result['files']:
print(f"File: {file_info['name']}")
print(f" Confidence: {file_info['confidence_score']:.2%}")
print(f" Download: {file_info['full_path']}")
else:
print(f"Split failed: {result.result['message']}")
Async Examples
Concurrent Processing
Process multiple documents at once with async:
import asyncio
from unsiloed_sdk import AsyncUnsiloedClient
async def main():
async with AsyncUnsiloedClient(api_key="your-api-key") as client:
# Process 3 documents concurrently
results = await asyncio.gather(
client.parse_and_wait(file="doc1.pdf"),
client.parse_and_wait(file="doc2.pdf"),
client.parse_and_wait(file="doc3.pdf"),
)
for i, result in enumerate(results, 1):
print(f"Document {i}: {result.total_chunks} chunks")
asyncio.run(main())
Async Extract
import asyncio
from unsiloed_sdk import AsyncUnsiloedClient
async def main():
schema = {
"type": "object",
"properties": {
"company": {
"type": "string",
"description": "Company name"
},
"amount": {
"type": "number",
"description": "Total amount"
}
},
"required": ["company", "amount"],
"additionalProperties": False
}
async with AsyncUnsiloedClient(api_key="your-api-key") as client:
result = await client.extract_and_wait(
file="invoice.pdf",
schema=schema
)
# Access extracted values with confidence scores
print(f"Company: {result.result['company']['value']}")
print(f"Amount: {result.result['amount']['value']}")
asyncio.run(main())
Error Handling
from unsiloed_sdk import (
UnsiloedClient,
AuthenticationError,
QuotaExceededError,
InvalidRequestError
)
try:
with UnsiloedClient(api_key="your-api-key") as client:
result = client.parse_and_wait(file="document.pdf")
except AuthenticationError:
print("Invalid API key")
except QuotaExceededError as e:
print(f"Quota exceeded. Remaining: {e.response_data}")
except InvalidRequestError as e:
print(f"Invalid request: {e.message}")
Which Client Should I Use?
| Use Case | Client | Example |
|---|---|---|
| Scripts, notebooks | UnsiloedClient |
client.parse_and_wait(file="doc.pdf") |
| Flask, Django | UnsiloedClient |
client.extract_and_wait(file="invoice.pdf", schema=schema) |
| FastAPI, async apps | AsyncUnsiloedClient |
await client.classify_and_wait(file="doc.pdf", categories=[...]) |
| Concurrent processing | AsyncUnsiloedClient |
await asyncio.gather(...) |
API Methods
Both UnsiloedClient (sync) and AsyncUnsiloedClient (async) have the same methods:
Parse:
parse()- Start parse jobget_parse_result(job_id)- Check job statusparse_and_wait()- Parse and wait for completion ⭐ Most common
Extract:
extract()- Start extract jobget_extract_result(job_id)- Check job statusextract_and_wait()- Extract and wait for completion ⭐ Most common
Classify:
classify()- Start classify jobget_classify_result(job_id)- Check job statusclassify_and_wait()- Classify and wait for completion ⭐ Most common
Split:
split()- Start split jobget_split_result(job_id)- Check job statussplit_and_wait()- Split and wait for completion ⭐ Most common
Tip: Use the
*_and_wait()methods for simpler code. They handle polling automatically.
Examples
Check the examples/ directory for complete working examples:
parse_example.py- Document parsingextract_example.py- Data extractionclassify_example.py- Document classificationsplit_example.py- Document splitting
Support
- Documentation: https://docs.unsiloed.ai/
- API Reference: https://docs.unsiloed.ai/api-reference
- Support Email: support@unsiloed.com
- Issues: https://github.com/unsiloed/unsiloed-sdk-python/issues
License
MIT License - see LICENSE file for details.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file unsiloed_sdk-0.1.4.tar.gz.
File metadata
- Download URL: unsiloed_sdk-0.1.4.tar.gz
- Upload date:
- Size: 102.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5a418d29c7c82b67beb1b9e2b462dbd48c43a362618fe6bddbe56f149557d792
|
|
| MD5 |
82ff3bebd6e6d454a19485d08557a32e
|
|
| BLAKE2b-256 |
dc1a7c5b6b12ce2c7588bd9089b17e4f5714c9867f4f2e7d71d26c5dbf97ede8
|
File details
Details for the file unsiloed_sdk-0.1.4-py3-none-any.whl.
File metadata
- Download URL: unsiloed_sdk-0.1.4-py3-none-any.whl
- Upload date:
- Size: 13.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7035f225dd8638611c3bdd58eb8dfae38b34acb487dafeb62f8717447a972d62
|
|
| MD5 |
26833b7e32c342b7b72adff155e3ee74
|
|
| BLAKE2b-256 |
e104d70ba7c0f7a4db848c98b7e7cc92b4cab417b7d5a92f5fbb1ac24e6ace15
|