Skip to main content

PyStreamPDF

Reduce RAG costs 50-70%. Extract only what matters from PDFs.

Stop sending entire documents to LLMs. PyStreamPDF analyzes structure, identifies relevant sections, and extracts only critical content. Cut token costs 50-70% while improving retrieval accuracy.

PyPI Python 3.10+ Tests: 523 Passing License: Proprietary


30-Second Start

from pystreampdf import Document

# Extract only relevant content from PDF
doc = Document("financial_report.pdf")

# Smart content extraction
relevant = doc.extract_relevant("revenue for Q3 2024")
print(f"Extracted {len(relevant)} chunks")
print(f"Token savings: {relevant.token_savings:.0%}")  # 60% savings

# Send only relevant parts to LLM
for chunk in relevant:
    response = llm.query(chunk, "What was Q3 revenue?")

Why PyStreamPDF?

The Problem:

  • You send entire PDFs to LLMs (wasteful, expensive)
  • RAG systems retrieve too much content
  • Token costs skyrocket on large documents
  • No way to know which parts actually matter

The Solution:

  • Intelligent document analysis finds relevant sections
  • Semantic chunking with context awareness
  • 50-70% reduction in token usage
  • Better retrieval accuracy (less noise)

Key Features

  • Intelligent Extraction: Find relevant sections automatically
  • Semantic Chunking: Context-aware splitting, not just word count
  • Multi-Format: Text, tables, images, charts, OCR
  • Token Budgeting: Allocate tokens by document type
  • Smart Caching: L1 memory + L2 disk (avoid reprocessing)
  • Metadata Preservation: Keep tables, images, structure
  • Production-Ready: 523 tests, type-safe API

Real-World Use Cases

Financial Documents:

# Extract relevant sections from annual report
doc = Document("10-K_2024.pdf")
revenue_sections = doc.extract_relevant("revenue")
earnings_sections = doc.extract_relevant("earnings")

# 70% fewer tokens than sending whole PDF
for section in revenue_sections:
    summary = llm.query(section, "What was total revenue?")

Legal Contracts:

# Find clauses without reading everything
doc = Document("contract.pdf")
liability = doc.extract_relevant("liability", "indemnification")
print(f"Found in {len(liability)} sections")

Research Papers:

# Extract methodology and results
doc = Document("paper.pdf")
methods = doc.extract_relevant("methods", "experiment")
results = doc.extract_relevant("results", "findings")

Token Savings

Document Size Full PDF Tokens PyStreamPDF Savings
Annual Report 200 pages 50K 15K 70%
Contract 50 pages 12K 4K 67%
Research Paper 30 pages 8K 2K 75%

Results: Lower costs + better retrieval accuracy + faster responses


Installation

pip install pystreampdf
# or with uv
uv pip install pystreampdf

Documentation


License

Proprietary License - Free to use with explicit attribution. See LICENSE.


PyStreamPDF v2.1.0 | Intelligent PDF processing for AI | Python 3.10+ | 523 tests


See [TOKEN_BUDGET_MULTIPLIERS.md](docs/TOKEN_BUDGET_MULTIPLIERS.md) for comprehensive guide.

## Quick Start: Document Extraction

```python
from pystreampdf import SemanticChunker, PDFCache, TokenBudgetConfig

# Setup budget config
budget_config = TokenBudgetConfig(
    base_budget=800,
    rules=[BudgetRule("complex", 1.1, match_fields=["filename"])]
)

# Initialize cache with budget management
cache = PDFCache(
    memory_limit_mb=500,
    disk_cache_dir="./pdf_cache",
    token_budget_config=budget_config
)

# Process PDF with intelligent chunking
def extract_pdf(pdf_path):
    chunks, preview, title, pages = cache.get_or_process(
        pdf_path,
        processor_fn=extract_chunks
    )
    return chunks

# Use semantic chunker directly
chunker = SemanticChunker(target_tokens=500)
chunks = chunker.chunk_content(text, element_type=ElementType.TEXT)

Budget Configuration

Default Settings

  • Minimum budget: 500 tokens (fixed)
  • Maximum budget: 1000 tokens (fixed)
  • Default base: 1000 tokens
  • Multiplier range: 0.5 - 1.5 (common)

Dynamic Adjustment via Multipliers

Use keyword-based rules to scale within the 500-1000 range without changing hard limits. See examples for financial, legal, and summary documents.

MCP 2.0 Integration

Enable MCP tools on port 8780 (see MCP_QUICKSTART.md for details).

AI systems discover all 207 tools across 18 projects:

  • 12 PyStreamPDF tools for document processing
  • 195 tools from 17 integrated projects
  • Multi-project workflows with intelligent orchestration
  • 60-75% reduction in context usage through query optimization

Documentation

Part of Unified Platform

18 projects, 207 tools, 18 simultaneous MCP endpoints (8765-8782).

All tools discoverable via MCP protocol in a single connection.

Version History

v2.0.0 (Current)

  • ✅ MCP 2.0 Support
  • ✅ Integrated with 17 other projects
  • ✅ 207 unified MCP tools
  • ✅ Intelligent orchestration
  • ✅ Production-ready (wheels only)

License

MIT


MCP 2.0 Mega-Platform | v2.0.0 | Wheels-Only Distribution

Release files for PyStreamPDF 2.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for PyStreamPDF 2.1.2
File Size Uploaded
pystreampdf-2.1.2.tar.gz 98.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for PyStreamPDF 2.1.2
File Interpreter ABI Platform
pystreampdf-2.1.2-cp313-cp313-macosx_11_0_arm64.whl CPython 3.13 CPython 3.13 macOS 11.0+ ARM64 Details

Total release size: 1.6 MB

Release files / pystreampdf-2.1.2.tar.gz

Download URL pystreampdf-2.1.2.tar.gz
Size 98.1 kB
Tags Source
SHA-256 checksum
How to use checksums
6b151faa32d2d439831b4ae97138b23a22b3699305d4130bc392d9b16a623757
BLAKE2b-256 checksum
How to use checksums
639f4c9bfda2c47f3dbc6ca49732bdac127bb13fd8f8ba7ede5fbb8786209242
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / pystreampdf-2.1.2-cp313-cp313-macosx_11_0_arm64.whl

Download URL pystreampdf-2.1.2-cp313-cp313-macosx_11_0_arm64.whl
Size 1.5 MB
Tags CPython 3.13 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
16659cc66d71e07bfe97b2cf90cf087a0e47f9ac40f63e0af59c9e6013d3fe4d
BLAKE2b-256 checksum
How to use checksums
db5d0df6961d2cda6026181db36f5036b2094d5e1bcf939b6b142cf75e50294c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

2.3.0

2 release files

2.2.1

1 release file

2.2.0

2 release files

This release

2.1.2 This release

2 release files

2.1.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page