Skip to main content

PyStreamPDF

Reduce RAG costs 50-70%. Extract only what matters from PDFs.

Stop sending entire documents to LLMs. PyStreamPDF analyzes structure, identifies relevant sections, and extracts only critical content. Cut token costs 50-70% while improving retrieval accuracy.

PyPI Python 3.10+ Tests: 523 Passing License: Proprietary


30-Second Start

from pystreampdf import Document

# Extract only relevant content from PDF
doc = Document("financial_report.pdf")

# Smart content extraction
relevant = doc.extract_relevant("revenue for Q3 2024")
print(f"Extracted {len(relevant)} chunks")
print(f"Token savings: {relevant.token_savings:.0%}")  # 60% savings

# Send only relevant parts to LLM
for chunk in relevant:
    response = llm.query(chunk, "What was Q3 revenue?")

Why PyStreamPDF?

The Problem:

  • You send entire PDFs to LLMs (wasteful, expensive)
  • RAG systems retrieve too much content
  • Token costs skyrocket on large documents
  • No way to know which parts actually matter

The Solution:

  • Intelligent document analysis finds relevant sections
  • Semantic chunking with context awareness
  • 50-70% reduction in token usage
  • Better retrieval accuracy (less noise)

Key Features

  • Intelligent Extraction: Find relevant sections automatically
  • Semantic Chunking: Context-aware splitting, not just word count
  • Multi-Format: Text, tables, images, charts, OCR
  • Token Budgeting: Allocate tokens by document type
  • Smart Caching: L1 memory + L2 disk (avoid reprocessing)
  • Metadata Preservation: Keep tables, images, structure
  • Production-Ready: 523 tests, type-safe API

Real-World Use Cases

Financial Documents:

# Extract relevant sections from annual report
doc = Document("10-K_2024.pdf")
revenue_sections = doc.extract_relevant("revenue")
earnings_sections = doc.extract_relevant("earnings")

# 70% fewer tokens than sending whole PDF
for section in revenue_sections:
    summary = llm.query(section, "What was total revenue?")

Legal Contracts:

# Find clauses without reading everything
doc = Document("contract.pdf")
liability = doc.extract_relevant("liability", "indemnification")
print(f"Found in {len(liability)} sections")

Research Papers:

# Extract methodology and results
doc = Document("paper.pdf")
methods = doc.extract_relevant("methods", "experiment")
results = doc.extract_relevant("results", "findings")

Token Savings

Document Size Full PDF Tokens PyStreamPDF Savings
Annual Report 200 pages 50K 15K 70%
Contract 50 pages 12K 4K 67%
Research Paper 30 pages 8K 2K 75%

Results: Lower costs + better retrieval accuracy + faster responses


Installation

pip install pystreampdf
# or with uv
uv pip install pystreampdf

Documentation


License

Proprietary License - Free to use with explicit attribution. See LICENSE.


PyStreamPDF v2.1.0 | Intelligent PDF processing for AI | Python 3.10+ | 523 tests


See [TOKEN_BUDGET_MULTIPLIERS.md](docs/TOKEN_BUDGET_MULTIPLIERS.md) for comprehensive guide.

## Quick Start: Document Extraction

```python
from pystreampdf import SemanticChunker, PDFCache, TokenBudgetConfig

# Setup budget config
budget_config = TokenBudgetConfig(
    base_budget=800,
    rules=[BudgetRule("complex", 1.1, match_fields=["filename"])]
)

# Initialize cache with budget management
cache = PDFCache(
    memory_limit_mb=500,
    disk_cache_dir="./pdf_cache",
    token_budget_config=budget_config
)

# Process PDF with intelligent chunking
def extract_pdf(pdf_path):
    chunks, preview, title, pages = cache.get_or_process(
        pdf_path,
        processor_fn=extract_chunks
    )
    return chunks

# Use semantic chunker directly
chunker = SemanticChunker(target_tokens=500)
chunks = chunker.chunk_content(text, element_type=ElementType.TEXT)

Budget Configuration

Default Settings

  • Minimum budget: 500 tokens (fixed)
  • Maximum budget: 1000 tokens (fixed)
  • Default base: 1000 tokens
  • Multiplier range: 0.5 - 1.5 (common)

Dynamic Adjustment via Multipliers

Use keyword-based rules to scale within the 500-1000 range without changing hard limits. See examples for financial, legal, and summary documents.

MCP 2.0 Integration

Enable MCP tools on port 8780 (see MCP_QUICKSTART.md for details).

AI systems discover all 207 tools across 18 projects:

  • 12 PyStreamPDF tools for document processing
  • 195 tools from 17 integrated projects
  • Multi-project workflows with intelligent orchestration
  • 60-75% reduction in context usage through query optimization

Documentation

Part of Unified Platform

18 projects, 207 tools, 18 simultaneous MCP endpoints (8765-8782).

All tools discoverable via MCP protocol in a single connection.

Version History

v2.0.0 (Current)

  • ✅ MCP 2.0 Support
  • ✅ Integrated with 17 other projects
  • ✅ 207 unified MCP tools
  • ✅ Intelligent orchestration
  • ✅ Production-ready (wheels only)

License

MIT


MCP 2.0 Mega-Platform | v2.0.0 | Wheels-Only Distribution

Release files for PyStreamPDF 2.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for PyStreamPDF 2.1.1
File Interpreter ABI Platform
pystreampdf-2.1.1-cp313-cp313-macosx_11_0_arm64.whl CPython 3.13 CPython 3.13 macOS 11.0+ ARM64 Details

Release files / pystreampdf-2.1.1-cp313-cp313-macosx_11_0_arm64.whl

Download URL pystreampdf-2.1.1-cp313-cp313-macosx_11_0_arm64.whl
Size 1.6 MB
Tags CPython 3.13 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
f4294623511c0c000757c423244fab25e4c72663583859e5a19d9ee214bc20ed
BLAKE2b-256 checksum
How to use checksums
49c9733b5d8c9fd6edbb26ea3aa36091a9d10d24c4cc0a8afadbdeb43050c9e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

2.3.0

2 release files

2.2.1

1 release file

2.2.0

2 release files

2.1.2

2 release files

This release

2.1.1 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page