Skip to main content

PyStreamPDF

Extract only what matters from PDFs, so you send less to LLMs.

Stop sending entire documents to LLMs. PyStreamPDF parses PDF structure and splits content into semantic, token-budget-aware chunks so you can send only the relevant parts of a document to a model. See "Token Savings" below for how much that saves in practice — it depends on your documents.

PyPI Python 3.9+ Tests: 557 Passing License: Proprietary


30-Second Start

from pystreampdf import SemanticChunker, ElementType

with open("financial_report.pdf", "rb") as f:
    text = f.read().decode("latin-1", errors="ignore")  # or use a real PDF text extractor first

chunker = SemanticChunker(target_tokens=500)
chunks = chunker.chunk_content(
    text, element_type=ElementType.TEXT, page_start=1, page_end=1
)
print(f"Split into {len(chunks)} semantic chunks")

There is no Document/.extract_relevant() class in this package — SemanticChunker (shown above) and PDFCache (see "Quick Start: Document Extraction" below) are the real, exported API. If you built against a Document class from an earlier version of these docs, it never actually existed in source; see Known Issues.


Why PyStreamPDF?

The Problem:

  • You send entire PDFs to LLMs (wasteful, expensive)
  • RAG systems retrieve too much content
  • Token costs skyrocket on large documents
  • No way to know which parts actually matter

The Solution:

  • Semantic chunking with context awareness (real, in SemanticChunker)
  • Token-budget-aware caching (real, in PDFCache / TokenBudgetConfig)
  • Reduces the volume of text you send per query — see "Token Savings" below for what's measured vs. illustrative

Key Features

  • Real encryption/permission detection: Backed by PDFium's actual document security APIs — not a stub. Encrypted, password-protected, and permission-restricted PDFs are detected and reported accurately; opening with a wrong/missing password fails closed rather than silently succeeding.
  • Honest failure signaling: If a PDF can't be parsed (corrupt, truncated, unsupported), you get a real error — never fabricated placeholder content.
  • Intelligent Extraction: Find relevant sections automatically
  • Semantic Chunking: Context-aware splitting, not just word count
  • Multi-Format: Text, tables, images, charts, OCR
  • Token Budgeting: Allocate tokens by document type
  • Smart Caching: L1 memory + L2 disk (HMAC-signed on disk — no unverified deserialization)
  • Metadata Preservation: Keep tables, images, structure
  • Production-Ready: 557 passing tests (2 skipped when optional OCR system deps aren't installed)

Real-World Use Cases

See examples/basic_parse.py for the Rust-backed pystreampdf.open() path (document/page/structure inspection — requires the compiled _core extension, see Known Issues) and examples/token_budget_and_cache_example.py for the pure-Python SemanticChunker + PDFCache + TokenBudgetConfig path shown above. Both are real, runnable scripts — copying them is more reliable than a README snippet for an evolving API.


Token Savings

A note on how token counts are computed: by default, token counts use a len(text) / 4 heuristic (a common rule of thumb, but approximate — it can be off by a meaningful margin, especially for code or non-English text). Install the optional tiktoken extra for exact BPE token counts:

pip install "pystreampdf[tiktoken]"

When tiktoken is installed, pystreampdf.tokenizer.is_exact() returns True and counts use the real cl100k_base encoding; otherwise it falls back to the heuristic. Check pystreampdf.tokenizer.TOKENIZER_MODE to see which mode produced a given number.

Actual token savings depend entirely on your documents and how narrowly you scope extract_relevant-style queries — there's no committed benchmark result in this repo backing a specific percentage (the earlier "70%"/"10-50x" style claims in this README and in __init__.py's docstring weren't measured against anything checked in). Run examples/token_budget_and_cache_example.py against your own documents to get a real number for your use case.


Security Notes

  • Encryption & permissions: pystreampdf._core.PyPdfDocument.is_encrypted(), .permissions(), and .open_with_password() are backed by PDFium's real document security APIs. A wrong or missing password on an encrypted PDF raises an error — it never falls back to an unauthenticated parse.
  • Cache integrity: the on-disk (L2) cache is HMAC-SHA256 signed using a key generated on first use and stored with owner-only (0600) permissions. A tampered or foreign cache file fails signature verification and is discarded before any deserialization occurs.
  • MCP/DAB connector: binds to 127.0.0.1 with no cross-origin access and read-only permissions by default. Wider exposure requires an explicit allow_remote=True opt-in — review the security implications first, since the connector has no authentication of its own.

Installation

pip install pystreampdf

# With exact token counting (tiktoken)
pip install "pystreampdf[tiktoken]"

Documentation

Quick Start: Document Extraction

from pystreampdf import SemanticChunker, PDFCache, TokenBudgetConfig, BudgetRule

# Setup budget config
budget_config = TokenBudgetConfig(
    base_budget=800,
    rules=[BudgetRule("complex", 1.1, match_fields=["filename"])]
)

# Initialize cache with budget management
cache = PDFCache(
    memory_limit_mb=500,
    disk_cache_dir="./pdf_cache",
    token_budget_config=budget_config
)

# Process PDF with intelligent chunking
def extract_pdf(pdf_path):
    chunks, preview, title, pages = cache.get_or_process(
        pdf_path,
        process_fn=extract_chunks
    )
    return chunks

# Use semantic chunker directly
chunker = SemanticChunker(target_tokens=500)
chunks = chunker.chunk_content(text, element_type=ElementType.TEXT, page_start=1, page_end=1)

Budget Configuration

Default Settings

  • Minimum budget: 500 tokens (fixed)
  • Maximum budget: 1000 tokens (fixed)
  • Default base: 1000 tokens
  • Multiplier range: 0.5 - 1.5 (common)

Dynamic Adjustment via Multipliers

Use keyword-based rules to scale within the 500-1000 range without changing hard limits. See examples for financial, legal, and summary documents.

MCP Integration (optional, local by default)

PyStreamPDF ships an optional MCP/DAB connector (pystreampdf._mcp_connector) exposing document-processing tools (extract text/tables/images, OCR, structure detection, metadata, chunking, citations, validation). It's opt-in, binds to 127.0.0.1 only, and requires allow_remote=True to be exposed beyond localhost. See examples/mcp_pystreampdf.py.

Known Issues

  • Earlier versions of this README documented a Document class with .extract_relevant() and a .token_savings attribute — verified against source: this class has never existed in this package. Fixed in this pass to describe the real, exported API (SemanticChunker, PDFCache, TokenBudgetConfig, and the Rust-backed pystreampdf.open()).
  • The published PyPI wheel is the only wheel on PyPI — macOS arm64 only, no sdist. On any other platform or Python version, pip install will fail without a local Rust toolchain to build from source.
  • The published version (2.2.1) is ahead of what's tagged in this repo's Cargo.toml/init.py (2.2.0) Resolved: 2.3.0 (this release) is unambiguously ahead of both 2.2.0 and 2.2.1.
  • The Rust-backed pystreampdf.open() / pystreampdf.load_index() API (used in examples/basic_parse.py) silently becomes None if the compiled _core extension isn't available (e.g. a from-source install without maturin develop) — falling back to SemanticChunker/PDFCache in that case, per the code above, not a hard error.
  • The "Tests & Build" GitHub Actions workflow was red for several weeks due to two CI infrastructure bugs (not the test suite itself), both now fixed and confirmed green in CI: run 32611116541. First, dtolnay/rust-toolchain@v1 required an explicit toolchain input the workflow didn't provide, failing both jobs before any tests ran — fixed by pinning to dtolnay/rust-toolchain@stable. Second, once that was fixed, maturin develop failed in the Python-test job because actions/setup-python doesn't provide an active virtualenv, which maturin requires — fixed by creating and activating a .venv before the build step. Actual CI output as of 2.2.0: pytest — 536 passed, 2 skipped (Python 3.10, 3.11, and 3.12, each identical); cargo test — 23 passed, 0 failed. (The Rust suite has 23 tests, not 536 — an earlier version of this section conflated the pytest count with the Rust one.) As of 2.3.0: pytest — 557 passed, 2 skipped (21 new tests added for tests/test_mcp_tools.py).

License

Proprietary License — Free to use with explicit attribution. See LICENSE.


PyStreamPDF v2.3.0 | Intelligent PDF processing for AI | Python 3.9+ | 557 passing tests

Release files for PyStreamPDF 2.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for PyStreamPDF 2.3.0
File Size Uploaded
pystreampdf-2.3.0.tar.gz 112.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for PyStreamPDF 2.3.0
File Interpreter ABI Platform
pystreampdf-2.3.0-cp311-cp311-macosx_11_0_arm64.whl CPython 3.11 CPython 3.11 macOS 11.0+ ARM64 Details

Total release size: 1.7 MB

Release files / pystreampdf-2.3.0.tar.gz

Download URL pystreampdf-2.3.0.tar.gz
Size 112.5 kB
Tags Source
SHA-256 checksum
How to use checksums
9b3945808a81a60ee8b4cd3cacd98bb45ca740050efada141f3d6b148085e829
BLAKE2b-256 checksum
How to use checksums
24e5ceb25301c4acfcca3bf937de5ca8d9af0b71ba39c89a3b19ec7df0d2619a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release files / pystreampdf-2.3.0-cp311-cp311-macosx_11_0_arm64.whl

Download URL pystreampdf-2.3.0-cp311-cp311-macosx_11_0_arm64.whl
Size 1.6 MB
Tags CPython 3.11 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
c5f1a6dd56fc0c2fb9843d5b6c90e663dd1a0922c7278f18e7a7522adedf7022
BLAKE2b-256 checksum
How to use checksums
e90245acc3f41763819c3e69204d5a8a96e384ca4fba1032139a35f08118e4bb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release history Release notifications | RSS feed

This release

2.3.0 This release

2 release files

2.2.1

1 release file

2.2.0

2 release files

2.1.2

2 release files

2.1.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page