Skip to main content

pdfmuse

English · 中文

crates.io PyPI npm CI live demo license

▶ Live playground — drag a PDF, watch it parse in your browser (nothing is uploaded)

pdfmuse playground: original PDF ↔ pdfmuse reconstruction

Deterministic PDF/DOCX parser for RAG / LLMs — one Rust core, with Python, Node & WASM bindings that produce byte-identical output.

pdfmuse is a precision pre-layer for AI/RAG: it extracts everything a file actually contains — text with exact coordinates, fonts, vector rules, tables, links — fast, robustly, and identically across every binding. It stops cleanly at the ML boundary: OCR and visual layout inference are left to a pluggable backend, so the core stays deterministic with zero ML dependencies. It is not another probabilistic vision model.

Why pdfmuse

Complete Keeps the finest-grained chars + coordinates; never silently drops content.
Fast Zero-copy streaming Rust core with a custom O(1) object parser + content tokenizer and per-page parallelism.
Robust A broken page/object never sinks the doc — returns structured errors, never panics (fuzz-tested).
Deterministic Same input → same output. No probabilistic models, no time/RNG in the core path.
Consistent Python / Node / WASM call one Rust core; output is byte-identical (CI-enforced).
CJK first-class CID/Type0 fonts + CMap/ToUnicode in the main path; compatibility codepoints NFKC-normalized for clean search.

Performance

Two things matter for a RAG pre-layer: speed, and whether it keeps the content. Both are measured on a public, reproducible corpus — 61 arXiv papers across 8 fields (large, dense PDFs — a deliberately hard case), so you can rerun the exact benchmark:

python benches/fetch_corpus.py --out /tmp/corpus      # the same PDFs, from a fixed manifest
pip install "pdfmuse==0.1.10" "pymupdf==1.28.0" "pdfplumber==0.11.10"
python benches/compare.py --dir /tmp/corpus

Text extraction (to_text, median of 7 runs after warm-up; PyMuPDF 1.28 / MuPDF 1.29, pdfplumber 0.11, macOS arm64, 65 papers):

vs speedup (geomean) win rate worst case
PyMuPDF ~7.7× faster 65 / 65 (100%) still 2.5× faster
pdfplumber ~150× faster 65 / 65 (100%) 69×

pdfmuse is faster on every file in this corpus — including a 22 MB paper (9× faster) and a plot-heavy one that draws 18k marker glyphs. Content is preserved: median 100% of PyMuPDF's non-whitespace characters (n=65).

to_text() / to_markdown() return a string straight from the Rust core (no full-IR deserialization). The full parse() — chars + bboxes + tables, far more than text — costs only ~2.3× the to_text time, still under PyMuPDF on most files. The native Node binding is ~as fast as the Rust core; WASM ~1.7×.

Honest limit — reading order: extraction is complete (100% of chars) and deterministic, but flattening a 2-D page to 1-D text is where the hard cases live. Single-column, tables, and clean two-column read correctly; dense two-column academic PDFs with very tight gutters can still interleave the columns (a known geometric edge — see docs/ / issue tracker). Eyeball any file with examples/visual_check.py.

Install

# Rust
cargo add pdfmuse-core
# Python (abi3 wheels)
pip install pdfmuse
# Node
npm install @pdfmuse/node   # native binding
# WASM (browser + Node, no native binary)
npm install @pdfmuse/core   # ESM `import` in the browser, `require` on Node >= 18

Usage

CLI (debug/inspection):

pdfmuse parse report.pdf --format md      # structured Markdown (headings, tables)
pdfmuse parse report.pdf --format json    # full IR (chars, bboxes, blocks, warnings)

Rust:

let data = std::fs::read("report.pdf")?;
let doc = pdfmuse_core::parse(&data, None)?;                 // auto-detect PDF/DOCX
for page in &doc.pages {
    for ch in &page.chars { /* ch.text, ch.bbox {x0,y0,x1,y1}, ch.size */ }
}
let md = pdfmuse_core::to_markdown(&doc);
let chunks = pdfmuse_core::chunk(&doc);                      // RAG chunks + {page, bbox, heading_path}

Python:

import pdfmuse
data = open("report.pdf", "rb").read()
text = pdfmuse.to_text(data)         # plain text — fast path (~1.3ms, no full-IR json.loads)
md = pdfmuse.to_markdown(data)       # structured Markdown — headings (PDF & DOCX) + tables
doc = pdfmuse.parse(data)            # full IR: doc.pages[i].chars/blocks with bboxes
clean = pdfmuse.to_text(data, drop_boilerplate=True)  # strip running headers/footers

Node:

const { toText, toMarkdown, parse } = require("@pdfmuse/node");
const data = fs.readFileSync("report.pdf");
const text = toText(data);           // plain text — fast path
const clean = toText(data, undefined, true);  // strip running headers/footers
const doc = parse(data);             // full IR (typed Document)

WASM (browser — digital PDFs; scanned pages return a NeedsOcr warning to hand off server-side):

import init, { to_text, parse } from "@pdfmuse/core";
await init();
const text = to_text(new Uint8Array(bytes));         // plain text
const doc = JSON.parse(parse(new Uint8Array(bytes))); // full IR

Integrations

  • LangChain — langchain-pdfmuse: a PdfmuseLoader with single / page / elements modes. In elements mode each chunk carries section-aware metadata (heading_path, bbox, category) — reproducible chunks for RAG.

    from langchain_pdfmuse import PdfmuseLoader
    docs = PdfmuseLoader("report.pdf", mode="elements").load()
    
  • LlamaIndex — llama-index-readers-pdfmuse: a PdfmuseReader with the same modes and section-aware metadata.

    from llama_index.readers.pdfmuse import PdfmuseReader
    docs = PdfmuseReader(mode="elements").load_data("report.pdf")
    
  • Haystack — pdfmuse-haystack: a PdfmuseConverter component (text / markdown) for Haystack 2.x pipelines.

    from pdfmuse_haystack import PdfmuseConverter
    docs = PdfmuseConverter(mode="markdown").run(sources=["report.pdf"])["documents"]
    

Scope boundary

In the core (deterministic): text + coordinates/font/size/color · vector rules & rects · line/paragraph/column clustering · heading detection (font-size + numbering) · running header/footer detection + opt-in removal · ruled & whitespace-aligned table reconstruction · full DOCX structure · JSON / Markdown / RAG-chunk output.

Out of the core (pluggable VisionBackend): scanned-page OCR · borderless-table structure recognition · heading/body/caption classification. Text-less (scanned/image) pages are flagged NeedsOcr and left for a backend — see docs/adr/0001-pdf-engine-strategy.md.

Guarding this boundary is what keeps pdfmuse fast, stable, and distinct from vision models.

Layout

crates/
  pdfmuse-core/     pure-Rust core: PDF/DOCX → unified IR (parser, tokenizer, layout, output)
  pdfmuse-python/   PyO3 (abi3) binding
  pdfmuse-node/     napi-rs binding
  pdfmuse-wasm/     wasm-bindgen binding
  pdfmuse-cli/      debug CLI (`pdfmuse`)
tests/{corpus,snapshots}   golden corpus + insta snapshots
tests/parity/              cross-binding byte-identical gate (Python == Node == WASM)
examples/visual_check.py   render original ↔ coordinate reconstruction for QA
fuzz/                      cargo-fuzz targets (never-panic)

Testing gates

  • Snapshot tests (insta + tests/corpus)
  • Cross-binding parity CI — Python/Node/WASM output byte-identical (a red gate blocks merge)
  • Robustness — mutated/garbage input never panics (tests/robustness.rs + fuzz/)
  • CJK correctness suite

Status

Core is feature-complete (milestones M0–M4 + real-world hardening M4.5): PDF + DOCX → unified IR → JSON / Markdown / RAG chunks, three byte-identical bindings, encryption, CJK. Currently in M5 · polish & release. Roadmap and tasks live in Linear (project pdfmuse).

License

Dual-licensed under MIT or Apache-2.0, at your option.

Metadata

Release files for pdfmuse 0.1.12

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for pdfmuse 0.1.12
File Interpreter ABI Platform
pdfmuse-0.1.12-cp38-abi3-win_amd64.whl CPython 3.8 abi3 Windows x86-64 Details
pdfmuse-0.1.12-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.8 abi3 Linux glibc 2.17+ x86-64 Details
pdfmuse-0.1.12-cp38-abi3-macosx_11_0_arm64.whl CPython 3.8 abi3 macOS 11.0+ ARM64 Details

Total release size: 2.6 MB

Release files / pdfmuse-0.1.12-cp38-abi3-win_amd64.whl

Download URL pdfmuse-0.1.12-cp38-abi3-win_amd64.whl
Size 787.6 kB
Tags CPython 3.8 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
9235f81d563b052be3fa6766cfa6c39014f314f76514c88fcffc731529c47ac8
BLAKE2b-256 checksum
How to use checksums
2aa3c9e997c7abce3aa24ce06cd4d0d4ce99be344e99ecbcfcbeb271eb02aa34
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.

Transparency log

Release files / pdfmuse-0.1.12-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL pdfmuse-0.1.12-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 938.4 kB
Tags CPython 3.8 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
78585067f792801ef9c65665cc8967897f453fd18703e1fb3bab9ce4d23e000b
BLAKE2b-256 checksum
How to use checksums
839e7cd180189d1d39f4d0c569cbc844e92abf898625e32daeacde0a6b8cf71c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.

Transparency log

Release files / pdfmuse-0.1.12-cp38-abi3-macosx_11_0_arm64.whl

Download URL pdfmuse-0.1.12-cp38-abi3-macosx_11_0_arm64.whl
Size 835.7 kB
Tags CPython 3.8 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
93360ab9a1d4189a7cf45266423b80cfe8cfc4850cc938152dd0ca4b88b729ce
BLAKE2b-256 checksum
How to use checksums
0b924715b5cb5dac3083b2313c4d298f79a3e26c3d737c216e199319ec58c004
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.12 This release

3 release files

0.1.11

3 release files

0.1.9

3 release files

0.1.8

3 release files

0.1.7

3 release files

0.1.6

3 release files

0.1.5

3 release files

0.1.4

3 release files

0.1.3

3 release files

0.1.2

3 release files

0.1.1

3 release files

0.1.0

3 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page