pdfspine
An Apache-2.0-licensed, pure-Rust reimplementation of PyMuPDF (fitz), with PyO3 Python bindings.
🦴 Part of the
spinefamily — framework-free backend engines, each the spine of a domain: zero framework lock-in, Protocol-ized seams, offline-capable. pdfspine is the PDF spine (this repo); ragspine is the RAG spine (deterministic dual-channel retrieval + agent orchestration).🤖 For AI agents / LLMs: before using this library, read
llms.txt(concise index) andpython/pdfspine/_llms/docs/(full API / recipes / gotchas); afterpip installthey ship atsite-packages/pdfspine/_llms/.
Status: alpha / pre-1.0, but the core is feature-complete. pdfspine can already parse/repair/decrypt PDFs, extract text & tables, search, edit / merge / split / save (incl. byte-exact incremental), encrypt, annotate, fill & flatten forms, redact (destructively), open image files as documents, render pages to images, and OCR (Tesseract + a pure-Rust PaddleOCR engine, stronger on CJK). 88.7% (682 / 769) of the PyMuPDF 1.24 public API is implemented and tested (climbing), with 1349+ Rust tests + 593+ Python tests green. Text extraction is at fitz parity (and beats fitz on Arabic / RTL), rendering is near-parity and ~1.74× faster, and the pure-Rust PaddleOCR engine beats fitz on CJK scans (see Accuracy). Now on PyPI:
pip install pdfspine(see Install); or build from source.
Why pdfspine?
PyMuPDF is excellent, but it is AGPL-3.0 (or a commercial license from Artifex) — a non-starter for many closed-source products, SaaS backends, and permissively-licensed open-source projects.
pdfspine is a drop-in-shaped, permissively-licensed (Apache-2.0) alternative:
- Apache-2.0 throughout — permissive, with an explicit patent grant. The
dependency graph is gated by
cargo-denyto exclude GPL / AGPL / LGPL / MPL / SSPL from the shipped wheel. License cleanliness is CI-enforced, not a promise. - Pure Rust, no C blob. Self-contained wheels, no system
zlib/C linkage, no bundled prebuilt engine (the differentiator vs pdfium-based wrappers). import fitzcompatible (opt-in). A compatibility shim lets much existing PyMuPDF code run unmodified — available asimport pdfspine.fitz as fitz, or registered under the globalfitz/pymupdfnames with one call topdfspine.install_fitz_shim(). A default install is collision-safe: it does not claim those global names, so it coexists with a real PyMuPDF in the same environment. A machine-readableCOMPAT.tomldocuments every symbol's status.- Memory-safe by construction.
#![forbid(unsafe_code)]in every first-party crate except the single audited PyO3 FFI chokepoint. - Clean-room. No code, tests, or fixtures derived from MuPDF / PyMuPDF / any AGPL source.
What works today
| Area | Capabilities |
|---|---|
| Read | open (file/bytes), malformed-PDF repair, encrypted PDFs (RC4 / AES-128 / AES-256, R2–R6) |
| Text | get_text (text/words/blocks/dict/rawdict/json/html/xhtml/xml), search_for, TextPage, fonts/images inventory |
| Tables | find_tables with merged-cell detection → extract() / to_markdown() / to_html() |
| Edit & save | full + byte-exact incremental save, garbage collection, page insert/delete/copy/move/select, insert_pdf merge, metadata/XMP, TOC, links, encryption write |
| Annotate | all common annotation types with /AP appearance streams; AcroForm read / fill / flatten + Widget; destructive redaction (verified content removal) |
| Render | get_pixmap (vector + text + image + shadings via a tiny-skia rasterizer), Pixmap (buffer-protocol/numpy), DisplayList, get_svg_image |
| Images | open PNG/JPEG/TIFF/GIF/BMP/WEBP as documents, convert_to_pdf, image-XObject decode (DCT/CCITT/JBIG2/JPX), extract_image |
| Markdown | markdown_to_pdf() — a pdfspine original extension (not a PyMuPDF API): CommonMark + GFM tables / strikethrough / task lists → PDF via a deterministic pure-Rust layout engine; local & data:-URI images (never the network); optional user TTF via font= / cjk_font= (CJK) |
| Layers | Optional Content Groups read/write (get_ocgs / add_ocg / set_layer) |
| OCR | pluggable engine: Tesseract adapter and a pure-Rust PaddleOCR engine (PP-OCRv5, weights from the shared ocrspine-models package, stronger on CJK) → searchable-sandwich PDF |
| CLI | pdfspine info / text / render / merge / split / pages / images / toc |
Planned next: reading-order accuracy improvements, Type1/Type3 glyph rendering,
broader CJK coverage. See PRD.md / docs/ROADMAP.md.
Out of scope: digital-signature creation.
Install
pip install pdfspine
pdfspine is on PyPI. OCR works out of the box: the PP-OCRv5 weights ship in
the shared ocrspine-models data
package — a runtime dependency pip pulls in automatically — so the wheel itself
stays lean and no longer embeds them. To build from source instead, see
Build & install.
Quick start
import pdfspine
doc = pdfspine.open("input.pdf")
print(len(doc), "pages", doc.metadata)
page = doc[0]
print(page.get_text()) # plain text
print(page.search_for("invoice")) # list[Rect]
page.get_pixmap(dpi=150).save("page1.png") # render to image
tables = page.find_tables()
for t in tables.tables:
print(t.to_markdown()) # or t.to_html() for merged cells
doc.save("output.pdf", garbage=4, deflate=True)
# Markdown → PDF (pdfspine original extension — not part of the PyMuPDF surface)
pdfspine.markdown_to_pdf("# Title\n\nHello **Markdown**!").save("hello.pdf")
Existing PyMuPDF code often runs unchanged via the opt-in compat shim:
import pdfspine.fitz as fitz # the shim, no global-name collision
doc = fitz.open("input.pdf")
text = doc[0].get_text("dict")
# Or make the literal `import fitz` resolve to the shim (one-time opt-in):
import pdfspine
pdfspine.install_fitz_shim()
import fitz # now -> pdfspine's fitz shim
A default install does not claim the global fitz / pymupdf names, so it
is safe alongside a real PyMuPDF; install_fitz_shim() uses setdefault and
never clobbers a PyMuPDF you imported first.
Command line:
pdfspine info report.pdf
pdfspine text report.pdf --pages 1-3 --format json -o out.json
pdfspine render report.pdf --dpi 200 -o images/
pdfspine merge a.pdf b.pdf -o merged.pdf
Accuracy
Validated against an objective ground-truth harness and with PyMuPDF (fitz) as
the differential oracle (clean-room: the AGPL oracle is run locally only and never
committed). See docs/BENCHMARKS.md and the
conformance/gt/ reports for the dated, reproducible evidence.
- Text extraction is at fitz parity on born-digital corpora, and beats fitz on Arabic / RTL (correct bidi reordering).
- Rendering is near-parity with fitz (page-image SSIM ~0.945) and ~1.74× faster after a font-cache fix.
- OCR beats fitz on CJK scans: the pure-Rust PaddleOCR engine (PP-OCRv5, with
weights from the shared
ocrspine-modelspackage) outperforms fitz's OCR path on Chinese/Japanese/Korean documents. - Real-corpus robustness: open rate 100%, 0 panics/hangs, re-saved files
100%
qpdf --check-clean across the public-domain US-government corpus.
Remaining accuracy work (multi-column reading order, Type1/Type3 glyph rendering,
broader CJK) is tracked in docs/PRD-NEXT.md.
Build & install
Requirements: Rust (pinned to 1.96.0 by rust-toolchain.toml), Python ≥
3.11, maturin ≥ 1.7. uv
recommended.
uv venv .venv && source .venv/bin/activate
maturin develop # build + install the extension in-place
python -c "import pdfspine; print(pdfspine.__version__)"
# redistributable wheel:
maturin build --release # -> target/wheels/
Building from source needs a C/asm compiler. The bundled pure-Rust PaddleOCR engine depends on
tract, which compiles target-specific assembly kernels at build time: a C compiler (cc/clang) on Linux/macOS, or the MSVC Build Tools (incl.ml64.exe) on Windows. Prebuilt PyPI wheels need none of this. To build a fully C-free library, compile the Rust crates with--no-default-features(drops thepaddle-ocrfeature). The wheel no longer embeds the OCR models — they ship in the sharedocrspine-modelspackage (a runtime dependency).
Architecture
A Cargo workspace with a strict dependency DAG; the Python bindings touch exactly one façade crate, and core logic is split into independently testable units.
py-bindings (PyO3 cdylib -> pdfspine._core, abi3-py311)
│
▼
pdf-api facade / re-exports
┌──────────┬───┴────┬──────────┐
▼ ▼ ▼ ▼
pdf-text pdf-edit pdf-image pdf-render
│ │ │ │
└────┬─────┘ │ (fonts, text)
▼ │
pdf-fonts ◄────────┘
▼
pdf-core ◄──────── pdf-crypto
| Crate | Responsibility |
|---|---|
pdf-core |
object model, lexer/parser, xref, repair, filters, writer, geometry |
pdf-crypto |
Standard security handler (RC4 / AES-128 / AES-256) |
pdf-fonts |
font mapping (encodings / ToUnicode / CMap / widths) |
pdf-text |
content-stream interpreter, get_text, search, find_tables |
pdf-edit |
page ops, merge, annotations / forms, metadata / TOC, redaction, OCG |
pdf-image |
image documents, image-XObject codecs, Pixmap |
pdf-render |
tiny-skia rasterizer → Pixmap, DisplayList, SVG |
pdf-api |
unified ergonomic façade |
py-bindings |
PyO3 wrappers → the _core extension module |
Develop / test
cargo fmt --all --check
cargo clippy --workspace --all-targets --all-features -- -D warnings
cargo test --workspace
maturin develop && pytest python/tests # Python tests
python conformance/run_validation.py … # real-corpus accuracy harness
pdfspine is built strictly test-first (red → green → refactor → harden); the
per-function test plan is in docs/test-case-catalog.md.
Documentation
Guide + API reference + PyMuPDF migration guide: build the docs site with
mkdocs serve (see mkdocs.yml / docs/). The
authoritative design lives in PRD.md.
License
Apache-2.0 — see LICENSE and NOTICE. All third-party
dependencies are permissive (MIT / Apache-2.0 / BSD / Zlib / …); the shipped graph
is CI-verified free of copyleft.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pdfspine-0.4.0.tar.gz.
File metadata
- Download URL: pdfspine-0.4.0.tar.gz
- Upload date:
- Size: 4.7 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ec1ee0692881e7291df9e4ac29af573c01af45ad156a100a4edb72bffd21a8f5
|
|
| MD5 |
8ac5c4557be5fe1f14f6667083ecbcaa
|
|
| BLAKE2b-256 |
f853e12c03632098aaf1ceb058f1fc5429aea4e6d933d41e095f2e1475785fc2
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-win_amd64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-win_amd64.whl
- Upload date:
- Size: 12.1 MB
- Tags: CPython 3.11+, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
32440991907fe7870255580c5284c69b1032b222aad0b5ee2621f0b5e5b952dd
|
|
| MD5 |
65d68586d4611512c81d1c7c9d60d024
|
|
| BLAKE2b-256 |
ca71bc89117f803dfc37f4675090ea4943bcc7c6f84533d9788972aa6eb0b6f7
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-manylinux_2_28_aarch64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-manylinux_2_28_aarch64.whl
- Upload date:
- Size: 10.9 MB
- Tags: CPython 3.11+, manylinux: glibc 2.28+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
16801d2d87f2ed509ea6a94e680729c63a0a20c1185e5e142cc94b84de6d48fe
|
|
| MD5 |
7c1e2fdae92b049b3bf5bdb206cc9761
|
|
| BLAKE2b-256 |
fcb2544e81424bedd8d2ece899bcbae0fce85222b0934168b54ffed8708ed90c
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 12.1 MB
- Tags: CPython 3.11+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
60ebbabb5cfd5d75a90a46fa8e880b4a17f08becb3f7c52f27815fe56ed299aa
|
|
| MD5 |
b86fdac92b681752afa92a1f42bc4415
|
|
| BLAKE2b-256 |
8de849dbf5feba5f30d68460f363aa7b75d3dd53a0502e50a9193dc65fa65e21
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 10.9 MB
- Tags: CPython 3.11+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ef997b3bf26ce5ba688935c7f65ebfdb9f81610d1722a8e452b0abf6d0135638
|
|
| MD5 |
af154533f304956a624e37f8b6d64578
|
|
| BLAKE2b-256 |
2e6b02a87b7b726361e69802ff4d7b9d4c7dc1a302a62d47ec5b41ab1c28e1ff
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 11.7 MB
- Tags: CPython 3.11+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c918ba7f783e65d6713e68f8d3a15c6e34573819ab9ada2dcde531659cd40e3f
|
|
| MD5 |
9210de93c919ea5745743391e810f9e5
|
|
| BLAKE2b-256 |
57d10d11f3f352d66a03ab050b0da272371382139406c004e3e17490f2f30efc
|
File details
Details for the file pdfspine-0.4.0-cp311-abi3-macosx_10_12_x86_64.whl.
File metadata
- Download URL: pdfspine-0.4.0-cp311-abi3-macosx_10_12_x86_64.whl
- Upload date:
- Size: 12.6 MB
- Tags: CPython 3.11+, macOS 10.12+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5c3432ce73eaed321c5f7e831a07fecfade75f965dec37a240506a039e1281f9
|
|
| MD5 |
becbe3f8a451e60ee5351a8fcfdfbde2
|
|
| BLAKE2b-256 |
7cf094256242fec90ef71ad4c36b9150b2fb3c7e4bd3cc16741cbe6ab8f1b260
|