Skip to main content

pdfboss

A PDF engine written from scratch in Rust — parse, extract text, rasterize to PNG. One core, a CLI, and pythonic bindings.

CI python-ci PyPI Rust 2021 MIT OR Apache-2.0


Motivation

Reading a PDF shouldn't mean linking a C library. pdfboss is a clean-room reader built straight from the ISO 32000 specification: no C dependencies, no bindings to anyone else's engine — just safe Rust with a small, obvious API. The same core powers a CLI and a native Python extension, so a script and a service share one implementation.

It is a lenient reader: real-world files are damaged, and pdfboss recovers rather than refuses — reconstructing broken cross-reference tables, tolerating wrong stream lengths, and skipping garbage operators instead of erroring out.

Install

Python

pip install pdfboss

Prebuilt abi3 wheels (CPython ≥ 3.12) for Linux and macOS; no toolchain required.

Rust

cargo add pdfboss-core pdfboss-text pdfboss-render pdfboss-aio pdfboss-tui   # library crates
cargo install pdfboss-cli                                                    # the `pdfboss` binary

Usage

CLI

pdfboss info    report.pdf                 # version, page count, sizes, metadata
pdfboss text    report.pdf --page 2        # extract text (omit --page for all)
pdfboss render  report.pdf --page 1 -o page.png --scale 2.0
pdfboss obj     report.pdf 5               # pretty-print object 5

Explorer subcommands, each accepting a local path or an http(s):// URL (range-fetched, never downloaded whole):

pdfboss json    report.pdf                    # dump the document as a JSON value tree
pdfboss hex     report.pdf obj:5              # hexdump the file or a selected element
pdfboss q       report.pdf '.header.version'  # jq-style queries over the JSON tree
pdfboss tui     report.pdf                    # interactive terminal explorer

Python

import pdfboss

doc = pdfboss.Document("report.pdf")       # or Document(data=raw_bytes)
print(doc.page_count, doc.version, doc.metadata)

page = doc[0]
print(page.width, page.height, page.rotation)
text = page.extract_text()                 # or doc.extract_text() for all pages
png  = page.render(scale=2.0)              # PNG bytes

for element in doc.elements():             # lazy: physical + logical, byte spans included
    print(element.kind, element.span)

# Async access over files or http(s) URLs, without reading the whole document.
doc = await pdfboss.AsyncDocument.open_url("https://example.com/report.pdf")
async for element in doc.elements():
    print(element.kind, element.value)

Rust

use pdfboss_core::Document;

let doc = Document::open("report.pdf")?;
let page = doc.page(0)?;

let text = pdfboss_text::extract_text(&doc, &page)?;
let pixmap = pdfboss_render::render_page(&doc, &page, 2.0)?;
pixmap.save_png("page.png")?;

What's inside

Crate Responsibility
pdfboss-core Tokenizer, object model, stream filters, cross-references, object streams, document & page tree, content-stream operators
pdfboss-text Simple and CID/Type0 fonts, standard encodings, ToUnicode CMaps, positional text extraction
pdfboss-render Anti-aliased vector rasterizer — paths, fills, strokes, clipping, color, images — to RGBA/PNG
pdfboss-aio Async I/O: range-fetching document access over files or HTTP, without reading the whole file
pdfboss-cli The pdfboss command-line tool
pdfboss-tui Interactive terminal explorer (pdfboss tui), built on pdfboss-aio
pdfboss-py PyO3 extension module (pdfboss._pdfboss) built with maturin

Supported: classic, stream, and hybrid cross-references with recovery scanning · object streams · FlateDecode, LZWDecode, ASCII85Decode, ASCIIHexDecode, RunLengthDecode + PNG/TIFF predictors · DCTDecode (JPEG) images · CCITTFaxDecode scans — Group 3 one-dimensional, Group 3 mixed and Group 4 coding (ITU-T T.4/T.6) · JBIG2Decode scans — generic regions, symbol dictionaries and text regions, arithmetic- or Huffman-coded, MMR-coded generic regions and collective bitmaps, with or without /JBIG2Globals · Standard-handler decryption — RC4 and AES-128/256 (empty user password) · page-tree attribute inheritance · text extraction with ToUnicode and WinAnsi/MacRoman/Standard encodings · rasterization of paths, fills (nonzero & even-odd), strokes, transforms, clipping, image/form XObjects, and embedded-TrueType glyph outlines · lazy element iteration over physical (objects, xref sections, trailer, with byte spans) and logical (pages, fonts, images, annotations, content operators) elements.

Benchmarks

Text and parsing

Against other Python PDF libraries over 40 real-world PDFs (best-of-3 per file, aggregated over the files every library handled; pages/sec, higher is faster):

pdfboss vs. Python PDF libraries

pdfboss is the fastest library measured on both operations — including against the C-backed PyMuPDF. On text extraction it reaches 1,539 pages/s versus PyMuPDF's 279 (≈5.5×), and 25–80× the pure-Python readers. On open + parse it reaches 19,114 pages/s versus PyMuPDF's 3,766 (≈5×): lazy page-tree loading means opening a document reads only its declared page count instead of parsing every page dictionary up front. Rendering is not compared — pdfboss's rasterizer does not yet paint every glyph, so timing it against full renderers would be misleading.

Numbers are machine-dependent; reproduce with benchmarks/bench.py.

Scanned documents

Scans are the other half of the world's PDFs, and they are a different workload: one full-page bilevel image per page, JBIG2- or CCITT-coded, with no text operators at all. Rendering is comparable there — with no glyphs to paint, every library draws the same picture — so it gets its own benchmark, over a 544-page JBIG2 book (1994 × 2832 samples per page) rasterized to PNG at 1:1.

Library pages/sec Ink on page 1
pdfboss 53.7 4.83%
pdfplumber (via pdfium) 35.5 4.87%
PyMuPDF 35.0 4.82%
pypdfium2 32.7 4.85%

pdfboss is the fastest of the four here, at about 1.5× the C-backed renderers — and the only one of them with no C in it. Compare the four rows against each other rather than against another machine's: all four are timed in one pass, and the ratio between them held to within 3% across runs whose absolute numbers varied by a fifth.

What is left is the codec itself. Four fifths of the time goes to the JBIG2 arithmetic decoder and the context formation feeding it, and that part is a serial dependency chain — every decision needs the interval state the previous one wrote, and every pixel's context contains the pixels just decoded — so it neither vectorizes nor parallelizes. The rest was arithmetic that did not need doing: expanding a packed scan into eight times its size in RGBA before sampling a fraction of it, blending opaque pixels through an alpha formula that returns them unchanged, and walking bitmaps a pixel at a time where a row of bytes would do.

The ink column is what makes the timings mean anything: a library that cannot decode a scan's codec usually hands back a blank page instead of raising, and a blank page benchmarks superbly. Agreeing coverage says all four decoded the same picture. They do not agree pixel for pixel — each downsamples 1994 × 2832 samples onto a 462 × 663 page with its own resampling.

Reproduce with benchmarks/bench_scans.py.

Limitations

Rendered pages paint the outlines of embedded TrueType glyphs (Type0/CIDFontType2 under Identity, and simple /TrueType fonts via their cmap). Text in other fonts (CFF/Type1 programs, the standard 14, subset fonts without a usable cmap) is still positioned but not drawn.

JBIG2Decode covers generic regions (all four templates, with TPGDON, arithmetic or MMR-coded), symbol dictionaries and text regions in both the arithmetic and the Huffman variant, and custom code table segments. That is what scanners actually emit, but it is not the whole standard, and the rest is refused rather than approximated — a stream using refinement or aggregate coding, pattern dictionaries or halftone regions fails with a message naming the feature, so a scan that will not decode says why on the first try.

Not yet supported (they error or degrade gracefully, and are on the roadmap): password-protected documents (the empty user password is handled for both RC4 and AES) · non-TrueType glyph outlines (CFF/Type1) · shadings and tiling patterns · JPXDecode (JPEG 2000) · the JBIG2 features listed above · soft masks and blend modes · annotation appearance streams.

Rendering is lenient: content pdfboss cannot read is skipped so the rest of the page still rasterizes. It says so rather than passing the result off as a faithful render — pdfboss render prints a warning line per dropped item on stderr and annotates its summary, the TUI preview raises a status-bar notice, and the libraries expose the detail through render_page_reporting (Rust) and Page.render_reporting() (Python), which return the pixels plus a report of everything dropped or approximated.

The sync and async APIs are not at parity on encryption: Document/Page (and the CLI's info/text/render/obj) decrypt empty-user-password RC4/AES files transparently, as above. AsyncDocument (pdfboss tui, any http(s):// target, and the Python AsyncDocument) currently rejects every encrypted document outright, real password or not — async decryption parity is a tracked follow-up.

Development

cargo test --workspace          # Rust test suite
cargo clippy --workspace --all-targets -- -D warnings
maturin develop                 # build the Python extension into your venv
pytest                          # Python integration tests

License

Dual-licensed under either of

at your option. Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you shall be dual-licensed as above, without any additional terms or conditions.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdfboss-0.7.1.tar.gz (3.1 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pdfboss-0.7.1-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.6 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.17+ x86-64

pdfboss-0.7.1-cp312-abi3-macosx_11_0_arm64.whl (2.3 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file pdfboss-0.7.1.tar.gz.

File metadata

  • Download URL: pdfboss-0.7.1.tar.gz
  • Upload date:
  • Size: 3.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for pdfboss-0.7.1.tar.gz
Algorithm Hash digest
SHA256 5d116e20565c0be03507068f2818a3afabc345b627ad2a7f8a86b1d01af2ed1a
MD5 926267e6ddca83df2491dcc2b19897c2
BLAKE2b-256 662b10fea0281b8d99b9f6823aa8d1cb452f3e81595e5d87102261d4a2186e2d

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.7.1.tar.gz:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdfboss-0.7.1-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pdfboss-0.7.1-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 dc3403b03fd456c0841948c3015d03917a0f16319c25da08585090fd56f301f2
MD5 82ab12cf7d5bdc483f6fa40d42e65bd0
BLAKE2b-256 eda59fd4246165c2161a63189354f5d998feff389839c8171619cb3b8a56eb31

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.7.1-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdfboss-0.7.1-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pdfboss-0.7.1-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 a9087e0ac96bc9cf81bac3da0e332d1dd48f90e5512bd2fe31ed4d64c237f3c1
MD5 add6471a9912aa2b7dbe943a307b7735
BLAKE2b-256 e7f18fd15b3a79d09d124b81c53b23a5781d1bfd6daa933e0c1ac5a94f9a5576

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.7.1-cp312-abi3-macosx_11_0_arm64.whl:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.2.0

3 files

1.1.0

3 files

1.0.0

3 files

0.25.0

3 files

0.24.0

3 files

0.23.0

3 files

0.22.0

3 files

0.21.1

3 files

0.21.0

3 files

0.20.0

3 files

0.19.1

3 files

0.19.0

3 files

0.18.0

3 files

0.17.1

3 files

0.17.0

3 files

0.16.0

3 files

0.15.0

3 files

0.14.0

3 files

0.13.0

3 files

0.12.1

3 files

0.12.0

3 files

0.11.0

3 files

0.10.0

3 files

0.9.0

3 files

0.8.0

3 files

0.7.2

3 files

This release

0.7.1 This release

3 files

0.7.0

3 files

0.6.0

3 files

0.5.0

3 files

0.4.1

3 files

0.4.0

3 files

0.3.0

3 files

0.2.1

3 files

0.1.0

3 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page