Skip to main content

pdfboss

A PDF engine written from scratch in Rust — parse, extract text, rasterize to PNG. One core, a CLI, and pythonic bindings.

CI python-ci PyPI Rust 2021 MIT OR Apache-2.0


Motivation

Reading a PDF shouldn't mean linking a C library. pdfboss is a clean-room reader built straight from the ISO 32000 specification: no C dependencies, no bindings to anyone else's engine — just safe Rust with a small, obvious API. The same core powers a CLI and a native Python extension, so a script and a service share one implementation.

It is a lenient reader: real-world files are damaged, and pdfboss recovers rather than refuses — reconstructing broken cross-reference tables, tolerating wrong stream lengths, and skipping garbage operators instead of erroring out.

Install

Python

pip install pdfboss

Prebuilt abi3 wheels (CPython ≥ 3.12) for Linux and macOS; no toolchain required.

Rust

cargo add pdfboss-core pdfboss-text pdfboss-render pdfboss-aio pdfboss-tui   # library crates
cargo install pdfboss-cli                                                    # the `pdfboss` binary

Usage

CLI

pdfboss info    report.pdf                 # version, page count, sizes, metadata
pdfboss text    report.pdf --page 2        # extract text (omit --page for all)
pdfboss render  report.pdf --page 1 -o page.png --scale 2.0
pdfboss obj     report.pdf 5               # pretty-print object 5

Explorer subcommands, each accepting a local path or an http(s):// URL (range-fetched, never downloaded whole):

pdfboss json    report.pdf                    # dump the document as a JSON value tree
pdfboss hex     report.pdf obj:5              # hexdump the file or a selected element
pdfboss q       report.pdf '.header.version'  # jq-style queries over the JSON tree
pdfboss tui     report.pdf                    # interactive terminal explorer

Python

import pdfboss

doc = pdfboss.Document("report.pdf")       # or Document(data=raw_bytes)
print(doc.page_count, doc.version, doc.metadata)

page = doc[0]
print(page.width, page.height, page.rotation)
text = page.extract_text()                 # or doc.extract_text() for all pages
png  = page.render(scale=2.0)              # PNG bytes

for element in doc.elements():             # lazy: physical + logical, byte spans included
    print(element.kind, element.span)

# Async access over files or http(s) URLs, without reading the whole document.
doc = await pdfboss.AsyncDocument.open_url("https://example.com/report.pdf")
async for element in doc.elements():
    print(element.kind, element.value)

Rust

use pdfboss_core::Document;

let doc = Document::open("report.pdf")?;
let page = doc.page(0)?;

let text = pdfboss_text::extract_text(&doc, &page)?;
let pixmap = pdfboss_render::render_page(&doc, &page, 2.0)?;
pixmap.save_png("page.png")?;

What's inside

Crate Responsibility
pdfboss-core Tokenizer, object model, stream filters, cross-references, object streams, document & page tree, content-stream operators
pdfboss-text Simple and CID/Type0 fonts, standard encodings, ToUnicode CMaps, positional text extraction
pdfboss-render Anti-aliased vector rasterizer — paths, fills, strokes, clipping, color, images — to RGBA/PNG
pdfboss-aio Async I/O: range-fetching document access over files or HTTP, without reading the whole file
pdfboss-cli The pdfboss command-line tool
pdfboss-tui Interactive terminal explorer (pdfboss tui), built on pdfboss-aio
pdfboss-py PyO3 extension module (pdfboss._pdfboss) built with maturin

Supported: classic, stream, and hybrid cross-references with recovery scanning · object streams · FlateDecode, LZWDecode, ASCII85Decode, ASCIIHexDecode, RunLengthDecode + PNG/TIFF predictors · DCTDecode (JPEG) images · CCITTFaxDecode scans — Group 3 one-dimensional, Group 3 mixed and Group 4 coding (ITU-T T.4/T.6) · JBIG2Decode scans — arithmetic-coded generic regions, symbol dictionaries and text regions, MMR-coded generic regions, with or without /JBIG2Globals · Standard-handler decryption — RC4 and AES-128/256 (empty user password) · page-tree attribute inheritance · text extraction with ToUnicode and WinAnsi/MacRoman/Standard encodings · rasterization of paths, fills (nonzero & even-odd), strokes, transforms, clipping, image/form XObjects, and embedded-TrueType glyph outlines · lazy element iteration over physical (objects, xref sections, trailer, with byte spans) and logical (pages, fonts, images, annotations, content operators) elements.

Benchmarks

Text and parsing

Against other Python PDF libraries over 40 real-world PDFs (best-of-3 per file, aggregated over the files every library handled; pages/sec, higher is faster):

pdfboss vs. Python PDF libraries

pdfboss is the fastest library measured on both operations — including against the C-backed PyMuPDF. On text extraction it reaches 1,539 pages/s versus PyMuPDF's 279 (≈5.5×), and 25–80× the pure-Python readers. On open + parse it reaches 19,114 pages/s versus PyMuPDF's 3,766 (≈5×): lazy page-tree loading means opening a document reads only its declared page count instead of parsing every page dictionary up front. Rendering is not compared — pdfboss's rasterizer does not yet paint every glyph, so timing it against full renderers would be misleading.

Numbers are machine-dependent; reproduce with benchmarks/bench.py.

Scanned documents

Scans are the other half of the world's PDFs, and they are a different workload: one full-page bilevel image per page, JBIG2- or CCITT-coded, with no text operators at all. Rendering is comparable there — with no glyphs to paint, every library draws the same picture — so it gets its own benchmark, over a 544-page JBIG2 book (1994 × 2832 samples per page) rasterized to PNG at 1:1.

Library pages/sec Ink on page 1
pdfplumber (via pdfium) 59.6 4.87%
PyMuPDF 56.2 4.82%
pypdfium2 55.9 4.85%
pdfboss 42.7 4.83%

Here pdfboss is the slowest of the four, at roughly 0.7× the C-backed renderers — and the only one of them with no C in it. About two thirds of its time is the JBIG2 arithmetic decoder, which is the honest cost of decoding the format rather than delegating it.

The ink column is what makes the timings mean anything: a library that cannot decode a scan's codec usually hands back a blank page instead of raising, and a blank page benchmarks superbly. Agreeing coverage says all four decoded the same picture. They do not agree pixel for pixel — each downsamples 1994 × 2832 samples onto a 462 × 663 page with its own resampling.

Reproduce with benchmarks/bench_scans.py.

Limitations

Rendered pages paint the outlines of embedded TrueType glyphs (Type0/CIDFontType2 under Identity, and simple /TrueType fonts via their cmap). Text in other fonts (CFF/Type1 programs, the standard 14, subset fonts without a usable cmap) is still positioned but not drawn.

JBIG2Decode covers the arithmetic half of the format plus MMR: generic regions (all four templates, with TPGDON, arithmetic or MMR-coded), symbol dictionaries, and text regions. That is what scanners actually emit, but it is not the whole standard, and the rest is refused rather than approximated — a stream using Huffman-coded symbol dictionaries or text regions, custom Huffman tables, refinement or aggregate coding, pattern dictionaries or halftone regions fails with a message naming the feature, so a scan that will not decode says why on the first try.

Not yet supported (they error or degrade gracefully, and are on the roadmap): password-protected documents (the empty user password is handled for both RC4 and AES) · non-TrueType glyph outlines (CFF/Type1) · shadings and tiling patterns · JPXDecode (JPEG 2000) · the JBIG2 features listed above · soft masks and blend modes · annotation appearance streams.

Rendering is lenient: content pdfboss cannot read is skipped so the rest of the page still rasterizes. It says so rather than passing the result off as a faithful render — pdfboss render prints a warning line per dropped item on stderr and annotates its summary, the TUI preview raises a status-bar notice, and the libraries expose the detail through render_page_reporting (Rust) and Page.render_reporting() (Python), which return the pixels plus a report of everything dropped or approximated.

The sync and async APIs are not at parity on encryption: Document/Page (and the CLI's info/text/render/obj) decrypt empty-user-password RC4/AES files transparently, as above. AsyncDocument (pdfboss tui, any http(s):// target, and the Python AsyncDocument) currently rejects every encrypted document outright, real password or not — async decryption parity is a tracked follow-up.

Development

cargo test --workspace          # Rust test suite
cargo clippy --workspace --all-targets -- -D warnings
maturin develop                 # build the Python extension into your venv
pytest                          # Python integration tests

License

Dual-licensed under either of

at your option. Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you shall be dual-licensed as above, without any additional terms or conditions.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdfboss-0.6.0.tar.gz (3.1 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pdfboss-0.6.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.6 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.17+ x86-64

pdfboss-0.6.0-cp312-abi3-macosx_11_0_arm64.whl (2.3 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file pdfboss-0.6.0.tar.gz.

File metadata

  • Download URL: pdfboss-0.6.0.tar.gz
  • Upload date:
  • Size: 3.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for pdfboss-0.6.0.tar.gz
Algorithm Hash digest
SHA256 535ae74001bbe6b9bb9529c4d1d73075a401ef01ccab85634ac1dd1ceef9a4f1
MD5 f59a4b8ea8dd03eeee89c8bb71cbf1f9
BLAKE2b-256 b7a256c879aad6a711ae8246360f0feecb58abc4f9962638dbe61064c9e304d9

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.6.0.tar.gz:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdfboss-0.6.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pdfboss-0.6.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 eadb1f2327ebf72a00e411da856c0a63e2daef8ba362dfe447cfce5fa16ec0bb
MD5 4971fbd11d599662935876d49298c545
BLAKE2b-256 0311f617dde1ea4fa78d12af77d5cf1180be8819d7f5d5e5360bf9ca7fb7edfd

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.6.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdfboss-0.6.0-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pdfboss-0.6.0-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 1978a9682a5da6694a4ebe1a170acbc6a2b7e1d01a2a49f0d8d438e84c38316c
MD5 885a202dc7b2b40fef453d8e8a8ad77a
BLAKE2b-256 31f27a84f33e6a5656f5eeeef4c1e191a7f3b056542205aba9bff3a16eadab74

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfboss-0.6.0-cp312-abi3-macosx_11_0_arm64.whl:

Publisher: release-please.yaml on 4thel00z/pdfboss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.2.0

3 files

1.1.0

3 files

1.0.0

3 files

0.25.0

3 files

0.24.0

3 files

0.23.0

3 files

0.22.0

3 files

0.21.1

3 files

0.21.0

3 files

0.20.0

3 files

0.19.1

3 files

0.19.0

3 files

0.18.0

3 files

0.17.1

3 files

0.17.0

3 files

0.16.0

3 files

0.15.0

3 files

0.14.0

3 files

0.13.0

3 files

0.12.1

3 files

0.12.0

3 files

0.11.0

3 files

0.10.0

3 files

0.9.0

3 files

0.8.0

3 files

0.7.2

3 files

0.7.1

3 files

0.7.0

3 files

This release

0.6.0 This release

3 files

0.5.0

3 files

0.4.1

3 files

0.4.0

3 files

0.3.0

3 files

0.2.1

3 files

0.1.0

3 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page