Pure-Rust PDF extraction that distills documents into clean, LLM-ready HTML — for LLMs and RAG, built on lopdf
Project description
distillPDF
Turn any PDF into clean, LLM-ready HTML or Markdown — structure-aware, pure-Rust, MIT-licensed.
📓 New — OCR for scanned PDFs: detect image-only pages and OCR them into clean HTML or a compact searchable PDF — with a bundled, offline fast engine (no extra, no download) or an optional accurate VLM. See the OCR example notebook »
distillpdf reads a PDF and reconstructs its structure — reading order, headings,
paragraphs, lists, tables, and figures — then emits compact, semantic HTML or
Markdown (or plain text) ready to feed to an LLM or a RAG pipeline. No styling noise, no
layout junk: just the content a model needs. Markdown is produced from the same HTML, so
both formats benefit from every extraction improvement.
It's built on lopdf and shipped to Python via
PyO3 + maturin as a small, self-contained
wheel — a lightweight, permissively licensed alternative to AGPL/heavyweight extractors
(PyMuPDF, pdfminer, Unstructured), with no system dependencies and no Python runtime deps.
🧪 Early release (
0.0.3) — testers wanted. The API is small and may still change. If you have PDFs that come out wrong, please open an issue with the file (or a description) — real-world documents are exactly what this needs to get better.
Install
pip install distillpdf
Prebuilt wheels; no compiler or system libraries required. Installing also puts a
distillpdf command on your PATH.
Command line
Convert a PDF to clean HTML or Markdown in one command:
distillpdf paper.pdf # HTML to stdout
distillpdf paper.pdf -o paper.html # ...or to an HTML file
distillpdf paper.pdf -o paper.md # ...or Markdown (inferred from the .md extension)
distillpdf paper.pdf --markdown # Markdown to stdout
distillpdf *.pdf -o out/ # batch: out/<name>.html per input
distillpdf paper.pdf -o p.html --image-mode external # lean HTML + an img/ folder
distillpdf paper.pdf --image-mode drop # replace images with placeholder text
distillpdf paper.pdf --mode page # page-first HTML (default is section-first)
distillpdf paper.pdf --no-toc # omit the table-of-contents nav
distillpdf paper.pdf --text # plain text instead of HTML
distillpdf paper.pdf --toc # print the table of contents
distillpdf paper.pdf --section abstract
distillpdf scan.pdf --ocr # OCR a scanned PDF → scan.searchable.pdf (bundled, no extra)
(Also available as python -m distillpdf.)
Quickstart
import distillpdf
doc = distillpdf.open("paper.pdf") # or distillpdf.from_bytes(data)
# By default, these WRITE a file and return 1 (deriving the name from the PDF):
doc.to_html() # → paper.html (one self-contained file)
doc.to_markdown() # → paper.md + an img/ folder of figures
doc.to_html("out.html") # ...or to a specific path / directory
doc.to_html("out.html", image_mode="external") # ...lean HTML + an img/ folder
# Pass return_string=True to get the rendered text back instead of writing:
html = doc.to_html(return_string=True) # self-contained HTML string (inline images)
md = doc.to_markdown(return_string=True) # ...or Markdown (built from the same HTML)
# rendering options work the same on both:
doc.to_html(mode="page", toc=False)
text = doc.extract_text() # plain text, in reading order
toc = doc.toc() # [(level, title, page, anchor_id), ...]
abstract = doc.section("abstract") # targeted section extraction (returns a string)
Markdown
to_markdown() is a transform of the very HTML to_html() produces — so every
processor improvement (clipping, heading detection, tables, front-matter) flows into
Markdown automatically, with no second renderer to keep in sync.
doc.to_markdown() # string (images embedded inline)
doc.to_markdown("paper.md") # writes paper.md
doc.to_markdown("paper.md", image_mode="external") # paper.md + paper's img/fig_NN_slug.ext
doc.to_markdown(image_mode="drop") # drop images (caption-only placeholders)
image_mode controls figures — identically for to_html() and to_markdown():
image_mode |
result |
|---|---|
"embed" (default) |
inline base64 data: URIs — one self-contained string/file |
"external" |
extract each figure to img/fig_NN_slug.ext (vectors as .svg) and reference it; only when writing to a file (a returned string falls back to "embed") |
"drop" |
replace images with placeholder text |
Because both formats run through the same converter, "external" produces the same
img/ layout whether you write .html or .md.
Output modes
By default, logical sections are first-order: every heading becomes its own nested
<section id="sec-…">, so you can pull a whole section as one block (great for RAG / LLM
chunking), and page numbers are dropped.
distillpdf.open("paper.pdf").to_html(return_string=True)
# <section id="sec-abstract"><h2>Abstract</h2><p>…</p></section>
distillpdf.open("paper.pdf").section("methods") # → the <section id="sec-methods"> block
Pass mode="page" for the page-faithful structure instead — each page wrapped in
<section data-page="N" id="page-N">, with page numbers in the TOC:
distillpdf.open("paper.pdf").to_html(mode="page", return_string=True)
Want compact, text-only output? image_mode="drop" replaces each embedded image with a
lightweight <image N> placeholder (captions and figure anchors are kept):
distillpdf.open("paper.pdf").to_html(image_mode="drop", return_string=True)
# <figure id="fig-1"><image 1><figcaption>…</figcaption></figure>
Pass toc=False to skip the auto table-of-contents <nav> (heading anchors are still
emitted, so #section links and doc.section(...) keep working):
distillpdf.open("paper.pdf").to_html(toc=False, return_string=True)
Rendering options
open() only loads the PDF; the rendering options live on to_html() and
to_markdown() (and mode on toc()/section()), since that's where the content is
actually extracted:
| Option | Default | Effect |
|---|---|---|
path= |
None |
where to write: a file, or a directory to place <source-stem>.html/.md in. None writes <source>.html/.md next to the PDF. (Ignored when return_string=True.) |
return_string= |
False |
True returns the rendered string and writes nothing; the default writes a file and returns 1 |
mode= |
"section" |
"page" wraps each page in <section data-page="N"> and numbers TOC entries; the default groups content into nested <section id="sec-…"> and drops page info |
image_mode= |
"embed" |
"embed" inline data: URIs (self-contained); "external" an img/ folder (when writing to a file); "drop" placeholder text |
toc= |
True |
False omits the <nav> table of contents (section/heading anchors still emitted) |
Raw pieces
Need the structured data instead of HTML?
doc.extract_tables() # cell grids (handles multi-level / colspan headers)
doc.extract_images() # embedded images, with raw bytes
doc.extract_links() # hyperlinks with targets
doc.extract_fonts() # font inventory
doc.page_count() # number of pages
The .dpdf document model
distillPDF builds a typed element tree per document — reading order, headings, the section
tree, tables, figures, OCR provenance — and normally renders it once and throws it away.
Distilling persists that analysis to a .dpdf file (a zip of model.json + image
assets). Distill once → re-render HTML / Markdown / text from the file, forever and
byte-identically, with no source PDF and no re-analysis. The model is also queryable on its
own — for a single document, an agent with file tools needs no corpus or vector store.
import distillpdf
distillpdf.open("case.pdf").distill("case.dpdf") # analyse once → a durable model file
doc = distillpdf.load("case.dpdf") # reopen it — no source PDF needed
doc.info() # counts, OCR state, asset profile (as data)
doc.toc() # [(level, title, page, section-id), ...]
print(doc.section("sec-methods")) # one section as markdown (the whole subtree)
hits = doc.find("indemnification") # lexical search…
print(len(hits.hits), "in", hits.searched_blocks, "blocks;",
len(hits.no_text_pages), "pages had no text") # …with HONEST coverage, never a silent miss
doc.embed() # build the semantic index (BAAI/bge-m3 vectors)
res = doc.search("limits on the seller's liability") # semantic search over chunks…
for h in res.hits[:3]:
print(h["chunk_id"], round(h["score"], 3), h["section"], h["block_ids"]) # ids thread into read
doc.to_markdown("case.md") # fidelity re-render from the model (byte-identical
doc.to_html("case.html") # to to_markdown/to_html on the original PDF)
It distributes itself as an agent CLI — any shell with distillpdf installed can drive a
.dpdf, no SDK or server. Every listing emits ids that thread into the next call, read
carries navigation breadcrumbs and resumable truncation, and find ends with a coverage line:
$ distillpdf case.pdf -o case.dpdf # distill (a .dpdf output path, not HTML)
distillpdf: wrote case.dpdf
$ distillpdf case.dpdf info # pages, sections, tables, OCR state, assets
case.pdf (schema v0, distillpdf 0.0.32)
pages: 1564 sections: 42 blocks: 9210
$ distillpdf case.dpdf toc # section tree: ids → read targets
sec-introduction Introduction (p1-4)
sec-methods Methods (p5-12)
$ distillpdf case.dpdf read sec-methods # one section as markdown + breadcrumbs
…
prev: sec-introduction · next: sec-results · parent: —
$ distillpdf case.dpdf find "fls. 249" --pages xii-xv # scoped lexical search
b0421 [sec-methods] p13 (fls. 249) …matched «fls. 249» phrase…
searched 318 blocks across 4 pages
$ distillpdf case.dpdf embed # build the semantic index (BAAI/bge-m3)
distillpdf: embedded 612 chunks into space 'e1' (BAAI/bge-m3, dim 1024, backend vendored)
$ distillpdf case.dpdf search "limits on liability" # semantic search; ids thread into read
semantic search over 612 chunks (model BAAI/bge-m3, space e1)
c0207 score=0.7193 [sec-indemnification] p41-42
Neither party's aggregate liability shall exceed the fees paid in the twelve months…
blocks: b1180 b1181 b1182
Semantic search is opt-in and self-contained: embed chunks the document (consecutive
same-section blocks, ~400 tokens each — derived from blocks, like the indexes, so no text is
duplicated) and stores BAAI/bge-m3 vectors inside the .dpdf as a binary member; search
ranks chunks by cosine similarity. The vectors are byte-identical to what kglite's bge-m3 stack
produces (distillpdf uses kglite's embedder directly when installed, else a faithful vendored
twin of it). The embedding runtime is an optional dependency (pip install onnxruntime tokenizers huggingface_hub) — a missing one yields the exact install line, never a silent
miss; find --semantic is an alias for search. If the document's blocks change (re-distill),
a stale space is dropped loudly and info flags it.
Asset profiles make the size/sharing trade-off an explicit choice, never a surprise:
distill(assets="figures") (the default) embeds figure images; assets="none" keeps text +
structure only (a few MB even for a 1,500-page scan — emailable); every dropped binary leaves
a named, regenerable stub (hash, dimensions, a recipe to re-extract it from the source PDF),
so a hole is observable, not silent.
🧪 Experimental (
schema_version 0). The.dpdfschema is not yet frozen — it stays at0until the first downstream cutover proves the shape; treat it as a working format, not a stable contract, and re-distill to pick up extraction improvements (a.dpdfis a snapshot of extractor quality at distill time). See docs/datamodel-design.md.
OCR — scanned PDFs
Image-only / scanned pages have no text to extract. distillPDF OCRs them and folds the recovered text back into the same HTML / Markdown / searchable PDF outputs — born-digital pages keep distillPDF's normal extraction. There are two tiers (full OCR setup guide » for per-OS install + GPU):
| Tier | Engine | Install | Speed | Quality | Notes |
|---|---|---|---|---|---|
| fast (default) | bundled Tesseract | none — in the wheel | ~0.8 s/page | char ~95% | offline, no download; flat text (no tables) |
| accurate | granite-docling VLM | install a runtime yourself (see below) | ~6 s/page (GPU) | char ~97% | structure + tables; downloads a model |
The fast tier works out of the box on a plain pip install distillpdf — no extra, no
PyTorch, no model download, fully offline. English, Portuguese and Norwegian ship in the
wheel; the document's language is auto-detected from a sample at the start of processing and
OCR runs in just that language (faster, more accurate). Pin it with
OcrConfig(languages=["eng"]), or point TESSDATA_PREFIX at your own tessdata for other
languages.
The accurate tier needs a heavier model runtime that's genuinely platform- and
hardware-specific, so there's no catch-all [ocr] extra — you pick a path and install it,
and distillpdf prints the exact commands: python -c "import distillpdf; print(distillpdf.ocr.install_help('granite'))". The short version (full OCR setup guide »):
# macOS (Apple Silicon) — Metal GPU via MLX, automatic:
pip install mlx-vlm "transformers>=4.57,<5" pillow
# Windows / Linux / Intel-Mac — PyTorch, no C++ compiler:
pip install torch "transformers>=4.57,<5" pillow
# …for an NVIDIA GPU, install the CUDA build instead (default torch is CPU-only and slow):
pip install torch --index-url https://download.pytorch.org/whl/cu124
# Lightweight alternative — GGUF (engine="granite-docling-gguf"):
pip install llama-cpp-python huggingface-hub pillow
(transformers is pinned <5: 5.x changed the idefics3 image processor and fails to load
granite-docling; >=4.57 is the floor that supports it.)
Then doc.run_ocr(engine="granite"). The model is public and downloads on first run — MLX
fetches it automatically, PyTorch/GGUF put it in a visible ./ocr_model/ folder
(OcrConfig(model_dir=…) overrides). No token needed for the default model; a gated repo needs
HF_TOKEN (env or .env) or OcrConfig(hf_token=…, store_token=True). The speed gap is real: a
509-page scan OCRs in ~6 min on the fast tier vs much longer on the accurate tier — and the
accurate tier on CPU is very slow (minutes/page), so use a GPU for it or stick with the fast tier.
From the command line — open → OCR (progress bar shown automatically) → write, no Python:
distillpdf scan.pdf --ocr # → scan.searchable.pdf (fast tier, bundled)
distillpdf scan.pdf --ocr --remove-raster # → reflowed clean text + figures, smaller file
distillpdf scan.pdf --ocr -o out.html # OCR'd HTML (use a .md path for Markdown)
distillpdf scan.pdf --ocr --ocr-engine accurate # granite-docling (install a runtime — see setup guide)
distillpdf --list-ocr-engines # show engines: name, tier, bundled, offline
Or from Python:
import distillpdf
doc = distillpdf.open("scanned.pdf")
doc.run_ocr() # fast tier by default — bundled, offline; cached on the document
# (a progress bar shows on a terminal — pass progress=False to silence)
# accurate tier (install a granite runtime first — see the setup guide / install_help):
# doc.to_html("out.html", ocr=True, engine="granite") # or run_ocr(engine="granite")
doc.to_pdf("out.pdf") # searchable PDF (reuses the cache — no second pass)
doc.to_html("out.html") # OCR text folded into clean HTML
doc.to_markdown("out.md") # …and Markdown
run_ocr OCRs each scanned page once and caches the result on the document, so every output
is rendered from a single pass. Trying it out? Pass only={1, 2, 3} to OCR just a few pages first. For a single output you can skip the explicit call and pass
ocr=True — e.g. doc.to_pdf("out.pdf", ocr=True) runs OCR (once) then writes. The render
methods (to_html / to_markdown / to_pdf) also work without OCR — they just warn
that scanned pages have no text and point you at run_ocr.
Detection handles real-world scans — images nested in Form XObjects, CCITT Group-4 fax and
Flate-wrapped JPEG encodings, and full-page rasters whose only text is an e-filing stamp.
Searchable-PDF modes (doc.to_pdf):
- keep the scan (default) — the original page image is preserved and the OCR text is added as an invisible, selectable layer over it. The scan always shows, so OCR errors never destroy content (best for archival/legal use).
remove_raster=True— pages are reflowed to clean visible text + cropped figures and the raster is dropped, for a much smaller file.
Tables are recovered natively by the accurate tier: granite-docling emits OTSL table
structure, which distillPDF renders as a real <table> (with <th> / colspan / rowspan) in
HTML and a gridded table (cell rules + shaded header row) in the searchable PDF. The reflow
also justifies granite's paragraph blocks for a typeset look. The fast (Tesseract) tier
produces flat text only — use the accurate tier when table structure matters.
Pick the tier for the job: the fast tier is great for "make this scan searchable, now"; the accurate tier is for structure-faithful extraction (tables, headings, reading order). See the OCR example notebook ».
Why distillPDF
- Structure, not just text. Two-column reading order, multi-level table headers mapped
onto a single grid (
colspan), vector figures transcoded to inline SVG (including rotated axis labels), an auto-generated table of contents, and named section extraction (doc.section("methods")). - LLM-ready output. Lean, class-free HTML — semantic markup a model can read directly,
with anchor ids so
toc()entries link straight into the document. - Small & permissive. Pure Rust on
lopdf, MIT-licensed, no system dependencies, no Python runtime dependencies. Drops into any pipeline without license headaches. - Fast. Native Rust extraction with a release build tuned for speed (LTO, single codegen unit).
Scope
In scope: text, table, image, and font extraction; an HTML/markdown output layer for RAG and LLM ingestion; and optional OCR for scanned pages with a searchable-PDF writer.
Out of scope (for now): page rendering.
Comparison
| distillPDF | PyMuPDF | pdfminer.six | Unstructured | |
|---|---|---|---|---|
| License | MIT | AGPL / commercial | MIT | Apache (heavy deps) |
| Structure-aware HTML | ✅ | partial | ❌ | ✅ |
| System deps | none | none | none | many |
| Implementation | Rust | C | Python | Python |
Contributing & feedback
This is a young project and feedback is the fastest way to improve it. The most useful things you can do:
- Try it on your PDFs and tell me where the output is wrong — open an issue.
- Star the repo if it's useful, so others can find it.
- PRs welcome — see the development notes below.
Development
The test suite lives in tests/ (pytest) and runs on CI. It needs only
distillpdf installed. CI runs entirely on data we own — a self-contained demo PDF
(tests/demo/, end-to-end structure check) and a synthetic table corpus
(tests/corpus_tables/). The third-party PDF corpora (tests/corpus*/) are gitignored, so
their tests self-skip on a fresh clone and run only when the corpora are present locally for
deeper coverage.
Build from source with maturin:
git clone https://github.com/kkollsga/distillpdf
cd distillpdf
maturin develop --release # build + install into the current venv
bash tests/run.sh # build distillpdf + run pytest
pytest tests/ -q # or just run the tests against an installed build
License
MIT — see LICENSE. Use it anywhere, including commercial and closed-source projects.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file distillpdf-0.0.33-cp38-abi3-win_amd64.whl.
File metadata
- Download URL: distillpdf-0.0.33-cp38-abi3-win_amd64.whl
- Upload date:
- Size: 9.1 MB
- Tags: CPython 3.8+, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
43577eea31caffee5d5fbc7b32b72210db804b4b0773ea9eefc23b090154ef2c
|
|
| MD5 |
b24acb19aeea1d94186161fa8cbc4133
|
|
| BLAKE2b-256 |
9994e2e788570ca4cb176af7a69addaa1e5ffcf4e555477b5c61aa0d7102219e
|
Provenance
The following attestation bundles were made for distillpdf-0.0.33-cp38-abi3-win_amd64.whl:
Publisher:
publish.yml on kkollsga/distillpdf
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
distillpdf-0.0.33-cp38-abi3-win_amd64.whl -
Subject digest:
43577eea31caffee5d5fbc7b32b72210db804b4b0773ea9eefc23b090154ef2c - Sigstore transparency entry: 1779535327
- Sigstore integration time:
-
Permalink:
kkollsga/distillpdf@448c2d71a57f35798263ea78b122281234033684 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/kkollsga
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@448c2d71a57f35798263ea78b122281234033684 -
Trigger Event:
push
-
Statement type:
File details
Details for the file distillpdf-0.0.33-cp38-abi3-manylinux_2_39_x86_64.whl.
File metadata
- Download URL: distillpdf-0.0.33-cp38-abi3-manylinux_2_39_x86_64.whl
- Upload date:
- Size: 8.4 MB
- Tags: CPython 3.8+, manylinux: glibc 2.39+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d648ace02066040cc0a4fba834f96756d0970a10e54548b55f29d980de05c08d
|
|
| MD5 |
7eedc221bbd869601de5e846f3fdbd75
|
|
| BLAKE2b-256 |
4f2191ef74e23beaef428b16930fbfebe7ad055bc98108d12987d05353d88256
|
Provenance
The following attestation bundles were made for distillpdf-0.0.33-cp38-abi3-manylinux_2_39_x86_64.whl:
Publisher:
publish.yml on kkollsga/distillpdf
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
distillpdf-0.0.33-cp38-abi3-manylinux_2_39_x86_64.whl -
Subject digest:
d648ace02066040cc0a4fba834f96756d0970a10e54548b55f29d980de05c08d - Sigstore transparency entry: 1779535171
- Sigstore integration time:
-
Permalink:
kkollsga/distillpdf@448c2d71a57f35798263ea78b122281234033684 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/kkollsga
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@448c2d71a57f35798263ea78b122281234033684 -
Trigger Event:
push
-
Statement type:
File details
Details for the file distillpdf-0.0.33-cp38-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: distillpdf-0.0.33-cp38-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 8.0 MB
- Tags: CPython 3.8+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5dd6bd103ce53f4c61f2f76ef7eccb1978facae7a782a93846a8e91229a51a9c
|
|
| MD5 |
3e2c6a84775c900e4d4a6ef0188a27a8
|
|
| BLAKE2b-256 |
5e2197364bc372f3817e3f570221fd04bc613c0ca9ac1d7d95dcc66654503be0
|
Provenance
The following attestation bundles were made for distillpdf-0.0.33-cp38-abi3-macosx_11_0_arm64.whl:
Publisher:
publish.yml on kkollsga/distillpdf
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
distillpdf-0.0.33-cp38-abi3-macosx_11_0_arm64.whl -
Subject digest:
5dd6bd103ce53f4c61f2f76ef7eccb1978facae7a782a93846a8e91229a51a9c - Sigstore transparency entry: 1779535233
- Sigstore integration time:
-
Permalink:
kkollsga/distillpdf@448c2d71a57f35798263ea78b122281234033684 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/kkollsga
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@448c2d71a57f35798263ea78b122281234033684 -
Trigger Event:
push
-
Statement type: