pdf-html
Early-stage geometry-first PDF → HTML converter for text-based PDFs — verbatim text, source typography, zero images.
One command in. One elegant, dependency-free HTML file out.
🎯 What it does
pdf-html is an early-stage, geometry-first PDF-to-HTML converter for text-based PDFs. It rebuilds the document as a single self-contained HTML5 file that mirrors the source document's look and structure:
-
✅ Best for: text-based PDFs with real text layers, headings, lists, tables, and multi-column layouts
-
⚠️ Not a universal OCR-first converter: scanned/image-heavy PDFs still need an OCR pre-pass or a dedicated workflow
-
🏷️ Heading hierarchy — font-size tiers become real
<h1>–<h6> -
🎨 Typography & color — the page CSS is derived from the document's own fonts, sizes, and palette
-
📊 Tables — ruled tables are rebuilt as real
<table>elements, including lists inside cells and rows that continue across page breaks -
📝 Lists — bullets and numbered items become nested
<ul>/<ol>(indent decides nesting) -
🧭 Reading order — multi-column layouts are re-linearized column by column
-
✂️ Page furniture — repeated running headers/footers and page numbers are stripped
-
🔒 Text is verbatim — never summarized, reworded, or reordered within a block
-
🚫 No images, ever — image content is dropped by design; text alone carries the document
🧭 Early-stage scope
This is an early-stage, geometry-first conversion tool for text-based PDFs. It is designed to be reliable and deterministic where the PDF has a usable text layer, but it is not a universal OCR-first converter for scanned documents, forms, or arbitrary image-heavy PDFs.
What it does well
- text-based PDFs with real text layers
- styled HTML with verbatim text
- heading hierarchy, lists, and multi-column layout reconstruction
- ruled table reconstruction with cell-aware content flow
What it does not yet do well
- OCR and scanned-document support are planned as separate, opt-in paths
- arbitrary image-heavy PDFs or forms with weak text layers
- pixel-perfect visual reconstruction of every layout edge case
⚡ Quick start
# install from PyPI (Python 3.10+)
pip install pdf-html
# or: uv pip install pdf-html
# convert
pdf-html report.pdf -o report.html
# open report.html in any browser — no external assets needed
If you are working from a cloned repository instead of the published package, install the local checkout with pip install -e . or uv pip install -e ..
🖥️ CLI reference
pdf-html INPUT.pdf -o out.html [--extractor pymupdf|pdftotext]
[--style auto|default] [--paginate] [--no-tables] [--no-callouts]
[--keep-headers] [--allow-scanned]
| Flag | Default | What it does |
|---|---|---|
-o, --output |
(required) | Path of the HTML file to write |
--extractor |
pymupdf |
Text extraction backend (pdftotext fallback planned) |
--style |
auto |
auto derives CSS from the document's own fonts/sizes/colors; default uses a clean built-in theme |
--paginate |
off | Wrap each PDF page in a <section class="sheet"> |
--no-tables |
off | Disable table reconstruction (table text flows as paragraphs) |
--keep-headers |
off | Keep repeated running headers/footers |
--allow-scanned |
off | Convert scanned/image PDFs instead of exiting with an OCR hint |
🏗️ Architecture
A deterministic pipeline — every stage is a small, pure, individually testable module:
🔍 Pipeline diagram — click to enlarge (click again to close) · open full screen ↗
---
config:
layout: elk
theme: neutral
---
flowchart LR
subgraph S1["📥 1 · Extract"]
direction TB
PDF@{ shape: doc, label: "📄 PDF" }
EX@{ shape: rect, label: "Extractor<br/>spans + font metadata" }
PDF --> EX
end
subgraph S2["🧹 2 · Clean & Profile"]
direction TB
HF@{ shape: rect, label: "Header/Footer<br/>stripping" }
SP@{ shape: hex, label: "Style Profiler<br/>body size · h1–h6 · palette" }
HF --> SP
end
subgraph S3["🧭 3 · Layout"]
direction TB
RO@{ shape: rect, label: "Reading Order<br/>column clustering" }
TR@{ shape: fr-rect, label: "Table Reconstructor<br/>per-cell pipeline" }
end
subgraph S4["🏗️ 4 · Structure"]
direction TB
SD@{ shape: div-rect, label: "Structure Detector<br/>headings · paragraphs · lists" }
LP@{ shape: rect, label: "List Parser<br/>nested ul/ol" }
AST@{ shape: bow-rect, label: "AST<br/>Document → Page → Blocks" }
SD --> LP --> AST
end
subgraph S5["🎨 5 · Render"]
direction TB
RN@{ shape: rect, label: "Renderer<br/>semantic HTML5 + inline CSS" }
OUT@{ shape: tag-doc, label: "🌐 out.html" }
RN --> OUT
end
EX --> HF
SP --> RO
SP --> TR
RO --> SD
TR --> SD
AST --> RN
| Module | Responsibility |
|---|---|
extractor.py |
TextExtractor ABC; PyMuPDF backend reads per-span size/weight/color/bbox and detects table regions |
header_footer.py |
Strips spans repeating on ≥ 60% of pages in the top/bottom 10% bands |
style_profiler.py |
Character-weighted font-size histogram → body size, heading tiers, color palette |
reading_order.py |
Column detection via x-gap clustering, with a card-grid fallback and straddle guard |
table_reconstructor.py |
Assigns spans to detected cells, runs the full pipeline inside each cell, merges cross-page rows |
structure.py |
Classifies lines into headings/paragraphs/list items from geometry + font cues |
list_parser.py |
Indent-based nesting; glyph style only picks ul vs ol; markers stripped, text verbatim |
ast.py |
Typed document model — Document → Page → Block, runs carry inline style |
renderer.py |
Single-file HTML5 with one <style> block and CSS variables from the profile |
📊 Table reconstruction highlights
The hardest part of PDF → HTML is tables. pdf-html:
- 🔍 Detects ruled tables geometrically (PyMuPDF
find_tables()) at extraction time - 📌 Assigns the page's styled spans to cells by bounding box — inline bold/color/size survive
- 🔄 Runs the normal line → paragraph → list pipeline inside every cell, so bullets in cells become real nested lists
- 🧵 Merges rows that continue across page breaks (empty-first-cell fragments) back into one row — even resuming mid-list-item
- 🏷️ Promotes a bold-only first row to a
<th>header row
🧭 Design principles
| Principle | Meaning |
|---|---|
| 🧮 Pure geometry, no AI | All structure is inferred from font metadata and bounding boxes. No NLP, no LLM, no document-specific regexes |
| 🔒 Text is sacred | Output text is verbatim; only & < > are escaped |
| 🚫 No images | Spans overlapping image rects are dropped; <img> is never emitted |
| 🪂 Graceful degradation | Heuristic failures only affect styling — never text content or order |
| 🔧 Tunable & testable | Every heuristic threshold is a named module-level constant with focused unit tests |
⚖️ How it compares
Every PDF converter picks a trade-off. pdf-html optimizes for semantic, reflowable, styled HTML with a verbatim-text guarantee — a square none of the established tools occupy:
| Tool | Output | Semantic structure | Keeps typography | Deterministic | Footprint |
|---|---|---|---|---|---|
| pdf-html | Self-contained HTML5 | ✅ h1–h6, ul/ol, table |
✅ CSS derived from the source | ✅ | ~30 MB (PyMuPDF only) |
| pdf2htmlEX | Pixel-faithful HTML | ❌ positioned glyphs | ✅ visually | ✅ | C++ toolchain |
| Poppler pdftohtml | Positioned divs / bare text | ❌ | ⚠️ partial | ✅ | system package |
| pymupdf4llm | Markdown for LLM ingestion | ⚠️ headings & lists | ❌ discarded | ✅ | ~30 MB |
| marker-pdf | Markdown/JSON via ML | ✅ | ❌ discarded | ❌ model-dependent | GB-scale models, GPU-friendly |
| docling | Markdown/HTML/JSON via ML | ✅ | ❌ discarded | ❌ model-dependent | GB-scale models |
| unstructured | Element JSON for RAG | ⚠️ element types | ❌ | ⚠️ | heavy optional deps |
| Adobe PDF Services | Structured JSON/HTML | ✅ | ⚠️ | ❌ | cloud API, paid |
When to choose pdf-html — you want a readable, reflowable document that still looks like the original, produced offline, reproducibly, with text you can trust character-for-character (text-based PDFs, tables included, even across page breaks).
When to choose something else — you need pixel-perfect visual replicas (pdf2htmlEX), OCR-heavy scanned document conversion (marker, docling), or RAG-oriented element JSON (unstructured).
🧪 Testing
56 tests cover every pipeline stage plus end-to-end CLI runs over deterministic fixture PDFs (report, two-column paper, slide deck, brochure, ruled table):
uv sync # dev deps (pytest, reportlab)
uv run pytest # run the suite
uv run python tests/fixtures/make_fixtures.py # regenerate fixture PDFs
Verbatim-ness is asserted mechanically: every source string drawn into a fixture must appear in the rendered HTML.
📁 Project structure
pdf-html/
├── src/pdf_html/
│ ├── cli.py # argparse CLI → pipeline → HTML
│ ├── extractor.py # PyMuPDF span + table-region extraction
│ ├── header_footer.py # repeated page-furniture stripping
│ ├── style_profiler.py # font-size histogram → style profile
│ ├── reading_order.py # column clustering & span ordering
│ ├── table_reconstructor.py # cell assignment, per-cell pipeline, row merging
│ ├── structure.py # heading / paragraph / list classification
│ ├── list_parser.py # nested list folding
│ ├── ast.py # typed document model
│ └── renderer.py # semantic HTML5 + derived CSS
├── tests/ # 56 tests + deterministic PDF fixtures
├── README.md
└── TUTORIAL.md # step-by-step usage guide
🛣️ Roadmap
- Ruled-table reconstruction with cross-page row merging
- Column-aware reading order with card-grid detection
- Repeated header/footer stripping
- Borderless-table detection (whitespace-gap heuristic)
- Callout/aside detection (
--no-calloutsflag already reserved) - Dependency-free
pdftotextfallback extractor -
colspan/rowspanfrom merged-cell geometry
🤝 Contributing
Contributions welcome! Ground rules:
- 🐍 Python 3.10+, type hints throughout
- 📦 PyMuPDF is the only hard runtime dependency
- 🧩 Keep heuristics small, pure, and individually testable; thresholds as named constants
- ✅ One feature per commit (
feat|fix|docs|refactor|chore: ...); update README/TUTORIAL with any user-facing change - 🧪
uv run pytestmust stay green — fixtures are the contract
🚀 Releases
This project is intentionally published as an early-stage v0.x tool: the core pipeline is solid for text-based PDFs, but it is not a universal PDF converter for scanned pages, forms, or OCR-heavy corpora.
# 1. bump version in pyproject.toml
uv build # 2. artifacts land in dist/
uv publish # 3. push to PyPI (or twine upload dist/*)
📄 License
MIT — see pyproject.toml.
Built with 🐍 + 📐 — early-stage geometry over guesswork.
If this project helped you, consider giving it a ⭐!
Metadata
Release files for pdf-html 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdf_html-0.1.1.tar.gz | 1.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdf_html-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.3 MB
Release files / pdf_html-0.1.1.tar.gz
| Download URL | pdf_html-0.1.1.tar.gz |
|---|---|
| Size | 1.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e5dfaba91491d6833ec1cdaa185afb9ed4cb37527d3a42231daab57439e83442
|
|
BLAKE2b-256 checksum How to use checksums |
bed0b8825b382598ac9780b503402fcc815b022840f49b169016c16834b9841f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.11
|
Release files / pdf_html-0.1.1-py3-none-any.whl
| Download URL | pdf_html-0.1.1-py3-none-any.whl |
|---|---|
| Size | 33.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
db7c6f8b4cbc4a96f3c4522b69f62e2c9cccab193b6bd9965ddc15a56062db77
|
|
BLAKE2b-256 checksum How to use checksums |
a798a8364aebbadde9a2b0c3dae5b9a3205e87e4e48bc66ae7083eeff2fdf3d2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.11
|