Skip to main content
pdf-html logo

pdf-html

Early-stage geometry-first PDF → HTML converter for text-based PDFs — verbatim text, source typography, zero images.

Python PyMuPDF License: MIT Tests Typed No LLM

One command in. One elegant, dependency-free HTML file out.


🎯 What it does

pdf-html is an early-stage, geometry-first PDF-to-HTML converter for text-based PDFs. It rebuilds the document as a single self-contained HTML5 file that mirrors the source document's look and structure:

  • ✅ Best for: text-based PDFs with real text layers, headings, lists, tables, and multi-column layouts

  • ⚠️ Not a universal OCR-first converter: scanned/image-heavy PDFs still need an OCR pre-pass or a dedicated workflow

  • 🏷️ Heading hierarchy — font-size tiers become real <h1>–<h6>

  • 🎨 Typography & color — the page CSS is derived from the document's own fonts, sizes, and palette

  • 📊 Tables — ruled tables are rebuilt as real <table> elements, including lists inside cells and rows that continue across page breaks

  • 📝 Lists — bullets and numbered items become nested <ul>/<ol> (indent decides nesting)

  • 🧭 Reading order — multi-column layouts are re-linearized column by column

  • ✂️ Page furniture — repeated running headers/footers and page numbers are stripped

  • 🔒 Text is verbatim — never summarized, reworded, or reordered within a block

  • 🚫 No images, ever — image content is dropped by design; text alone carries the document

🧭 Early-stage scope

This is an early-stage, geometry-first conversion tool for text-based PDFs. It is designed to be reliable and deterministic where the PDF has a usable text layer, but it is not a universal OCR-first converter for scanned documents, forms, or arbitrary image-heavy PDFs.

What it does well

  • text-based PDFs with real text layers
  • styled HTML with verbatim text
  • heading hierarchy, lists, and multi-column layout reconstruction
  • ruled table reconstruction with cell-aware content flow

What it does not yet do well

  • OCR and scanned-document support are planned as separate, opt-in paths
  • arbitrary image-heavy PDFs or forms with weak text layers
  • pixel-perfect visual reconstruction of every layout edge case

⚡ Quick start

# install (Python 3.10+)
uv pip install .        # or: pip install .

# convert
pdf-html report.pdf -o report.html

# open report.html in any browser — no external assets needed

🖥️ CLI reference

pdf-html INPUT.pdf -o out.html [--extractor pymupdf|pdftotext]
    [--style auto|default] [--paginate] [--no-tables] [--no-callouts]
    [--keep-headers] [--allow-scanned]
Flag Default What it does
-o, --output (required) Path of the HTML file to write
--extractor pymupdf Text extraction backend (pdftotext fallback planned)
--style auto auto derives CSS from the document's own fonts/sizes/colors; default uses a clean built-in theme
--paginate off Wrap each PDF page in a <section class="sheet">
--no-tables off Disable table reconstruction (table text flows as paragraphs)
--keep-headers off Keep repeated running headers/footers
--allow-scanned off Convert scanned/image PDFs instead of exiting with an OCR hint

🏗️ Architecture

A deterministic pipeline — every stage is a small, pure, individually testable module:

🔍 Pipeline diagram — click to enlarge (click again to close) · open full screen ↗
---
config:
  layout: elk
  theme: neutral
---
flowchart LR
    subgraph S1["📥 1 · Extract"]
        direction TB
        PDF@{ shape: doc, label: "📄 PDF" }
        EX@{ shape: rect, label: "Extractor<br/>spans + font metadata" }
        PDF --> EX
    end
    subgraph S2["🧹 2 · Clean &amp; Profile"]
        direction TB
        HF@{ shape: rect, label: "Header/Footer<br/>stripping" }
        SP@{ shape: hex, label: "Style Profiler<br/>body size · h1–h6 · palette" }
        HF --> SP
    end
    subgraph S3["🧭 3 · Layout"]
        direction TB
        RO@{ shape: rect, label: "Reading Order<br/>column clustering" }
        TR@{ shape: fr-rect, label: "Table Reconstructor<br/>per-cell pipeline" }
    end
    subgraph S4["🏗️ 4 · Structure"]
        direction TB
        SD@{ shape: div-rect, label: "Structure Detector<br/>headings · paragraphs · lists" }
        LP@{ shape: rect, label: "List Parser<br/>nested ul/ol" }
        AST@{ shape: bow-rect, label: "AST<br/>Document → Page → Blocks" }
        SD --> LP --> AST
    end
    subgraph S5["🎨 5 · Render"]
        direction TB
        RN@{ shape: rect, label: "Renderer<br/>semantic HTML5 + inline CSS" }
        OUT@{ shape: tag-doc, label: "🌐 out.html" }
        RN --> OUT
    end
    EX --> HF
    SP --> RO
    SP --> TR
    RO --> SD
    TR --> SD
    AST --> RN
Module Responsibility
extractor.py TextExtractor ABC; PyMuPDF backend reads per-span size/weight/color/bbox and detects table regions
header_footer.py Strips spans repeating on ≥ 60% of pages in the top/bottom 10% bands
style_profiler.py Character-weighted font-size histogram → body size, heading tiers, color palette
reading_order.py Column detection via x-gap clustering, with a card-grid fallback and straddle guard
table_reconstructor.py Assigns spans to detected cells, runs the full pipeline inside each cell, merges cross-page rows
structure.py Classifies lines into headings/paragraphs/list items from geometry + font cues
list_parser.py Indent-based nesting; glyph style only picks ul vs ol; markers stripped, text verbatim
ast.py Typed document model — Document → Page → Block, runs carry inline style
renderer.py Single-file HTML5 with one <style> block and CSS variables from the profile

📊 Table reconstruction highlights

The hardest part of PDF → HTML is tables. pdf-html:

  1. 🔍 Detects ruled tables geometrically (PyMuPDF find_tables()) at extraction time
  2. 📌 Assigns the page's styled spans to cells by bounding box — inline bold/color/size survive
  3. 🔄 Runs the normal line → paragraph → list pipeline inside every cell, so bullets in cells become real nested lists
  4. 🧵 Merges rows that continue across page breaks (empty-first-cell fragments) back into one row — even resuming mid-list-item
  5. 🏷️ Promotes a bold-only first row to a <th> header row

🧭 Design principles

Principle Meaning
🧮 Pure geometry, no AI All structure is inferred from font metadata and bounding boxes. No NLP, no LLM, no document-specific regexes
🔒 Text is sacred Output text is verbatim; only & < > are escaped
🚫 No images Spans overlapping image rects are dropped; <img> is never emitted
🪂 Graceful degradation Heuristic failures only affect styling — never text content or order
🔧 Tunable & testable Every heuristic threshold is a named module-level constant with focused unit tests

⚖️ How it compares

Every PDF converter picks a trade-off. pdf-html optimizes for semantic, reflowable, styled HTML with a verbatim-text guarantee — a square none of the established tools occupy:

Tool Output Semantic structure Keeps typography Deterministic Footprint
pdf-html Self-contained HTML5 ✅ h1–h6, ul/ol, table ✅ CSS derived from the source ✅ ~30 MB (PyMuPDF only)
pdf2htmlEX Pixel-faithful HTML ❌ positioned glyphs ✅ visually ✅ C++ toolchain
Poppler pdftohtml Positioned divs / bare text ❌ ⚠️ partial ✅ system package
pymupdf4llm Markdown for LLM ingestion ⚠️ headings & lists ❌ discarded ✅ ~30 MB
marker-pdf Markdown/JSON via ML ✅ ❌ discarded ❌ model-dependent GB-scale models, GPU-friendly
docling Markdown/HTML/JSON via ML ✅ ❌ discarded ❌ model-dependent GB-scale models
unstructured Element JSON for RAG ⚠️ element types ❌ ⚠️ heavy optional deps
Adobe PDF Services Structured JSON/HTML ✅ ⚠️ ❌ cloud API, paid

When to choose pdf-html — you want a readable, reflowable document that still looks like the original, produced offline, reproducibly, with text you can trust character-for-character (text-based PDFs, tables included, even across page breaks).

When to choose something else — you need pixel-perfect visual replicas (pdf2htmlEX), OCR-heavy scanned document conversion (marker, docling), or RAG-oriented element JSON (unstructured).

🧪 Testing

56 tests cover every pipeline stage plus end-to-end CLI runs over deterministic fixture PDFs (report, two-column paper, slide deck, brochure, ruled table):

uv sync                                        # dev deps (pytest, reportlab)
uv run pytest                                  # run the suite
uv run python tests/fixtures/make_fixtures.py  # regenerate fixture PDFs

Verbatim-ness is asserted mechanically: every source string drawn into a fixture must appear in the rendered HTML.

📁 Project structure

pdf-html/
├── src/pdf_html/
│   ├── cli.py                  # argparse CLI → pipeline → HTML
│   ├── extractor.py            # PyMuPDF span + table-region extraction
│   ├── header_footer.py        # repeated page-furniture stripping
│   ├── style_profiler.py       # font-size histogram → style profile
│   ├── reading_order.py        # column clustering & span ordering
│   ├── table_reconstructor.py  # cell assignment, per-cell pipeline, row merging
│   ├── structure.py            # heading / paragraph / list classification
│   ├── list_parser.py          # nested list folding
│   ├── ast.py                  # typed document model
│   └── renderer.py             # semantic HTML5 + derived CSS
├── tests/                      # 56 tests + deterministic PDF fixtures
├── README.md
└── TUTORIAL.md                 # step-by-step usage guide

🛣️ Roadmap

  • Ruled-table reconstruction with cross-page row merging
  • Column-aware reading order with card-grid detection
  • Repeated header/footer stripping
  • Borderless-table detection (whitespace-gap heuristic)
  • Callout/aside detection (--no-callouts flag already reserved)
  • Dependency-free pdftotext fallback extractor
  • colspan/rowspan from merged-cell geometry

🤝 Contributing

Contributions welcome! Ground rules:

  • 🐍 Python 3.10+, type hints throughout
  • 📦 PyMuPDF is the only hard runtime dependency
  • 🧩 Keep heuristics small, pure, and individually testable; thresholds as named constants
  • ✅ One feature per commit (feat|fix|docs|refactor|chore: ...); update README/TUTORIAL with any user-facing change
  • 🧪 uv run pytest must stay green — fixtures are the contract

🚀 Releases

This project is intentionally published as an early-stage v0.x tool: the core pipeline is solid for text-based PDFs, but it is not a universal PDF converter for scanned pages, forms, or OCR-heavy corpora.

# 1. bump version in pyproject.toml
uv build        # 2. artifacts land in dist/
uv publish      # 3. push to PyPI (or twine upload dist/*)

📄 License

MIT — see pyproject.toml.


Built with 🐍 + 📐 — early-stage geometry over guesswork.

If this project helped you, consider giving it a ⭐!

Metadata

Release files for pdf-html 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-html 0.1.0
File Size Uploaded
pdf_html-0.1.0.tar.gz 1.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdf-html 0.1.0
File Interpreter ABI Platform
pdf_html-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / pdf_html-0.1.0.tar.gz

Download URL pdf_html-0.1.0.tar.gz
Size 1.4 MB
Tags Source
SHA-256 checksum
How to use checksums
a6ff66b420d39163399fea36983a6291d417d1ddd8c5327900608790b77b33f4
BLAKE2b-256 checksum
How to use checksums
3400415e2688c79827bf5bb4ae128bec60a7afac11e32a79bbeca33de13f001f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release files / pdf_html-0.1.0-py3-none-any.whl

Download URL pdf_html-0.1.0-py3-none-any.whl
Size 33.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
df1a0fd8b76f0091b0bec57f00f306ea749372f33e56fb4645e79f494fb3d50a
BLAKE2b-256 checksum
How to use checksums
566f7173cd76565b04043c7d76d6639a39fe517366ecd38234efeadce55a686c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page