Skip to main content

foliodoc

Fast document conversion for Python (PDF / scans / images / DOCX → Markdown, JSON, text) on Windows, Linux and macOS. It solves the same problem as docling with no layout neural network and no GPU: exact text from the PDF text layer, OCR only where a page has none, and geometric reconstruction of reading order, headings, lists and tables. See Benchmarks for where that wins and where it doesn't.

from foliodoc import convert

doc = convert("report.pdf")        # or .png/.jpg/.tiff scan, or .docx
print(doc.to_markdown())
doc.tables                         # [[["Item", "Qty"], ["Apple", "3"]], ...]
doc.to_json()                      # blocks with type, text, bbox, page, heading level, cells
foliodoc report.pdf scan.png -o out/ --to md --timings

Install

pip install foliodoc              # digital PDFs and Word files
pip install "foliodoc[ocr]"       # + OCR for scans/images: Apple Vision on macOS, RapidOCR on every OS

Python 3.10+ on Windows, Linux and macOS. Individual engines: foliodoc[apple], foliodoc[rapidocr], foliodoc[tesseract] (the last also needs the tesseract binary). Documentation: https://meet2147.github.io/foliodoc/

How it gets its accuracy

Problem What foliodoc does
Most PDFs already contain the exact text Reads the text layer through pdfium: every character comes with its box, font size and weight. No OCR, 100% exact, about 10 ms per page. OCR is used only for pages with no text layer or a broken font encoding, and for large embedded images that have no text drawn over them.
Scanned pages ocr="auto" picks the best engine installed: Apple Vision on macOS (Neural Engine, on-device), RapidOCR (PaddleOCR v5 on ONNX Runtime, CPU) on Windows, Linux and macOS, then Tesseract. You can plug in your own engine with register_ocr_engine. Pages stream through a bounded pool of engines.
Skewed scans Projection-profile deskew (coarse-to-fine search, about 30 ms), applied only when it clearly helps.
OCR skips lines or returns half a line Every ink line on the page is checked against the OCR boxes. Lines that were skipped or only partly read are re-read from a tight crop, and the result replaces the old reading only if it is clearly more complete. It never adds duplicates.
OCR merges two table cells into one "line" Lines are split wherever the pixels show a real gap that is wide compared with the line's own word spacing. Justification stretches all spaces in a line equally, so prose isn't split. Boxes are then snapped to the ink.
OCR gives no font information Stroke thickness (average run length of ink pixels) separates bold and large headings from body text far more reliably than box height.
Reading order in multi-column layouts Recursive XY-cut that refuses horizontal cuts through a multi-column zone. Tables and figures are treated as single units.
Tables Ruled grids come from vector lines (or morphological line detection on scans). Rules-only "booktabs" tables are supported. Tables without rules are accepted only if the rows share column gutters and every column is aligned on its left edge, right edge or center, which prose never is.
Tables broken across pages or columns Stitched back together, including a lone header row, repeated headers, and continuation rows.
Page furniture Repeated headers and footers, matched fuzzily to tolerate OCR noise, plus page numbers are dropped from the body.
Hyphenation "experi-\nment" → "experiment".

Benchmarks

Two benchmarks, both reproducible from bench/. Every number below comes from the final code. Both systems ran one after the other on an Apple M2, warm (model loading excluded), with default settings unless a row says otherwise. Docling is version 2.63.

Summary

Where foliodoc wins Where docling wins
Speed Digital PDFs: ~20× faster (0.08 vs 1.66 s/page). Scans with Apple Vision on both: 5–10× faster. Cross-platform OCR on small forms (FUNSD): docling 4.2 s vs foliodoc 6.7 s per form.
Text accuracy Digital PDFs (slightly), scanned forms and receipts with Apple Vision, receipt key fields (total found 97% vs 77–93%). Real page images from OmniDocBench (books, papers, slides), and forms with the cross-platform OCR.
Structure Synthetic tables split across pages or columns. Real-world layout: tables, headings, list items, header/footer removal (DocLayNet), and table structure (OmniDocBench). Docling's neural layout and table models are much better here.
Robustness 1 failure in 7,822 documents. 214 failures in 7,822: docling rejects large images (over Pillow's 179-megapixel "decompression bomb" limit after its internal upscaling).

foliodoc is not more accurate than docling across the board. It is much faster, it reads text as well or better on digital PDFs, forms and receipts, and it loses clearly on layout and table structure in real documents.

Real-world: 7,822 public documents

Dataset Docs What it is Ground truth
DocLayNet v1.2 test 4,999 Digital PDF pages: financial reports, manuals, papers, laws, tenders, patents Human layout labels + the page's text
OmniDocBench 1,651 Page images: books, papers, slides, notes, newspapers (English and Chinese) Reading-order text + table HTML
SROIE 2019 973 Scanned receipts Text lines + company/date/address/total
FUNSD 199 Noisy scanned forms Words

Default configuration (on macOS both systems use Apple Vision for OCR; docling through ocrmac):

Dataset Metric foliodoc docling
DocLayNet (4,999 digital PDFs) Text token F1 0.939 0.931
Table / heading / list-item detection F1 (IoU ≥ 0.5) 0.35 / 0.41 / 0.36 0.86 / 0.86 / 0.87
Page headers/footers leaking into text ↓ 32% 7%
Seconds per page (mean / median) 0.08 / 0.03 1.66 / 0.82
OmniDocBench, English pages Text edit distance ↓ (all 755 pages) 0.288 0.282
Text edit distance ↓ (703 pages both processed) 0.298 0.229
Table TEDS (703 pages both processed) 0.343 0.493
OmniDocBench, all 1,651 pages Seconds per page 0.99 7.94
SROIE (973 receipts) Token F1 0.691 0.595
Token F1 (840 receipts both processed) 0.702 0.640
Total / date found (840 both processed) 97.1% / 95.8% 89.5% / 89.3%
Seconds per receipt 0.41 4.01
FUNSD (199 forms) Token F1 0.899 0.856
Seconds per form 0.46 2.26
All Failed documents 1 214

Cross-platform configuration. Both systems use RapidOCR (PaddleOCR models on ONNX Runtime, CPU), the OCR they use on Windows and Linux. This ran on a subset: all FUNSD forms, the SROIE test split, and the English and mixed-language OmniDocBench pages that both finished (493). DocLayNet needs no OCR, so its results above apply on every OS.

Dataset Metric foliodoc + RapidOCR docling + RapidOCR
FUNSD (199) Token F1 0.810 0.844
Seconds per form 6.71 4.18
SROIE test (347) Token F1 0.601 0.620
Total / date found 97.4% / 91.6% 93.4% / 91.9%
Seconds per receipt 4.66 5.71
OmniDocBench English + mixed (493) Text edit distance ↓ 0.324 0.296
Table TEDS 0.461 0.697
Seconds per page 8.56 11.81

Notes on the real-world numbers:

  • 102 DocLayNet pages are excluded from the text metric for both systems because DocLayNet's own ground-truth text is garbled on them (PDFs with broken font encodings, e.g. Ó Ç ä ä Ê).
  • Text inside figures is excluded on both sides. Docling's layout model was trained on DocLayNet, so DocLayNet is home ground for it.
  • Chinese pages: both systems ran with default (Latin-script) OCR settings and score poorly on them. foliodoc accepts languages=["zh-Hans", "en-US"] for Apple Vision; that setting was not benchmarked.
  • A failed document counts as empty output. Every document that crashed or hung was recorded as a failure; none were silently skipped.
  • The cross-platform run was cut short to save time: OmniDocBench covers 493 of the planned 871 pages. Both systems are scored on exactly the same pages.
  • Bugs found in foliodoc while running this benchmark were fixed before the final numbers: effective font size, word spacing in PDFs without space glyphs, letter-spaced and rotated text, page-border frames, content outside the visible page, a pdfium threading crash, a Vision stall, and an XY-cut infinite loop on mirrored glyphs.

Reproduce:

python bench/real_prep.py funsd sroie doclaynet omnidocbench    # after downloading the datasets to bench/real/
python bench/real_run.py foliodoc doclaynet                        # resumable; one JSON line per document
bench/.venv-docling/bin/python bench/real_run.py docling doclaynet
python bench/real_run.py folio-rapidocr funsd                   # cross-platform OCR configuration
python bench/real_score.py > bench/real/results.json

Synthetic: exact ground truth

bench/corpus.py generates documents where every character, heading and table cell is known: reports, two-column papers, invoices and dense pages, each as a digital PDF, a clean 300 dpi scan, and a noisy 200 dpi scan (skew, blur, noise, JPEG). The table shows the held-out set (seed 1000, generated after tuning; 36 files, 78 pages).

Input System Character error ↓ Word F1 Table structure Table cells Headings F1 Sec/page ↓
Digital PDF foliodoc 0.03% 1.000 0.917 0.972 1.000 0.011
docling 9.70% 0.987 0.528 0.858 1.000 1.65
Clean scan (Apple Vision) foliodoc 0.38% 0.998 0.917 0.965 0.989 0.49
docling 1.61% 0.989 0.528 0.847 0.996 1.55
Noisy scan (Apple Vision) foliodoc 0.17% 0.996 0.917 0.969 0.996 0.42
docling 1.75% 0.993 0.528 0.854 1.000 2.11
Clean scan (RapidOCR) foliodoc 0.99% 0.989 0.917 0.958 0.945 4.7
docling 7.70% 0.928 0.528 0.827 0.986 5.9
Noisy scan (RapidOCR) foliodoc 0.25% 0.994 0.917 0.939 0.985 5.8
docling 8.00% 0.946 0.528 0.815 1.000 5.8

On these synthetic documents docling's character error is high for two reasons confirmed by diffing: it reorders blocks in two-column layouts, and on justified text it cuts off the last word of some lines ($94,118.11 → $94,11). The synthetic set is narrower than real documents, which is why the real-world results above are the ones to trust.

Layout

foliodoc/pdf.py       pdfium text layer (chars → segments with words), vector rulings, images, scan extraction
foliodoc/ocr.py       OCR engines behind one interface (apple / rapidocr / tesseract)
foliodoc/raster.py    deskew, rule detection, weak-line re-read, ink-gap splitting, stroke-based styling
foliodoc/layout.py    rows, ruled/unruled tables, XY-cut reading order, paragraphs, headings, lists, furniture, stitching
foliodoc/docx.py      DOCX straight from XML
foliodoc/convert.py   orchestration, streaming OCR pool
bench/             corpus generator, scorer, runner, per-document inspector

Metadata

Release files for foliodoc 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for foliodoc 0.1.0
File Size Uploaded
foliodoc-0.1.0.tar.gz 41.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for foliodoc 0.1.0
File Interpreter ABI Platform
foliodoc-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 79.4 kB

Release files / foliodoc-0.1.0.tar.gz

Download URL foliodoc-0.1.0.tar.gz
Size 41.9 kB
Tags Source
SHA-256 checksum
How to use checksums
9f6ff26d35eaac23fc3406a5f7b5f8111dfeed02ebcb4be80bfa86babcd0d12c
BLAKE2b-256 checksum
How to use checksums
ed30aca34a317b08948721c0e63283ed9fbacbd5b369d8d13b8040778c0b811a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.8

Release files / foliodoc-0.1.0-py3-none-any.whl

Download URL foliodoc-0.1.0-py3-none-any.whl
Size 37.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0f2e9315b6dec0c788e3e6d0d35a77ed878ecd3c69bd1021d288447e71a25f80
BLAKE2b-256 checksum
How to use checksums
1e3d8474117a14f85f2cced4438575cf8b57d8cbfce1dc15ad3ea97496a6d92a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page