Warp-Ingest
Warp-Ingest is a state-of-the-art, deterministic PDF parser. It turns a PDF into layout-aware structure — accurate word- and block-level bounding boxes, structural labels (section header, paragraph, list item, table row), and the relationships between them (the section/sub-section/paragraph parent ↔ child hierarchy) — and renders it as blocks, JSON, HTML, or an OpenContracts structural export.
It is rule-based, not model-based: structure comes from text coordinates, graphics, and font data — no GPU, no per-page rasterization, no training set. That makes it fast and predictable on long, text-layer documents (hundreds of pages), which is where vision parsers are slowest and least stable. See Rule-based vs. model-based for the rationale.
Warp-Ingest is a pure-Python rewrite of the nlmatics nlm-ingestor engine, with the
Java/Apache-Tika and Tesseract dependencies removed.
How the pipeline works
PDF ──pdfplumber──► per-word boxes + fonts ──► Tika-format XHTML ──► visual_ingestor ──► blocks / JSON / HTML / OpenContracts
(scanned pages ──rapidocr──► OCR words ──► same XHTML ──┘)
pdfplumber(MIT) extracts each word's real bounding box and font data in absolute, top-left-origin PDF points.- A scanned or sparse page is detected automatically and routed to the optional
rapidocr-onnxruntimeOCR backend (Apache-2.0, no Tesseract binary, no GPU), which emits the same word-box format — so a scanned page and a born-digital page flow through the identical layout engine. - The front-end emits an intermediate Tika-format XHTML (one
<p>per visual line, carrying per-word positions and fonts).visual_ingestor— the ~6,000-line rule engine — groups those lines into typed blocks, detects tables from rule-line graphics, strips repeating headers/footers, removes watermarks, and fixes reading order.
The word-level boxes use the same technique as OpenContracts (see docs/bbox_architecture.md).
What the parser produces
- Sections and sub-sections with their nesting levels.
- Paragraphs (lines joined into coherent blocks).
- The parent ↔ child links between sections and paragraphs.
- Tables, with the section each table sits in.
- Lists and nested lists.
- Content joined across page breaks.
- Removal of repeating headers and footers.
- Watermark removal.
- OCR with bounding boxes for scanned pages.
- An OpenContracts structural export: PAWLS
word tokens, one structural annotation per block, and the
parent_idheading hierarchy as explicit relationships.
Benchmarks
Warp-Ingest is scored on the official LlamaIndex ParseBench — 2,078 human-verified pages of real enterprise documents — by running it through the official framework with deterministic, rule-based metrics (no LLM-as-a-judge). Among the 8 deterministic, local, no-API parsers, Warp-Ingest is a top-2 result (2nd overall), it is the only local parser that carries real visual grounding, and it leads on Charts. Full numbers, methodology, and the reproduction commands are in benchmarks/parsebench/RESULTS.md.
Installation
# parser-only install
pip install "warp-ingest[parser]"
# hosted service runtime (FastAPI/uvicorn)
pip install "warp-ingest[service]"
# full service runtime with OCR support
pip install "warp-ingest[all]"
# install the project and dev/test tools with uv
# (dev includes service dependencies because the test suite covers the API)
uv sync --group dev
# include the optional OCR backend for scanned PDFs
uv sync --group dev --extra ocr
# full service runtime with OCR
uv sync --group dev --extra all
# one-time NLTK data download
uv run python -m nltk.downloader punkt punkt_tab stopwords
Running the service
# install with the `service` or `all` extra first
python -m warp_ingest.ingestion_daemon # or: ./run.sh (FastAPI/uvicorn, port 5001)
The launcher budgets concurrency automatically from the CPUs actually available
(CPU affinity and the container cgroup quota, so a docker --cpus / K8s
limits.cpu deployment uses exactly its slice): one uvicorn worker per
effective CPU, with the front-end page-striping pool and the OCR session
threads sized so the layers never oversubscribe the box. Every knob can be
overridden via environment variables:
| Env var | Default | Effect |
|---|---|---|
WARP_API_KEY |
abc123 |
API key required by /api/parse (send as X-API-Key or Authorization: Bearer) |
WARP_WEB_WORKERS (or WEB_CONCURRENCY) |
effective CPUs | uvicorn worker processes |
WARP_FE_WORKERS |
max(1, min(8, cpus // workers)) |
per-worker front-end page-striping pool (≤1 = serial) |
WARP_OCR_THREADS |
max(1, cpus // (workers × fe_workers)) |
onnxruntime intra-op threads per OCR session |
WARP_WORKER_PARSE_SLOTS |
1 |
concurrent parses allowed inside one worker |
WARP_HOST / WARP_PORT (or PORT) |
0.0.0.0 / 5001 |
bind address |
POST /api/parse with a file form field and the API key; the response body is
the parse result ({"page_dim": ..., "num_pages": ..., "result": ...}) and
errors are standard {"detail": ...} bodies. Interactive OpenAPI docs at /docs.
curl -H "X-API-Key: abc123" -F file=@document.pdf \
"http://localhost:5001/api/parse?render_format=all"
Query parameters (booleans accept true/false, 1/0, yes/no, on/off):
| Param | Values | Effect |
|---|---|---|
render_format |
all | json | html | opencontracts |
output rendering |
apply_ocr |
bool | force OCR on every page (scanned pages are OCR'd automatically regardless) |
disable_ocr |
bool | keep every page on its embedded text layer — no OCR for this request |
semantic_units |
bool | append the additive Semantic-Unit clause layer (render_format=opencontracts) |
GET / and GET /healthz are unauthenticated health endpoints (/healthz
reports the resolved concurrency settings and OCR availability). Warp-Ingest
parses PDF only; a non-PDF upload returns HTTP 415.
Docker
docker build -t warp-ingest .
docker run -p 5010:5001 -e WARP_API_KEY=change-me warp-ingest
Every GitHub release publishes a versioned multi-arch image to GHCR
(ghcr.io/open-source-legal/warp-ingest:X.Y.Z / :X.Y / :X / :latest,
cosign-signed) via .github/workflows/docker-publish.yml.
Library use
from warp_ingest.ingestor import pdf_ingestor
export = pdf_ingestor.parse_to_opencontracts("document.pdf") # OpenContracts export
markdown = pdf_ingestor.parse_to_markdown("document.pdf") # Markdown export
payload = pdf_ingestor.parse_to_markdown_payload("document.pdf") # Blocks + tables + geometry
layout = pdf_ingestor.parse_to_layout_predictions("document.pdf") # Generic layout predictions
ingestor = pdf_ingestor.PDFIngestor("document.pdf", {"render_format": "all"})
blocks = ingestor.blocks
The notebook pdf_visual_ingestor_step_by_step.ipynb walks the whole pipeline on a sample PDF.
Testing
No Java or Tika is needed.
make test # full pytest suite (unit + fixture parsing, incl. OCR)
uv run pytest tests/ # same, directly
Beyond unit tests, the suite includes cross-engine regression against the original Java/Tika engine (S-1), an OpenContracts-export regression, a Docling layout oracle, and vision-adjudicated structural-correctness suites — all floored against committed baselines so engine changes can only improve, never silently regress. See CLAUDE.md for the map of which suite guards what.
Rule-based vs. model-based
Over four years the nlmatics team evaluated many options, including a YOLO-based vision parser, and settled on the rule-based approach for these reasons:
- Speed. It is ~100× faster than a vision parser, which must rasterize every page (even text-layer ones). A vision parser is the better tool for scanned PDFs with no text layer or small form-like documents; for large text-layer PDFs spanning hundreds of pages, a rule-based parser is far more practical.
- No special hardware. It runs on CPU.
- Fixable. Vision-parser errors are fixed either by adding training examples (which can degrade previously-correct behavior) or by layering on rules anyway — at which point you are writing rules again.
Credits
The PDF parser visual_ingestor and its indent parsers were written by Ambika Sukla, with contributions from Reshav Abraham, Tom Liu (the original Indent Parser), and Kiran Panicker (parsing speed, table-parsing, indent-parsing, and reordering accuracy). The core line_parser was written by Ambika Sukla.
Thanks to the pdfplumber / pdfminer.six, pypdfium2, and RapidOCR open-source
communities, and to the Apache PDFBox and Tika developers whose XHTML format the engine
is built around.
History
Earlier versions depended on an nlmatics-modified Apache Tika (nlm-tika) on the JVM
plus Tesseract for OCR. That is gone — the pure-Python front-end
(pdf_plumber_parser.py and
ocr_parser.py) reproduces the same
intermediate XHTML contract, so the layout engine did not have to change.
Metadata
Release files for warp-ingest 2.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| warp_ingest-2.1.0.tar.gz | 779.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| warp_ingest-2.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.6 MB
Release files / warp_ingest-2.1.0.tar.gz
| Download URL | warp_ingest-2.1.0.tar.gz |
|---|---|
| Size | 779.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
edc190479af2bd6e1dfd171a6a354f33511c7062fcc2b442ddbb37c7b515f1bf
|
|
BLAKE2b-256 checksum How to use checksums |
32db3ccf58605ac388a7b7238731a4fed80f5e4cd942c4fb73a5c792a880a475
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency logRelease files / warp_ingest-2.1.0-py3-none-any.whl
| Download URL | warp_ingest-2.1.0-py3-none-any.whl |
|---|---|
| Size | 797.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
97c6d8ceefe78347ff9d9f28be784b3640d965dcd27ea1122fd9811cb213a4f1
|
|
BLAKE2b-256 checksum How to use checksums |
7227cad5330ca182fd20259d59dbc51d39bb0959f04b8305b3199d1c18ea69fb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency log