Skip to main content

VerifyDoc

The trust layer for document → structured-JSON extraction. Wrap any extractor — get back JSON where every field carries a calibrated confidence, a source grounding (page + bbox / char span), and an accept/review decision tuned to your error budget.

CI License: Apache-2.0 Python 3.11+ Code style: black

VerifyDoc demo: a silently-wrong total is caught by grounding and routed to review

Above: a real pipeline run (scripts/make_demo_gif.py). The extractor returned $1,432.50; the page says $1,234.50. Grounding support drops to 0.78, the field misses the accept threshold, and the reviewer is pointed at the exact source region. The other three fields are auto-accepted.

The problem

Modern document parsers read pages at 96%+ benchmark accuracy — and still emit fluent, plausible, silently-wrong values ($42.50$45.20) with no reliable per-field signal telling you which values to trust. Commercial APIs (Box, Azure, Textract) sell field-level confidence as a closed feature. No popular open-source parser leads with it. (full USP audit)

VerifyDoc doesn't compete with the parsers — it layers on top of any of them:

document + schema ─► ingest ─► extractor adapter ─► confidence ─► calibration
                                (any model)          signals       (fit on cal split)
                     ─► grounding ─► abstention policy ─► verified JSON + review UI
                        (bbox/span)   (target risk α)

At a chosen operating point, VerifyDoc auto-accepts as many fields as possible while holding the error rate among accepted fields below your target (e.g. ≤ 2%) — everything else is routed to review with its source location attached, so a human verifies in seconds instead of eyeballing every field.

Quickstart

pip install verifydoc          # core (text pipelines + eval harness)
pip install 'verifydoc[pdf]'   # + PDF/image ingestion
from verifydoc import verify

# ready-to-run sample lives in examples/
result = verify("examples/invoice.txt", schema="examples/invoice_schema.json")
for f in result.fields:
    print(f"{f.path:12} = {f.value!r:24} conf={f.confidence:.2f} {f.decision}")
    if f.grounding:
        print(f"             └─ page {f.grounding.page}, bbox {f.grounding.bbox}")
verifydoc extract examples/invoice.txt --schema examples/invoice_schema.json --threshold 0.8
streamlit run ui/streamlit_app.py     # review UI: green/red fields + click-through to source

See examples/ for the full runnable walk-through.

Schemas are plain JSON Schema, with each leaf optionally declaring how it is scored (the executable-schema pattern):

{
  "type": "object",
  "properties": {
    "invoice_id": {"type": "string"},
    "vendor":     {"type": "string", "x-scoring": "semantic"},
    "total":      {"type": "number", "x-numeric-tol": 0.01}
  }
}

What's inside

Layer Modules Status
Adapters (all model code isolated here) mock · text-search · RapidOCR · PaddleOCR · dots.ocr · Docling/MinerU output · API-VLM (OpenAI/Anthropic)
Confidence signals token-prob · verbalized · consensus (k-sample voting) · grounding-based · combined
Calibrators (fit on a dedicated split, never test) temperature · Platt · isotonic · histogram · split conformal with finite-sample risk guarantee
Grounding value → page/bbox/char-span attachment with support scores
Policy empirical & conformal accept thresholds for a target selective risk
Eval harness / VerifyDocBench scorer Field-F1 · exact · CER/WER · ANLS · TEDS/TEDS-Struct · GriTS · omission vs hallucination · ECE/Adaptive-ECE/MCE/Brier/NLL/TCE · RC/AURC/E-AURC/Coverage@Risk/AUROC/AUPR/FPR@95 · box IoU/span-F1/grounding-conditioned correctness · bootstrap CIs + paired tests

Every metric implements the exact definition in PROJECT.md §5 with a hand-computed numeric regression test (201 tests, eval/ coverage 98%).

Results on real documents

Two independent real OCR extractors (RapidOCR and PaddleOCR) on real CORD receipts and FUNSD forms, scored by the harness (full tables + reading):

Confidence signal ranks errors? CORD AUROC (RapidOCR / PaddleOCR)
learned combiner ✅ best 0.89 / 0.84
grounding ✅ strong 0.82 / 0.74
token-probability ~ moderate 0.69 / 0.68
verbalized / consensus ✗ uninformative 0.50 / 0.50
  • Grounding is a real trust signal: grounded fields are 84–85% correct vs ~1% for ungrounded (gap ≈ +0.84; box accuracy @IoU 0.5 = 0.75–0.78).
  • The abstention layer is honest: with a weak field-extractor the base error rate is high, so conformal abstention at a 2–5% budget correctly refuses to auto-accept — you report selective risk, not headline accuracy.
  • The synthetic slice (strong extractor) shows the other end: Coverage@2% ≈ 1.0.

The thesis holds on real data: grounding + a learned fusion rank errors; self-reported and single-sample-consensus confidence do not.

The benchmark

make results     # regenerates every table/figure in paper/generated from configs/

The harness runs signals × calibrators × the full metric suite with a document-level calibration split (disjointness asserted in code), bootstrap CIs, and a conformal-guarantee row. It ships a deterministic synthetic slice (runs in CI) plus CORD and FUNSD loaders with gold source boxes; extractor: dispatches to any adapter (rapidocr, paddleocr-vl, …) and dataset: to any slice. See the GPU runbook to reproduce the real-model rows. Core claims (grounding beats verbalized; conformal holds its guarantee) are also CI-enforced as unit tests, not just stated.

Why not just use the parser's own score?

Because it doesn't exist (Docling/MinerU/Marker), or it's a raw recognition score that was never calibrated against field-level correctness (PaddleOCR/dots.ocr). See docs/USP.md for the audit, and the reliability diagrams in paper/generated/ for what "calibrated" actually buys you.

Roadmap

  • v0.1 — library + CLI + harness + synthetic benchmark slice + UI
  • v0.2 — CORD + FUNSD real slices with gold boxes; learned combiner; 1000× faster grounder
  • v0.3 — real-model results (RapidOCR + PaddleOCR on CORD/FUNSD)
  • v0.4 — vendor-neutral API-VLM extractor (OpenAI/Anthropic) with k-sample consensus; compilable paper with auto-generated tables
  • Run the API-VLM row at scale (fair verbalized + consensus comparison; needs an API key)
  • dots.ocr via vllm; SROIE / DocILE / XFUND slices; human-labeled correctness + IAA
  • Paper submission (contributions welcome)

Documentation

Development

git clone https://github.com/bhaskargurram-ai/verifydoc && cd verifydoc
uv venv .venv && uv pip install -e ".[dev]"
make test lint typecheck     # all green before any PR (CI enforces)
make results                 # regenerate benchmark tables + LaTeX
make paper                   # compile the paper (needs a LaTeX toolchain)

Contributions welcome — see the issues tagged good-first-issue. All model-specific code goes in verifydoc/adapters/; a new extractor is one file.

Citation

@software{verifydoc2026,
  author = {Gurram, Bhaskar},
  title  = {VerifyDoc: Calibrated, Abstaining, Grounded Document Extraction},
  year   = {2026},
  url    = {https://github.com/bhaskargurram-ai/verifydoc}
}

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

verifydoc-0.4.0.tar.gz (493.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

verifydoc-0.4.0-py3-none-any.whl (70.9 kB view details)

Uploaded Python 3

File details

Details for the file verifydoc-0.4.0.tar.gz.

File metadata

  • Download URL: verifydoc-0.4.0.tar.gz
  • Upload date:
  • Size: 493.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for verifydoc-0.4.0.tar.gz
Algorithm Hash digest
SHA256 a853fd9d600844de1c844a10213f45ffb27dad76516f5f7662a1dc8a38dc592f
MD5 c0e33b561a6d3abf0c1daf2936d47334
BLAKE2b-256 0c3a8e648d07b5c9aeb0e06835c4c05f49028a7bb475a26a0f7f50dfb52e60a4

See more details on using hashes here.

Provenance

The following attestation bundles were made for verifydoc-0.4.0.tar.gz:

Publisher: release.yml on bhaskargurram-ai/verifydoc

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file verifydoc-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: verifydoc-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 70.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for verifydoc-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b395f194e71be9736fdd8b0d7842e0fd5a61e0474ecd78ca0cf31fe8b415d49f
MD5 fb19299d7992074e4c69f8ef6771a5a8
BLAKE2b-256 8d52b1b202d6710f80f37da5b8fb7d8629a535fe4fb088dff206604789da8ca6

See more details on using hashes here.

Provenance

The following attestation bundles were made for verifydoc-0.4.0-py3-none-any.whl:

Publisher: release.yml on bhaskargurram-ai/verifydoc

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

This release

0.4.0 This release

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page