Skip to main content

ExtractBench

Website arXiv Dataset License

ExtractBench is a benchmark for schema-guided extraction from enterprise documents. Given a document and a user-defined JSON Schema, a system must return schema-valid JSON with the correct values, every record of each repeated structure, missing fields marked null rather than invented, and source evidence for each value.

A schema defines one extraction task, and all documents of that type share it: one invoice schema covers invoices from every vendor, however different each one looks. Enterprises write a new schema for almost every workflow, so a system cannot be tuned to a fixed template. It has to handle schemas and documents it has never seen. And because agents increasingly act on extracted values before anyone reviews them, one truncated schedule or one invented value becomes a wrong payment or a wrong decision. The benchmark therefore scores completeness and traceability.

The benchmark covers 370 documents (4,869 pages) across 8 business domains and 67 document types, each type with its own schema. Every document is tagged along five independent axes: task challenge, perception challenge, table structure, length, and business domain, so a low score can be traced to its cause.

ExtractBench: JSON Schema and document in; JSON plus evidence out; scored on value F1, grounding F1, challenge tags, and cost

Leaderboard

Models and prices reflect each provider's official documentation; each system uses its recommended configuration.

Unified value F1 — the headline metric. Every score is an unweighted mean over documents; each document counts once, whatever its length. For raw data including per-split precision and recall, cost, and latency, see leaderboard.csv. Equal displayed Overall scores are ordered by lower cost per page. The best score in each Overall, Short, Medium, and Long column is bold; the second-best distinct score is underlined.

Rank Provider Category Overall Short Medium Long ¢ / Page
1 LlamaExtract Agentic Plus LlamaExtract 96.38 97.15 94.77 94.60 7.50¢
2 Reducto Deep Extract Specialized APIs 96.08 96.84 94.80 92.85 6.35¢
3 LlamaExtract Agentic LlamaExtract 96.03 96.81 94.21 95.05 3.12¢
4 Pulse (Effort) Specialized APIs 95.91 96.46 95.01 93.51 10.50¢
5 Claude Code (Sonnet 5.5 Evidence) Coding Agents 94.68 96.22 92.97 83.69 5.69¢
6 Codex (GPT-6 Sol Evidence) Coding Agents 94.57 95.99 92.42 87.25 9.80¢
7 Codex (GPT-5.6 Sol Evidence) Coding Agents 93.77 96.04 90.19 82.73 21.65¢
8 LlamaExtract Cost-Effective LlamaExtract 93.67 96.07 90.21 80.48 1.00¢
9 Codex (GPT-5.5) Coding Agents 93.57 95.68 91.15 78.88 27.83¢
10 Codex (GPT-5.5 Evidence) Coding Agents 93.35 95.59 88.72 87.97 33.63¢

Top 10 of 49 systems — full table in leaderboard.csv.

Grounding F1 — a field counts only when its value is accepted and it points at the right evidence: at word level the predicted box must overlap an accepted evidence box at IoU 0.5, at page level the cited page must be correct. Scored only over the documents that carry verified box ground truth.

RankProviderWord-level grounding F1Page-level grounding F1
OverallShortMediumLongOverallShortMediumLong
1LlamaExtract Agentic Plus81.2681.6381.1875.2289.9492.4586.5074.71
2LlamaExtract Agentic78.8977.6381.8680.8990.6492.5886.6887.09
3Codex (GPT-6 Sol Evidence)77.1177.6175.5378.3186.5087.2085.1284.82
4Claude Code (Opus 5.5 Evidence)71.5573.5765.5474.8776.9579.2771.3676.47
5Codex (GPT-6 Luna Evidence)65.7062.9171.1174.3683.9384.4183.4780.17
6Claude Code (Sonnet 5.5 Evidence)64.6164.5963.1871.2275.9977.9371.5074.98
7Reducto Extract (v4)54.9760.5349.3516.3075.2887.4757.3818.59
8Codex (GPT-5.6 Sol Evidence)54.6650.8764.5869.0183.6285.1280.7778.84
9LlamaExtract Cost-Effective53.6549.2064.8455.6680.0987.5366.4754.94
10Codex (GPT-5.5 Evidence)52.5753.1348.0161.8780.6483.2773.8779.90

Top 10 of 49 systems — full table in leaderboard.csv.

Inclusion criteria
  1. The model or API needs to be publicly accessible, either via open weights or a self-serve API that any user can sign up for.
  2. The benchmark run needs to finish within a reasonable time (roughly single-digit hours).
  3. We can adjust concurrency based on the provider's recommended settings, but providers should not require custom framework changes, so the evaluation stays fair across models.

Quick Start

Prerequisites: Create a .env file with the API key for the extraction system you want to evaluate (see Configuration).

# Install from PyPI (pick the extras for the systems you want to run)
pip install "llama-extract-bench[runners]"    # every provider SDK
pip install "llama-extract-bench[llamaextract]" # or just one, e.g. llamaextract / openai / anthropic / google

# Or, from a checkout of this repo
uv sync --extra runners

# Quick test run (6 documents — good for trying things out)
# (drop the `uv run` prefix if you installed from PyPI)
uv run extract-bench run llamaextract_agentic --test

# Full benchmark run (replace with any pipeline name, see "Available Pipelines" below)
uv run extract-bench run llamaextract_agentic

# View interactive reports in your browser
uv run extract-bench serve llamaextract_agentic
Rough cost of one full run

Costs use each provider's official listed price.

Category One full run Examples
Commercial VLM $10 – $49 GPT-5.4 Nano ~$10, Gemini 3.5 Flash ~$49
LlamaExtract $49 – $395 Cost-Effective ~$49, Agentic ~$152, Agentic Plus ~$395
Specialized APIs $170 – $487 Datalab ~$170, Reducto Deep Extract ~$309, Extend ~$487
Coding agents $787 – $1,355 Claude Code (Opus 4.8) ~$787, Codex (GPT-5.5) ~$1,355
Self-hosted open weights GPU time only Qwen3.6 35B, Gemma4 26B, NuExtract3, Lift Datalab 9B

These figures apply the reported mean cost per page to 4,869 pages. Per-system, per-split prices are in leaderboard.csv (Cost_Per_Page, Cost_Short, Cost_Medium, Cost_Long, in dollars per page).

Available Pipelines

A pipeline is an extraction system or configuration you want to evaluate. Run uv run extract-bench pipelines for the live list, or see docs/pipelines.md.

Paper baselines (the 14 systems on the leaderboard)
Pipeline name Name in paper
llamaextract_agentic_plus LlamaExtract Agentic Plus
llamaextract_agentic LlamaExtract Agentic
llamaextract_cost_effective LlamaExtract Cost-Effective
reducto_deep_extract Reducto Deep Extract
extend_extract_max Extend (Max Context)
datalab_parse_accurate_extract_balanced Datalab (Accurate + Balanced)
codex_code_extract_gpt_5_5_low Codex (GPT-5.5)
claude_code_extract_opus_4_8 Claude Code (Opus 4.8)
qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file Qwen3.6 35B (self-hosted)
gemma4_26b_vllm_extract_oneshot_structured_output_file Gemma4 26B (self-hosted)
nuextract3_extract NuExtract3 (self-hosted)
lift_extract Lift Datalab 9B (self-hosted)
gemini_3_5_flash_extract_oneshot_structured_output_file Google Gemini 3.5 Flash
openai_gpt_5_4_nano_extract_oneshot_structured_output_file OpenAI GPT-5.4 Nano

The four self-hosted pipelines need an endpoint you run yourself; see .env.example. Other configurations of the same systems are registered too (reducto_extract, extend_extract, codex_code_extract_gpt_5_5_high, the two-stage parse baselines, and more). Run uv run extract-bench pipelines for the full roster.

All three LlamaExtract tiers return word-level citation boxes, so both grounding metrics are meaningful on each. llamaextract_agentic_plus does it natively. llamaextract_cost_effective and llamaextract_agentic get there by running a parse at their own tier that emits word boxes, which is a second job. llamaextract_cost_effective_standard_bbox and llamaextract_agentic_standard_bbox are those two tiers without that parse pass: one job instead of two, but citations carry only block-level boxes, so word-level grounding scores near zero.

The parse and layout-detection rosters inherited from ParseBench are still registered and runnable by name; list them with extract-bench pipelines --parse, --layout, or --all. They are hidden from the default listing because this benchmark scores extraction; the two-stage extract pipelines use them internally as their parse stage.

Dataset

Hosted on HuggingFace: llamaindex/ExtractBench

The benchmark is split by document length, with one JSONL row per (document, schema) test case plus the source PDFs:

Split File Documents Pages Length
Short short.jsonl 252 615 ≤10 pages
Medium medium.jsonl 98 2,438 11–50 pages
Long long.jsonl 20 1,816 >50 pages
Total 370 4,869

The benchmark spans 8 business domains and 67 document types: finance and fund holdings, energy-sector regulatory forms, government procurement and customs, auto valuation, supply chain, healthcare remittance, legal and bankruptcy filings, and real estate.

Tag axes, sources, and ground truth

What each task challenge tests:

  • T1: long-list completeness. Recover every record of a repeated structure that can span many pages. Typical failures are truncation, duplicated or merged rows, hallucinated records, and values attached to the wrong record.
  • T2: needle-in-haystack. Find a small number of requested facts in a long document. T2 has few target records but many plausible mentions, only one of which is canonical; failures are missed targets, wrong occurrences, and unnormalized paraphrases.
  • T3: dense documents. Fill many fields from a document dense with labels, blanks, checkboxes, handwriting, and scan artifacts. The characteristic failure is over-extraction, inventing a value for a field that is actually blank, compounded by missed checkboxes and mislabeled fields.

A document can carry more than one task challenge.

The other axes are tagged independently of the task challenge:

  • Table structure — S1 merged headers, S2 header not at top / pivoted, S3 cross-page table, S4 enormous table, S5 table within a cell.
  • Perception challenge — P1 rotated or image-only capture, P2 scanned page images, and P3 handwriting. 38 documents are degraded re-captures of documents that also appear clean, so capture degradation is a paired measurement on the same documents.
  • Business domain — D1 finance (145), D2 energy (98), D3 government (49), D4 automotive (27), D5 supply chain (20), D6 healthcare (15), D7 legal (10), D8 real estate (6).

Sources. All documents come from public records: SEC and regulatory filings, government procurement and customs forms, court and agency exhibits (including tax forms such as W-2, 1040, K-1, and 1099-B), Texas Railroad Commission energy filings, and published business documents. 325 are real; 45 are synthetic long lists rendered from real layouts. PDF metadata has been stripped from every file.

Ground truth uses a method matched to each source: adjudicated agreement across independent extraction systems for real documents, values fixed before rendering for synthetic long lists, and human-verified values and boxes for forms. Each field's ground truth is an evidence list — the expected value plus any alternate acceptable readings, each with its source location — and scoring accepts a match against any listed reading.

The dataset is automatically downloaded when you run a pipeline. To manage it manually:

# Download the full dataset
uv run extract-bench download

# Download a small test dataset (6 documents, good for trying things out)
uv run extract-bench download --test

# Check whether the dataset has been downloaded and show summary statistics
uv run extract-bench status

Metrics

ExtractBench reports one metric for value accuracy and two for grounding. We score the grounding metrics only on fields that carry verified box ground truth.

  • Unified value F1 — whether the extracted values match the expected output, under one definition for scalar fields and arrays of records. Each output is flattened into cells, one per scalar field and per aligned record subfield, and a cell is correct when it matches its expected counterpart after normalization. Precision, recall, and F1 are computed over these cells per document; slices report unweighted document means.
  • Word-level grounding F1 — a field is grounded-correct only when its value is accepted and its predicted box overlaps any accepted box on its evidence list, at IoU 0.5.
  • Page-level grounding F1 — the same rule against the cited source page instead of the box: a field counts only when its value is accepted and the page it cites is correct.
How the unified value F1 is computed
  • Array alignment. A repeated structure compares as an unordered set of records: records pair by a globally optimal one-to-one assignment (the Hungarian algorithm) that minimizes mismatched cells. Unmatched expected records lower recall, surplus predictions lower precision.
  • Normalization. Matching is deterministic and mostly exact: dates in eight written formats canonicalize to ISO form, strings compare exactly after whitespace collapsing, and everything else uses plain equality, with no numeric tolerance and no LLM judge.
  • Missing values. An omitted key scores as an explicit null, every scalar field enters both denominators, and a correct null on a blank field is credited. Only repeated records move precision and recall apart, so a gap between them means records were dropped or invented.

Failed and missing documents score zero rather than being dropped, so a pipeline cannot raise its average by erroring out on the documents it finds hardest.

Usage

Running the Benchmark

The run command runs inference, evaluates against ground truth, and generates reports:

# Evaluate an extraction system on the whole benchmark
uv run extract-bench run <pipeline_name>

# Evaluate a single split only (short, medium, long)
uv run extract-bench run <pipeline_name> --group short

# Skip calling the extraction system — just re-evaluate existing results
uv run extract-bench run <pipeline_name> --skip_inference

# Control how many documents are processed in parallel
uv run extract-bench run <pipeline_name> --max_concurrent 10

# Run on the small test dataset only
uv run extract-bench run <pipeline_name> --test

Viewing & Comparing Results

# View reports in your browser (needed because browsers block PDF rendering from file:// URLs)
uv run extract-bench serve <pipeline_name>

# Compare two extraction systems side-by-side
uv run extract-bench compare <pipeline_a> <pipeline_b>

# Generate a leaderboard across all evaluated systems
uv run extract-bench leaderboard

# Leaderboard for specific systems only
uv run extract-bench leaderboard llamaextract_agentic llamaextract_cost_effective
Advanced Subcommands

For fine-grained control over individual steps:

# Run inference only (call the extraction system, don't evaluate)
uv run extract-bench inference run <pipeline_name> data/short --output_dir output

# Run evaluation only (on existing inference results)
uv run extract-bench evaluation run output/<pipeline_name> --test_cases_dir data

# Generate detailed HTML report from evaluation results
uv run extract-bench analysis generate_report --evaluation_dir ./output/<pipeline_name>
Evaluating Your Own System

To add a new extraction system, use Claude Code:

/integrate-pipeline <name> <API docs or SDK link>

This creates the provider, registers the pipeline, and updates docs. The skill definition lives in .claude/commands/integrate-pipeline.md and can be adapted for other AI coding agents.

Configuration

API Keys

Each pipeline calls a specific system's API. You only need the key for the system you want to evaluate. Add it to a .env file at the project root (see .env.example for the full list):

# Only add the keys you need. For example, to evaluate LlamaExtract:
LLAMA_CLOUD_API_KEY=...

# To evaluate OpenAI-based pipelines:
OPENAI_API_KEY=...

# To evaluate Anthropic-based pipelines (including claude_code_extract_*):
ANTHROPIC_API_KEY=...

# To evaluate Google-based pipelines:
GOOGLE_API_KEY=...

# Codex coding-agent pipelines authenticate the codex CLI:
CODEX_API_KEY=...

ExtractBench does not use LLM-as-a-judge; all value scoring is deterministic. Your keys only ever call the extraction system you are evaluating.

CLI Reference

Command Description
extract-bench run Evaluate an extraction system end-to-end (inference + evaluation + reports)
extract-bench download Download the benchmark dataset from HuggingFace
extract-bench status Check whether the dataset has been downloaded
extract-bench pipelines List extraction pipelines (--parse, --layout, --all for the rest)
extract-bench compare Compare results from two systems side-by-side
extract-bench leaderboard Generate a leaderboard across all evaluated systems
extract-bench serve View HTML reports in your browser (with PDF rendering support)

Advanced subcommands: inference, evaluation, analysis, pipeline, data

Output Structure
output/
├── _leaderboard.html                       # Cross-pipeline leaderboard
└── <pipeline_name>/
    ├── short/
    │   ├── *.result.json                    # Inference results
    │   ├── _evaluation_report.json          # Evaluation summary
    │   ├── _evaluation_report_detailed.html # Interactive detailed report
    │   ├── _evaluation_results.csv          # Per-example CSV
    │   └── _evaluation_report.md            # Markdown summary
    ├── medium/  (same structure)
    ├── long/    (same structure)
    ├── _errors.json                         # Per-document inference failures
    └── _metadata.json                       # Run metadata
Project Structure
src/extract_bench/
├── cli.py                           # Fire CLI entry point
├── pipeline/cli.py                  # End-to-end pipeline orchestration
├── data/
│   ├── download.py                  # HuggingFace dataset download
│   └── cli.py                       # Data management CLI
├── inference/
│   ├── runner.py                    # Batch inference with concurrency
│   ├── pipelines/                   # Pipeline registry (extract, parse, layout)
│   └── providers/                   # Provider implementations per product type
├── evaluation/
│   ├── runner.py                    # Parallel evaluation + failure penalties
│   ├── evaluators/                  # Product-specific evaluators
│   ├── metrics/extract/             # Unified value F1, grounding, record matching
│   └── reports/                     # CSV, HTML, markdown export
├── analysis/
│   ├── detailed_report.py           # Interactive per-split HTML report
│   └── comparison.py                # Pipeline comparison
├── test_cases/
│   ├── loader.py                    # Load test cases (JSONL or sidecar .test.json)
│   └── schema.py                    # TestCase types (Extract, Parse, LayoutDetection)
└── schemas/
    ├── pipeline_io.py               # InferenceRequest, InferenceResult
    ├── evaluation.py                # EvaluationResult, EvaluationSummary
    └── product.py                   # ProductType enum

Citation

@misc{zhang2026extractbenchbenchmarkschemaguidedenterprise,
  title={ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction},
  author={Boyang Zhang and Adrian Lyjak and Eli Stewart and Zhaoqi Li and Simon Suo},
  year={2026},
  eprint={2607.29677},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2607.29677},
}

Metadata

Release files for llama-extract-bench 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-extract-bench 1.0.0
File Size Uploaded
llama_extract_bench-1.0.0.tar.gz 1.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-extract-bench 1.0.0
File Interpreter ABI Platform
llama_extract_bench-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.4 MB

Release files / llama_extract_bench-1.0.0.tar.gz

Download URL llama_extract_bench-1.0.0.tar.gz
Size 1.2 MB
Tags Source
SHA-256 checksum
How to use checksums
2e70fe94875950e85134d19ee75747904ec824e3adce9f40e7922619c27fbda9
BLAKE2b-256 checksum
How to use checksums
05fe5b6f3188e6b6298fdcf7874df5a24fad887203a6c48f51d98f53a0e8f26b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / llama_extract_bench-1.0.0-py3-none-any.whl

Download URL llama_extract_bench-1.0.0-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
235d838423d773190228edffef9003d7b6d08c3100c379945527cd0a5d1eec62
BLAKE2b-256 checksum
How to use checksums
c95b9e906813e0d624acddc56e69fef7a32e07f6da0c5e2fe67e2ca7dd45b0f4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page