Skip to main content

groundgate

LLMs propose facts. groundgate decides which ones you can trust.

groundgate is a deterministic admission layer for LLM document extraction. You give it a document, a schema, and the facts an extractor proposed. Every fact comes back with one of three outcomes:

  • admitted: every check passed;
  • needs_verification: nothing failed, but a person should look, and the decision says where;
  • rejected: a check failed, with a stable reason code.

Every run writes a receipt with hashes of every input. Anyone can re-derive it byte for byte.

The groundgate review page for IRS Publication 590-A. Left: a Roth IRA limit flagged QUALIFIED_VALUE because the model cited the spousal-deduction sentence ("more than $236,000"). Right: the source text, with that citation outlined in amber and the correct Roth sentence, admitted, highlighted in green below it.

Why this exists

Current models get most extracted values right. In the benchmark below, 4,233 of 4,350 extractions were correct. The other 117, about 1 in 37, each cited a real place in the document, so nothing on the page tells you which ones are wrong.

Grounded extraction tools check that the quoted text exists in the source. They don't check the value you store. LangExtract is the clearest example, and a good tool. It aligns each extraction_text to the source and reports MATCH_EXACT, MATCH_FUZZY and so on. It never looks at the attributes that hold the typed value.

A real case from the benchmark: IRS Publication 560 states two limits a few lines apart.

The limit on elective deferrals, other than catch-up contributions, is $23,500 for 2025 ...
The limit on salary reduction contributions, other than catch-up contributions, is $16,500 for 2025 ...

The first is the 401(k) limit. The second is the SIMPLE plan limit. Asked for the 2025 elective deferral limit, 4 of 14 runs proposed both, and LangExtract marked both quotes MATCH_EXACT. groundgate sent both to a person with CONFLICTING_CANDIDATES, because two proposals for one field disagree. In Publication 15-B, one model stored "$1 million" as 1, also an exact match, and groundgate flagged it SCALE_WORD.

The same gap shows up on the document I used while building the library. I ran Gemini 3.6 Flash and GPT-OSS 120B through LangExtract on two pages of IRS Publication 590-A. LangExtract aligned all 48 extractions as MATCH_EXACT. groundgate then:

  • rejected Gemini's reading of a misprinted $252,0000 as 2,520,000;
  • flagged two GPT-OSS facts that took a Roth IRA limit from the spousal-deduction sentence; the numbers match, but the text says "more than" where the field says "at least";
  • listed a required field both models missed;
  • admitted 44 facts, all with the right value.

One admitted fact still cites the wrong rule. Its number is correct, its comparator matches, and no span check can tell. I checked every decision by hand; the details are in examples/irs-590a, which is also the page in the screenshot.

I measured the gap on purpose before writing the library. On 51 facts from two public-domain government documents, LangExtract's strictest setting (MATCH_EXACT only) accepted 100% of planted value and unit errors whose quote was correct. groundgate caught all of them and rejected none of the correct facts. The method and full tables are in spikes/m0_5.

LangExtract locates the text. groundgate checks the number in it.

Benchmark

Seven models (Gemini, GPT and Claude, each through its own CLI) ran through LangExtract on 30 public-domain FDA drug labels, NTSB accident reports and IRS publications, at two chunk sizes. A person checked every gold fact. These are spec 0.1's numbers on the first set, as published. Of 4,350 extractions, 117 were wrong: the wrong value, a value for a field the document doesn't state, the wrong unit, or not a number.

Pooled over all 14 runs: MATCH_EXACT accepted 98% of wrong extractions without review, groundgate 6.8%. Of correct extractions citing the right place, MATCH_EXACT rejected 4.6%, groundgate 0.3%. groundgate sent 12% of extractions to a person.

LangExtract, MATCH_EXACT groundgate
wrong extractions accepted without review 98.3% (115/117) 6.8% (8/117)
correct extractions citing the right place, rejected 4.6% (194/4,233) 0.3% (14/4,233)
extractions sent to a person 0% 12.1% (526/4,350)

Of the 526 extractions groundgate sent to a person, 92 were wrong; a random sample that size would have held about 14. groundgate also rejected 71 correct values that LangExtract had aligned to a place that doesn't state them: the value was right, the evidence wasn't.

Most catches (84 of 109) were CONFLICTING_CANDIDATES: a model proposed two values for a single-valued field, and groundgate sent both to review instead of picking one.

Six of the 8 that got through are the case groundgate says it cannot catch (see What it does not do). Five models took metoprolol's heart failure maximum (200 mg) as the hypertension maximum, and one took lisinopril's renal impairment dose as the usual starting dose. Each number is real and cited correctly; it belongs to another condition. The other two read levothyroxine's 1.6 mcg/kg/day as a flat 1.6 mcg, which the unit check allows. Putting all seven models through one gate caught one more, because when they were wrong they mostly agreed. Spec 0.2 goes after both: keyed fields check the condition, and a dose per kilogram no longer passes as a flat dose.

The benchmark also found a bug. When LangExtract's chunker cut "29.97" after "29.", groundgate read the span as 29, which its own spec forbids. It's fixed, and three wrong altimeter settings no longer get through.

Tables and planted-error results: bench/RESULTS.md. Method and caveats: bench/README.md. The write-up, with what got through and why: Grounding is a contract, not a citation.

Spec 0.2 on a second set

Spec 0.2's rules were written from the escapes above, so they can't be measured on the same documents. I picked 101 new ones by rules frozen before anyone read them: FDA labels with doses per indication or per kilogram, IRS publications and Federal Register rules with "increased from", "must not be more than" and "$1.2 million". The same seven models ran at both chunk sizes. Every candidate was decided twice, by spec 0.1 (the released 0.1.0 wheel) and by spec 0.2. Pooled over 14 runs on the 74 documents checked in full:

LangExtract, MATCH_EXACT groundgate, spec 0.1 groundgate, spec 0.2
wrong extractions accepted without review 85.7% (705/823) 30.7% (253/823) 8.8% (72/823)
correct extractions citing the right place, rejected 3.8% (230/6,080) 11.3% (690/6,080) 0.5% (28/6,080)
extractions sent to a person 0% 28.2% (1,950/6,903) 39.6% (2,730/6,903)

642 of 0.1's 690 false rejects are in the 19 documents picked for amounts like "$1.2 million", which 0.1 could not match. 82 of its 253 escapes were a right value under the wrong indication; 0.2 lets 10 of those through. It pays in review load. The 72 that still got through are listed one by one in bench/set2/RESULTS.md, and two rules that need work are open for the next version (#51, #52).

Quickstart

pip install groundgate
import groundgate as gg

text = "The IRA contribution limit is $7,000. The deduction phases out above $79,000."
schema = {
    "fields": {
        "ira_limit": {"type": "integer", "unit": "USD"},
        "phaseout_start": {"type": "integer", "unit": "USD"},
    }
}


def cite(field, value, quote):
    """A candidate citing the first place `quote` appears, as UTF-8 byte offsets."""
    start = text.encode().index(quote.encode())
    span = {"start": start, "end": start + len(quote.encode()), "text": quote}
    return {"field": field, "value": value, "unit": "USD", "evidence": span}


candidates = [
    cite("ira_limit", "7000", "$7,000"),
    cite("ira_limit", "70000", "$7,000"),  # a digit too many
    cite("phaseout_start", "79000", "$79,000"),  # "above" is not "at"
]
receipt = gg.admit(text, schema, candidates)
for d in receipt.decisions:
    print(f"{d.outcome:<19} {d.field:<15} {d.value:<6} {' '.join(d.codes)}".rstrip())

print(gg.verify(receipt.to_dict(), text, schema, candidates).ok)
admitted            ira_limit       7000
rejected            ira_limit       70000  VALUE_NOT_IN_EVIDENCE
needs_verification  phaseout_start  79000  QUALIFIED_VALUE
True

The core has no dependencies. groundgate[pdf] adds pdfminer.six for PDF text, and groundgate[langextract] installs LangExtract. docs/guide.md covers schemas, policies, candidates and the Python API.

With LangExtract

import langextract as lx
from groundgate.adapters.langextract import admit_document

result = lx.extract(text_or_documents=text, prompt_description=prompt, examples=examples)
receipt = admit_document(result, schema)

Each extraction's value and unit attributes are checked at the place LangExtract aligned its extraction_text. An extraction LangExtract could not align is rejected NO_EVIDENCE.

From the command line

groundgate extract p590a.pdf --pages 1-2 -o doc.txt --layout layout.json
groundgate admit   doc.txt schema.json candidates.json -o receipt.json
groundgate verify  receipt.json doc.txt schema.json candidates.json
groundgate report  receipt.json doc.txt schema.json candidates.json --layout layout.json -o report.html

extract turns a PDF, HTML, XML or text file into the NFC text everything else reads, and records the page and box of every word. Text a reader can't see (proof marks drawn off the page, hidden HTML) is dropped, and superscripts are marked so that 10⁹ reads 10^9, never 109. Pass - as the document to read it from a pipe. report refuses to render a receipt that doesn't re-derive from its inputs. examples/irs-590a runs the whole pipeline on an IRS publication with cached model output, so it reproduces without an API key.

What it checks

Checks run in a fixed order. The first failure rejects the fact.

Code Rejects when
CANDIDATE_INVALID the candidate is malformed
FIELD_UNKNOWN the field is not in the schema
NULL_STRING_LITERAL the value is "null", "none", "n/a"
TYPE_INVALID the value does not parse as the field's type
RANGE_INVALID the value is outside the field's bounds
UNIT_INVALID the unit is not the field's unit
KEY_INVALID the field is keyed by condition, and the key is missing or not one of its keys
NO_EVIDENCE no evidence is cited
SPAN_INVALID the span is outside the document or splits a character
VALUE_NOT_IN_EVIDENCE no number in the span equals the value
UNIT_NOT_IN_EVIDENCE the value is there, but not with the field's unit

A fact that passes every check can still be flagged for a person:

Code Flags when
NON_VERBATIM_EVIDENCE the quote differs from the text at the span
QUALIFIED_VALUE "up to", "approximately", "or more" changes the value, and the schema didn't declare it
SCALE_WORD "million", "lakh" and similar follow the value, and the value was given unscaled
KEY_NOT_AT_VALUE the text puts the value under another condition, such as another indication's heading
LOW_CONFIDENCE the extractor's confidence is below the policy minimum
CONFLICTING_CANDIDATES another proposal for the same field and key has a different value

Number matching is collision-safe. The span 500 mg inside 1,500 mg never reads as 500, a span that stops at 29. inside 29.97 never reads as 29, and a malformed number like the $252,0000 printed in IRS Publication 590-A never equals 252,000 or 2,520,000. When a span misses the value but its quote occurs exactly once elsewhere with the right value and unit, groundgate moves the evidence there and records EVIDENCE_REANCHORED.

The rules are in SPEC.md. conformance/ holds 31 language-neutral vectors that pin every code, so another implementation can prove it agrees.

Receipts

{"groundgate": "0.3",
 "document": {"id": "irs-p590a-2025-pages-1-2", "sha256": "sha256:da7797d569a5..."},
 "schema_sha256": "sha256:...", "policy_sha256": "sha256:...",
 "decisions": [{"candidate_id": "Gemini_3.6_Flash_Medium/0", "field": "ira_limit_2025",
                "outcome": "admitted", "codes": [], "value": "7000", "unit": "USD",
                "evidence": {"start": 2317, "end": 2323}, "candidate_sha256": "sha256:aaf08a9a..."},
               "..."],
 "coverage": [{"field": "roth_phaseout_single_2025_end", "code": "REQUIRED_FIELD_MISSING"}],
 "summary": {"admitted": 44, "needs_verification": 3, "rejected": 1},
 "receipt_sha256": "sha256:..."}

Hashes are SHA-256 over RFC 8785 canonical JSON, so a receipt means the same thing in any language. groundgate verify re-runs every decision from the inputs and fails on any difference: a changed document, schema, policy, candidate or outcome.

What it does not do

  • It does not judge meaning. If the model reports a number that really is in the sentence but belongs to another field, every span check passes: a model that reads an age-50 limit of $8,000 as the IRA limit cites a real "$8,000" with the right unit. A second proposer plus CONFLICTING_CANDIDATES catches it only when the models disagree, and in the benchmark they mostly agreed: six of the eight escapes were this case. When the field depends on a condition the schema can name, such as a drug's indications, a keyed field checks it (spec 0.2).
  • No dates, arrays of records, or cross-document checks in spec 0.3.
  • English number formats only. A decimal comma (1.234,56), space groups (1 234) and number words (five) have no value, so a fact that cites them is rejected. The guide lists what is read. Locale packs for other formats are a design (docs/design/locale-packs.md, #60).
  • No OCR. Scanned PDFs need a text layer first (for example ocrmypdf).
  • No model calls in the decision. groundgate never asks an LLM whether an LLM was right. A planned optional tier (#66) records a small model's answers as inputs, so a receipt still re-derives byte for byte.

Status

Alpha. groundgate 0.3.0 implements spec v0.3. CHANGELOG.md lists what 0.3 changes, each with its measure: a wider change rule (#51), abbreviation dots that no longer end a sentence for qualifiers (#63, #77), Indian digit grouping (#59), ASCII digits (#94), and keys that stop at a label line (#52) or read a table by its lines (#87). A receipt names the spec version it was decided under, and only that version verifies it.

Two designs are not yet code: locale packs for number formats outside English, and hybrid decisions, where a recorded model answer can send a fact to review or clear a measured kind of flag.

License

Apache-2.0.

Metadata

Release files for groundgate 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for groundgate 0.3.0
File Size Uploaded
groundgate-0.3.0.tar.gz 99.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for groundgate 0.3.0
File Interpreter ABI Platform
groundgate-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 148.8 kB

Release files / groundgate-0.3.0.tar.gz

Download URL groundgate-0.3.0.tar.gz
Size 99.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a9eb159f107b121d2eff63bcd825a02a34ef58576c74eabce7d717cc6afdbb11
BLAKE2b-256 checksum
How to use checksums
2f922b5a2e7e2bd32fb07436137ba69d7e9b39a4431fd9472b01b663d79e7491
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / groundgate-0.3.0-py3-none-any.whl

Download URL groundgate-0.3.0-py3-none-any.whl
Size 49.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f3b9b65090f40610374dc1c4c96ba12316af81a87b89f11510226e3c71dc3a1b
BLAKE2b-256 checksum
How to use checksums
f85572ca4e85705c63748029016909ff3217e9fa49ce4c9e36ef41e675c49a08
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.0

2 release files

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page