Skip to main content

GroundLens

Execution verification runtime for AI systems and agents


License

PyPI Docs Rust Python OpenSSF Best Practices OpenSSF Scorecard REUSE status SLSA


What it is · Quick start · Architecture · How it works · Engine · Runtime · Records · MCP server · Determinism · Examples · Docs · FAQ · Roadmap


What GroundLens is

GroundLens is an execution verification runtime for AI systems and agents. It turns observable AI execution into deterministic, policy-governed evidence that can be independently verified.

GroundLens provides a vendor-neutral runtime and evidence protocol for observing AI executions, evaluating claims, tool calls, actions and outcomes against composable verifiers and policies, and producing signed, reproducible evidence records.

The unit is the execution: an ordered sequence of steps an AI system or agent takes, from a model call and a retrieval to a tool call, an action with side effects and a human approval. GroundLens records each step, checks it, decides, and seals the run into a signed record anyone can verify offline. Verifying a single answer is the smallest case, a run with one claim.

  • an answer, and the claims inside it → PASS, REVIEW or FAIL
  • a tool call or an action → ALLOW, REVIEW or DENY

Where a guardrail blocks or scores an output in the moment and leaves nothing behind, GroundLens leaves signed, hash-chained evidence a third party can check without trusting you. It runs locally, needs no access to your weights, prompts or architecture, and is built for teams shipping AI answers and agents into regulated or high-stakes workflows who need proof, not a score.


Quick start

pip install groundlens

The package installs the engine, the groundlens and glv commands, and needs no other dependency. Numbers and rules are checked out of the box; the lexical verifier needs one optional download, shown at the end.

Verify an answer

Check a model's answer against the sources it was given, under a policy, and get a record you can keep.

from groundlens import verify

question = "What is the invoice total?"
source   = "The total amount due is 10,000 dollars, payable within 30 days of receipt."
answer   = "The invoice total is 1,000 dollars, due in 30 days."

record = verify(answer, [("invoice.pdf#p1", source)], question=question)

print(record.decision)     # 'FAIL'
print(record.report())
FAIL  policy=groundlens_default_v1  record=rec_350455f44e60_4dbfea8eb79c
  c2   groundlens.numeric   contradicted   0.00   nearest in invoice.pdf#p1: '10,000 dollars'

Ten thousand is not one thousand. A similarity score would rate the right answer and the wrong one alike; the numeric verifier compares the quantities exactly and points at the source number the answer lost to. The 30 days are supported in both, so they do not appear in the report: it shows only what a reviewer needs to look at.

Every verification is sealed:

record.content_hash            # 'sha256:…' — same input, policy and bundle → same hash, any machine
record.verify()                # recompute every hash and the Ed25519 signature, offline; raises if altered
record.regulatory_mapping      # the articles this decision concerns, under the policy

Verify a run

Give GroundLens an MCP execution trace and an execution policy. It records the run as a hash-linked event log, gates it, and seals a signed run record.

from groundlens import verify_run

record = verify_run(
    "examples/run/trace.jsonl",            # an MCP session, as JSON-RPC lines
    "examples/run/execution-policy.yaml",  # the rules for what the agent may do
    run_id="run_demo",
    system="invoice-agent",
)

print(record.gate)         # 'DENY'  — the run called shell.exec, which the policy forbids
print(record.breaches)     # ()      — nothing *ran* against the policy; the call was denied, not executed
print(record.record_hash)  # 'sha256:…'  — signed and chained, like an answer record

gate is the verdict over the whole run: ALLOW, REVIEW or DENY, rolled up from the strictest step. breaches is different and narrower: it lists actions that actually executed against the policy, an action the policy forbade or one that needed a human approval that never came. Here the forbidden tool was stopped, so the run is DENY with no breach. The same thing on the command line, with the real output:

glv run verify --trace examples/run/trace.jsonl --policy examples/run/execution-policy.yaml \
  --run-id run_demo --system invoice-agent --log runs.jsonl
# exit code 1  (0 ALLOW · 3 REVIEW · 1 DENY)

glv run check runs.jsonl
# ok  1 run records, chain intact, all signatures verify

A runnable version of both is under examples/run.

Enable the lexical verifier

The numeric and rules verifiers need nothing. The lexical verifier (word-level grounding against the sources) needs the base bundle: the multilingual encoder, its tokenizer and a manifest of hashes.

groundlens bundle pull base      # ≈470 MB, once; the only command that uses the network

With the bundle installed, the lexical verifier runs and its evidence appears in the record, one row per content word, each anchored to the source word it was scored against.

from groundlens import verify

record = verify(
    "El importe de la factura es de 10.000 euros, pagaderos en 30 días.",
    [("factura", "El importe total asciende a 10.000 euros, pagaderos en un plazo de 30 días.")],
    locale="es",
)

for e in record.evidence:
    if e.verifier_id == "groundlens.lexical":
        print(e.result, round(e.score, 2), e.source_text)
# for example — scores are a 0–1 contextual support from the frozen encoder
supported  0.93  importe
supported  0.90  pagaderos
supported  0.88  factura

The score is a contextual similarity, so the same word used differently scores lower, and the weakest anchor is what a reviewer reads first. The download is checked against a hash pinned in the engine and refuses anything else; in an isolated environment, copy the bundle directory by hand and point GROUNDLENS_BUNDLE_DIR at it.

The command line

Everything except the lexical verifier works with the base install alone.

groundlens verify --answer answer.txt --question question.txt \
  --source "invoice.pdf#p1=invoice.txt" --policy eu_ai_act_high_risk_v1 --log records.jsonl
groundlens record verify records.jsonl        # every hash, every link, every signature
groundlens report records.jsonl --out report  # report.md, report.json, README-auditor.md
groundlens policy lint policies/eu_ai_act_high_risk_v1.yaml
groundlens bundle status                      # is the base bundle installed, where, which hash

Exit codes: 0 PASS, 1 FAIL, 2 error, 3 REVIEW. The Rust binary glv exposes the same commands and adds execution verification: glv run verify seals an agent run (exit 0 / 3 / 1 on ALLOW / REVIEW / DENY) and glv run check verifies a log of run records offline.

Full documentation, including the API reference and concept guides, is at groundlens.readthedocs.io.


Architecture

Verifying a single answer is the smallest case, a run with one claim, so one contract covers both ends of the range:

  • an answer, and the claims inside it, gets PASS, REVIEW or FAIL from verifiers and a policy;
  • a tool call or an action gets ALLOW, REVIEW or DENY from an execution policy.

Either way the run is sealed into a signed, chained record. What the record keeps of the world is hashes, not content, so it is safe to hold in a regulated place while staying independently verifiable.

GroundLens sits beside your AI system, not inside it. It observes what the system produces and does, and never sees your weights, your prompts or your internal architecture, so independent verification is possible even in a bank or a sensitive deployment.

It reads a run from what an agent already emits. An agent driving its tools speaks the Model Context Protocol (MCP); GroundLens ingests those JSON-RPC messages and turns them into a run, recording hashes of the arguments and results, never the content itself. Recording a run needs no change to how the agent is built.

The engine and runtime are a Rust workspace, wrapped for Python, with no runtime dependencies; glv is the same code as a binary. No engine or runtime crate depends on an HTTP or TLS library, and a CI job fails the build if one ever appears. The only network operation in the project is one explicit command, bundle pull, which fetches the optional lexical model. Verification never reaches the network.

For the full design, the crate-by-crate layout, the core contracts (verifier, evidence, claim, policy, record, run) and the data flow, see ARCHITECTURE.md.


How it works

How a verification works

A verifier produces evidence, not truth: it reports what it measured and how sure it is, and none of them decides. A policy interprets the evidence and reaches the decision. The whole chain becomes a record: the input hashes, the verifiers and model hashes that ran, the evidence, the policy and its hash, the decision, the regulatory mapping, and the hash of the previous record, sealed with an Ed25519 signature. A log of records is an audit trail you can hand over as a file.


Engine

The engine verifies an answer and the claims inside it. A verifier produces evidence; a policy turns it into PASS, REVIEW or FAIL. What an agent did, its tool calls and actions, is decided by the execution policy in the Runtime.

verifier what it does example
groundlens.numeric numbers, currencies, percentages and physical units, compared exactly in base units answer 1,000 vs source 10,000 → contradicted; 1.2 km = 1200 m; 212 °F = 100 °C; $37.35 billion = a cell 37,350 under "in millions"
groundlens.rules your own symbolic rules, run as a verifier rule "an APR must be a percentage" → an APR written as a bare number is contradicted
groundlens.lexical whether each word of the answer is anchored in the sources, by contextual token similarity, reported as the weakest anchor word pagaderos anchored to the source and scored 0.90; a word with no support scores low and surfaces first
groundlens.nli whether a source entails, contradicts or is neutral to each statement in the answer statement "the fee is 0.75%" against a source saying 0.50% → contradiction at high confidence

groundlens.numeric and groundlens.rules are exact (bit-identical on any machine) and need no download. groundlens.lexical and groundlens.nli are reproducible (a pinned model, scores within a declared tolerance across machines) and run from a bundle: lexical from the base bundle, nli whenever the loaded bundle carries an entailment model. The base bundle ships the entailment model from v2 (see the roadmap). Semantic, geometric (SGI, DGI) and LLM-judge verifiers are planned; every one plugs into the same contract.

Locales matter for numbers: 1.234 is one thousand in Spanish and one and a bit in English. GroundLens reads en, es, ca, de, fr, it, pt, nl and Swiss formats, knows short and long scale words, and keeps every legitimate reading of an ambiguous numeral instead of guessing. The base bundle's encoder covers about a hundred languages.

A policy is a short YAML file you control. Two policies over the same evidence can reach different decisions, and both are correct: that is where your risk appetite lives, not in the engine. The bundled eu_ai_act_high_risk_v1 maps outcomes to Art. 15(1) (accuracy and robustness), Art. 14(4)(a) (human oversight) and Art. 12(1) (record keeping) of Regulation (EU) 2024/1689; every policy has a version and a hash, and the hash goes into every record it decides. Scores from statistical verifiers drift slightly between machines, so each threshold carries a guard band, and a score inside it is REVIEW everywhere.

record = verify(answer, sources, policy="eu_ai_act_high_risk_v1")
record.decision              # 'FAIL'
record.regulatory_mapping    # [{'article': 'Art. 15(1)', ...}, {'article': 'Art. 12(1)', ...}]

Runtime

The runtime verifies an execution. It records each step of a run as an event in a hash-linked log, a model call, a retrieval, a tool request and its result, an action, a human approval, and an execution policy decides what the agent may do.

An execution policy is a short, ordered list of rules. Each rule matches a tool call or an action and carries an effect: DENY stops the step, REVIEW holds it for a human, ALLOW lets it proceed. The first rule that matches decides; when none does, the default applies, so a conservative deployment denies anything it did not explicitly allow. The gate is pure rule matching, with the same exact guarantee as the numeric verifier.

id: eu_high_risk_v1
rules:
  - id: no-shell          # a shell tool is never allowed, from any server
    match: tool
    name: shell.exec
    effect: DENY
  - id: high-risk         # any action at or above high risk needs a human
    match: risk_at_least
    risk: high
    effect: REVIEW
default: ALLOW

After a run, GroundLens audits the whole log against the policy, rolls it up to a single verdict, and flags any action that ran against it: one the policy forbade, or one that needed a human approval that never came. The verdict and the breaches go into the signed record, so an auditor can replay a run and see whether the policy was honoured.


Evidence records

Whether GroundLens checked one answer or a whole run, the result is the same kind of artefact: a signed record, chained to the one before it, that anyone can verify offline.

record.content_hash     # same input, policy and bundle → same hash, on any machine
record.verify()         # recompute every hash and the Ed25519 signature, offline
Record.verify_chain(Record.read_log("records.jsonl"))

Change one byte anywhere in a record and verification fails. Append records to a JSON Lines log and each one carries the hash of the previous one. groundlens report turns a log into a human-readable report with a one-page guide for auditors.


MCP server

The same verification is available as an MCP server, so an agent (or any Model Context Protocol client, including Claude) can call GroundLens as a tool: check an answer, gate an execution, or verify a log of records. It is a thin layer over the engine and runs over stdio.

It is an optional extra, so the base package keeps its zero dependencies:

pip install "groundlens[mcp]"
groundlens-mcp                     # runs the server over stdio

Three tools:

tool what it does
verify_answer verify an answer against its sources under a policy; returns the decision, the evidence and the signed record
verify_run gate an MCP execution trace under an execution policy; returns ALLOW / REVIEW / DENY, any breaches and the run record
verify_records verify a log of records offline: every hash, every link, every signature

Point an MCP client at the groundlens-mcp command. See the docs for a client configuration example.


Determinism

GroundLens is deterministic where it can be, and reproducible where it cannot.

exact verifiers and the execution gate use no floating point: the same input gives the same result, bit for bit, on any machine. reproducible verifiers run a pinned model in f32 on a pure-Rust inference engine, and their scores stay within a declared tolerance across machines. Anything non_deterministic, such as an LLM judge, is recorded with its model, prompt hash and settings, and decides only if the policy allows it.

This is tested, not asserted: CI runs the invoice example, with and without the lexical channel, on Linux, macOS and Windows under a Turkish locale and a Pacific timezone, and compares the record hash with a committed value.


Examples

Two notebooks under examples/notebooks run in Google Colab:

  • Verify an AI answer against its sources: one example in English, German, French, Spanish and Italian, from pip install to a signed record, with a wrong number, a paraphrase and a policy change. Open In Colab
  • Evidence records for auditors: a log of verifications, chain verification, tamper detection, the EU AI Act mapping and the report an auditor receives. Open In Colab

And a shell example of a whole agent run under examples/run: a trace, an execution policy and a signed run record.


Contributions are welcome; see CONTRIBUTING.md and SECURITY.md.

groundlens.dev · Javier Marín, 2026 (javier@groundlens.dev)

Release files for groundlens 5.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for groundlens 5.2.0
File Size Uploaded
groundlens-5.2.0.tar.gz 208.4 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for groundlens 5.2.0
File
groundlens-5.2.0-cp310-abi3-win_amd64.whl CPython 3.10 abi3 Windows x86-64 Details
groundlens-5.2.0-cp310-abi3-manylinux_2_28_x86_64.whl CPython 3.10 abi3 Linux glibc 2.28+ x86-64 Details
groundlens-5.2.0-cp310-abi3-manylinux_2_28_aarch64.whl CPython 3.10 abi3 Linux glibc 2.28+ ARM64 Details
groundlens-5.2.0-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details

Total release size: 34.2 MB

Release files / groundlens-5.2.0.tar.gz

Download URL groundlens-5.2.0.tar.gz
Size 208.4 kB
Tags Source
SHA-256 checksum
How to use checksums
07ac5cc32425d8102afd774e88df650749acb8f0d23ba023eb2bb017236f258f
BLAKE2b-256 checksum
How to use checksums
8288f80d3818bbff9890ed74606c2b35f879f8310c3fb4fc24b246b482f446a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / groundlens-5.2.0-cp310-abi3-win_amd64.whl

Download URL groundlens-5.2.0-cp310-abi3-win_amd64.whl
Size 8.8 MB
Tags CPython 3.10 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
618cc0251375ee2f82c3be10cfea1d02281409b2e1b9a585b76b2192d7a7ab9e
BLAKE2b-256 checksum
How to use checksums
d92f8116027c87b9f0c14819b2ecb553be4a49200c9aca51e5f78cc0d5149c58
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / groundlens-5.2.0-cp310-abi3-manylinux_2_28_x86_64.whl

Download URL groundlens-5.2.0-cp310-abi3-manylinux_2_28_x86_64.whl
Size 9.4 MB
Tags CPython 3.10 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
a1ca563ad693a0a15dbe9eee4fc4961541d5671d497cad26379e34d480ac3262
BLAKE2b-256 checksum
How to use checksums
5ccadc71e0ece6998c4d1e8f2b768594e4d2f34f3cbe8ea1fd8b59f2980381a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / groundlens-5.2.0-cp310-abi3-manylinux_2_28_aarch64.whl

Download URL groundlens-5.2.0-cp310-abi3-manylinux_2_28_aarch64.whl
Size 8.3 MB
Tags CPython 3.10 Linux glibc 2.28+ ARM64 abi3
SHA-256 checksum
How to use checksums
d7ae2112fb201b1f6616717e16c5aec1bad00bdbcbc4e7800f4a0f3fadd4395d
BLAKE2b-256 checksum
How to use checksums
db3b2768bb81572ea0b1ab7ef720a1c9543296610d0ce8d0b0918eab1e8ca406
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / groundlens-5.2.0-cp310-abi3-macosx_11_0_arm64.whl

Download URL groundlens-5.2.0-cp310-abi3-macosx_11_0_arm64.whl
Size 7.6 MB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
7bfb6cea0b8a39dbec00ebdccd81daaf9c4c87ba88ef01d50acd116a6fa9c237
BLAKE2b-256 checksum
How to use checksums
23ba687241f380af5fcde7da68eee3a939fc1091a484a2dd88a0514f765e30df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

5.3.0

5 release files

This release

5.2.0 This release

5 release files

5.1.0

5 release files

5.0.0

5 release files

4.0.0

5 release files

3.1.0

2 release files

3.0.5

2 release files

3.0.1

2 release files

3.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page