Skip to main content

AgentVerity

Conservative baseline admission for AI agents with bounded decisions.

PyPI Python 3.10+ CI Coverage: 90%+ License: Apache-2.0

AgentVerity qualifies tests for model-backed components that choose from a finite, reviewed set of decisions. Examples include routers, approval or policy gates, and supervisors that select the next agent or tool. It compares the named decision, exposed as verdict, rather than harmless changes in explanation text.

Correctness and trajectory evaluators ask whether the agent behaved properly. AgentVerity asks whether that evidence is repeatable and whether the reviewed decision contract was exercised before the run becomes a regression baseline. A regression baseline is a reviewed run saved as the expected behaviour for testing later versions.

Use another evaluator for open-ended chat with no reviewed decision or ordered tool-path contract. AgentVerity is alpha, with documented pre-1.0 guarantees.

Read the design story: Introducing AgentVerity: What Does a Green Agent Test Prove?

Is it a fit?

AgentVerity fits when all three conditions hold:

  • each run exposes a named decision or a reviewed tool or handoff path
  • repeated trials can start from equivalent state in isolated sessions
  • you can supply deliberately varied test inputs that should reach different valid decisions

Good targets include support and payment routers, fraud triage, approval and policy gates, incident routing, multi-agent supervisors, and bounded tool selectors. In a larger agent, test the step that owns the decision or the final pipeline decision. Test both when each is a release contract.

It is not an end-to-end quality evaluator for chat, RAG answers, generated content, coding agents, or research agents. If one of those systems also emits a reviewed route, approval, escalation, or tool path, AgentVerity can qualify that decision layer, not the open-ended work around it.

See the applicability checklist and exact limits.

Try it

pip install agentverity
from agentverity import from_callable, run

def route(ticket: str) -> dict[str, str]:
    # Deliberate defect: every ticket takes the same route.
    return {"text": "route: general", "verdict": "general"}

agent = from_callable(route)
result = run(agent, inputs=[
    "my card was charged twice",
    "the app crashes on login",
    "where is my refund",
    "the checkout button is the wrong colour",
])

print(result.headline)
NOT TRUSTWORTHY - the agent answered 'general' on 100% of the probes,
so a pass says more about the probe set than about the agent.

Those deliberately varied test inputs form the probe set. AgentVerity asks:

  • Decision stability: does one case reach the same decision across isolated reruns?
  • Observed decision spread: do the cases reach more than one decision?
  • Declared decision coverage: did the suite include every required decision, did the agent return them, and did it emit an unknown decision?

A green quality score answers a different question: whether the selected answers were right. AgentVerity checks how much evidence that score rests on. It complements DeepEval, promptfoo, AgentCore Evaluations, and ordinary assertions rather than replacing them.

Together, the checks guard against two failure modes:

  • Vacuous green: every supplied assertion passes, but the test inputs reach only one decision.
  • Regression trap: that narrow run becomes the baseline, so later changes to untested decisions remain invisible.

AgentVerity qualifies the evidence from this run. It does not certify the agent as correct, safe, or fully covered.

Why rerun counts are harder than they look

Three or five reruns chosen by convention can support the wrong conclusion. In one deterministic example, 36 non-overlapping rerun comparisons produced no decision changes. That was still too little evidence to certify a change rate below 5%. The honest result was undecided, not unstable. Certifying that threshold with no observed changes needed 73 comparisons.

The decision rule is:

  • pair reruns without reusing an output
  • calculate a 95% Wilson interval around the observed decision-change rate
  • report deterministic when the upper bound is below the tolerated rate
  • report stochastic when the lower bound is above the tolerated rate
  • report undecided when the interval spans that tolerance

Wilson intervals are established statistics. AgentVerity's design choice is to use one as a three-outcome release rule rather than force an underpowered run into stable or unstable. Non-overlapping pairs avoid making the sample look larger than the target calls justify, while the requested tolerance determines the call budget before execution.

The default balanced precision calculates that budget automatically.

See the executable helper, arithmetic, and exact API mapping.

Why not write a rerun loop in pytest?

A loop can collect repeated calls. The difficult part is deciding what those calls support: forming non-overlapping comparisons, sizing them from a tolerance, preserving undecided, enforcing stricter route targets, and keeping terminal, JSON, JUnit, telemetry, and snapshot admission consistent. If your evaluation platform already implements that policy, keep it. AgentVerity packages the policy for teams that would otherwise maintain it themselves.

The evidence gate

The evidence gate refuses to save a baseline until calls complete, decisions are stable enough, the probe set crosses a decision boundary, and a person approves the reference outputs.

The bundled payment-dispute example runs two test sets:

python examples/payment_dispute_gate.py
Probe set Exact-match Verdict stability Declared coverage Baseline
Narrow, 6 duplicate-charge cases ✅ 6/6 ✅ verdict-deterministic ❌ 1/6 required routes ❌ REFUSED
Repaired, 6 dispute categories ✅ 6/6 ✅ verdict-deterministic ✅ 6/6 required routes ✅ ADMITTED

Both score 6/6. The narrow set correctly tests one route, but it covers only one of the six routes declared as required. The repaired set reaches all six and can be saved as a versioned snapshot.

Declare that test contract in Python:

from agentverity import DecisionCase, DecisionContract, DecisionSuite, run

suite = DecisionSuite(
    contract=DecisionContract(
        allowed={"duplicate_charge", "refund_delay", "card_security"},
        critical={"card_security"},
    ),
    cases=(
        DecisionCase("I was charged twice", "duplicate_charge"),
        DecisionCase("My refund is late", "refund_delay"),
        DecisionCase("I do not recognise this card payment", "card_security"),
    ),
)

result = run(agent, suite=suite)

expected records the route each case is intended to exercise. AgentVerity checks intended and observed route coverage separately. Your assertions or quality evaluator still decide whether each returned route was correct.

Declaring a suite also splits the existing stability evidence by intended route. A noisy route is named instead of being averaged together with quieter ones, and the report lists the decision pairs that changed. A route proven stochastic blocks snapshot admission even when the pooled result looks deterministic.

Use stability_targets={"card_security": 0.05} when a route needs its own release condition, then inspect the zero-change cost before calling the agent:

agentverity plan --suite examples/route_stability_plan.json

A targeted route that remains undecided blocks snapshot admission. An explicit budget remains a hard cap, so an unaffordable plan is refused before any agent call.

Breadth remains separate from repeats. Declare minimum_cases={"card_security": 3} when review policy requires several cases for a route. This counts written cases and does not pretend three paraphrases are semantically diverse. When relations are enabled, the report also names routes that no transformation actually changed. Those are relation checks that were not exercised, not 0% violation results.

Create one through the CLI:

agentverity snapshot \
  --agent examples/payment_dispute_gate.py:build_agent \
  --suite examples/payment_decisions.json \
  --output baseline.json \
  --accept-reference

The same checks run before agentverity check reports differences as regressions. Snapshot files retain SHA-256 input fingerprints rather than raw prompts.

Where it fits

AgentVerity is an evaluation runner, not serving-path middleware. It admits or refuses evidence produced beside the evaluator a team already uses:

reviewed cases ---> agent ---> correctness / trajectory evaluator ---> quality result
       |
       +---- isolated reruns ---> agent ---> AgentVerity stability + contract
                                                   |
quality result ------------------------------------+----> release policy
                                                         |          |
                                                       admit      refuse
                                                         |
                                                regression baseline in CI
                                                         |
                                               synthetic production canary

Use it while developing, on a pull request, before release, or as a scheduled synthetic canary. Do not repeat live customer requests. Results can leave as text, JSON, JUnit XML, or one privacy-minimised OpenTelemetry span.

Read how per-route evidence works, with worked examples.

See CI, telemetry, lifecycle, and multi-agent integration.

Measured AgentCore canary

The optional production example combines a Strands payment router on Amazon Bedrock, DeepEval route-quality checks, AgentVerity, AgentCore Runtime, and CloudWatch.

A real AgentCore canary passes DeepEval quality, AgentVerity evidence, and cloud health checks before its baseline is admitted

At its declared 10% canary tolerance, the London run recorded 6/6 correct routes, no changes across 36 repeat pairs, all six routes reached, and 78 successful cloud calls with no errors or throttles. Its first run was stable but only 5/6 correct, so the example now requires both quality and evidence before snapshot admission.

The v0.9 contract path was then rerun directly through Bedrock after the old runtime was removed. It again scored 6/6 with 0/36 route changes, and reported all six required routes intended and observed with no unknown decision.

This is deployment proof, not an AWS requirement. The zero-dependency callable works with any stack.

Run the production example · Read the measured result

Scope

A trustworthy AgentVerity result means that, for the supplied probe set, optional decision contract, and chosen tolerance:

  • repeated isolated calls provided enough evidence about decision stability
  • observed decisions did not collapse onto one highly dominant route
  • every required decision was intended and observed when a contract was given
  • no observed decision fell outside that contract
  • execution completed, and at least one requested relation genuinely changed an input. Per-route relation gaps remain explicit diagnostics

That is a minimum dynamic adequacy check, not exhaustive branch coverage. Without a decision contract, AgentVerity only checks observed diversity. With one, it reports required, intended, observed, missing, and unknown decisions. This is required-decision presence, not comprehensive behavioural-boundary coverage. It does not establish that every boundary was tested or that critical routes are correct. Keep labelled correctness cases beside it. critical marks consequence for coverage reporting, while stability_targets separately declares any route-specific statistical policy.

It also does not judge answer correctness, prove safety, store traces, host a dashboard, or monitor production traffic. Static tools remain useful for declared branches, route schemas, and expected labels. AgentVerity measures the decisions a model-backed or black-box target actually returns.

Documentation

Development

pip install -e ".[dev]"
python -m pytest -q
ruff check .

CI covers Python 3.10 through 3.14, lint, package construction, and the generated README evidence. A coverage job enforces at least 90% statement coverage, and the branch-protection CI gate requires that job to pass.

Status and licence

Alpha. Pin a minor series for production use, for example agentverity~=0.10.0. Patch releases preserve the public API.

Apache-2.0. Contributions are welcome through the pull-request workflow.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentverity-0.11.0.tar.gz (642.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentverity-0.11.0-py3-none-any.whl (66.0 kB view details)

Uploaded Python 3

File details

Details for the file agentverity-0.11.0.tar.gz.

File metadata

  • Download URL: agentverity-0.11.0.tar.gz
  • Upload date:
  • Size: 642.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.11.0.tar.gz
Algorithm Hash digest
SHA256 034426c05023b49a5c7ae1887b15e414cd3fee6e30e0e76756639ca6a5582c62
MD5 fa797d04963d4611046123422305b6c2
BLAKE2b-256 1a9d99099c1af4185cdd876e5bbb30c2593c8e7339b36cceb2b502cc0e07e6ad

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.11.0.tar.gz:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agentverity-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: agentverity-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 66.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 096563bb4997bac0b603c7bc4e64896ff7ca73c75e53430f4b3c66351f686f63
MD5 bc8aaf75fa613ee6dc87657b8c1d313d
BLAKE2b-256 1414f2137ba5f45c04ae62bab893b601cd2ccb7f3a0d68d403014959a0b080f1

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.11.0-py3-none-any.whl:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page