Skip to main content

AgentVerity

Your agent test passed. Would it pass again?

PyPI Python 3.10+ CI Coverage: 90%+ License: Apache-2.0

A Python library and CLI that reruns the decisions your agent makes and reports whether they are repeatable enough, and varied enough, to trust as a regression baseline: the reviewed run you save as expected behaviour and compare future releases against.

Keep using Promptfoo, DeepEval, or your own assertions to decide whether each answer is correct. AgentVerity reuses those same results to find unstable routes, missing decisions, and runs too small to support a reliable baseline. It can read runs you already collected, so you do not pay for the same calls twice.

What you run Question it answers
Promptfoo, DeepEval, your assertions Was this answer acceptable?
AgentVerity Are the repeated answers stable and covered enough to save as expected behaviour?
LangSmith, AgentCore, your observability What happened during this run, and in production?
AgentMandate What is this agent permitted to do, and did this release widen it?

See it in a full release gate: agent-release-gate runs AgentVerity beside authority analysis on one agent, offline, including the case where a route shows zero flips over six pairs and an interval running to 39%.

The 60-second problem

Consider a routing workflow: a model classifies each payment dispute and directs it to one of six specialist queues. Promptfoo runs six cases 26 times. Its configured quality checks accept either fraud queue for one ambiguous card-security case. All 156/156 assertions pass.

Would you save that run as the expected behaviour for future releases?

AgentVerity reuses the same Promptfoo export and finds this:

4. STABILITY BY ROUTE
   route              cases  pairs  flips  95% CI            result
   card_security          1     13      8  [0.355, 0.823]    stochastic
   cash_withdrawal        1     13      0  [0.000, 0.228]    undecided
   duplicate_charge       1     13      0  [0.000, 0.228]    undecided
   ...
   flip pairs:
     card_security <-> merchant_dispute  x8

The contract check passes, but the decision switches between card_security and merchant_dispute in 8 of 13 paired reruns.

Why you should care

  • Those labels send work to different queues, controls, and owners.
  • A moving reference makes later regression failures noisy and hard to trust.
  • One pooled score hides which route is moving.
  • Zero observed changes do not prove a quiet route is stable when the sample is too small.

AgentVerity names the unstable route and leaves the five underpowered routes undecided. It will not freeze this run as a baseline.

Try it without model calls

The repository includes that recorded Promptfoo run:

git clone --depth 1 https://github.com/mrwersa/agentverity.git
cd agentverity
python -m pip install agentverity
agentverity assess \
  --promptfoo examples/promptfoo_bridge/results.json \
  --suite examples/payment_decisions.json

The last command performs arithmetic over saved decisions. It makes no model or provider calls.

Already using DeepEval? Pass the same precomputed LLMTestCase objects to evidence_from_deepeval. Any harness can use the small neutral evidence format.

Keep your existing evaluator for correctness and trajectory quality. AgentVerity checks whether the repeated results are stable and complete enough for a regression baseline.

If you would rather AgentVerity make the calls itself, install the adapter for your framework. The core has no agent library as a dependency, so this is the only place one is needed:

pip install "agentverity[strands]"     # Strands Agents
pip install "agentverity[langgraph]"   # LangGraph

Where to integrate it

AgentVerity is a test and release step, not serving-path middleware.

Stage Use it for
Local development Diagnose a moving or one-route-only test set
Pull request Qualify a candidate baseline and publish JUnit
Pre-release Refuse unstable, incomplete, or underpowered evidence
Scheduled canary Recheck reviewed synthetic cases and emit OpenTelemetry

One real integration combines DeepEval quality, AgentVerity evidence, and Amazon AgentCore health before admitting a baseline:

A real AgentCore canary combines DeepEval quality, AgentVerity evidence, and cloud health before baseline admission

Never repeat live customer requests. Use reviewed synthetic cases in CI, before release, or on a schedule.

Where it sits in the evaluation loop

  1. Explore capability. Use challenging cases to learn what the agent can and cannot do.
  2. Evaluate quality. Keep Promptfoo, DeepEval, or your current assertions for outcomes and trajectories.
  3. Qualify the evidence. Import repeated decisions into AgentVerity without calling the model again.
  4. Fix what is missing. Repair moving routes, add missing cases, or collect the reruns needed for an honest conclusion.
  5. Promote a regression case. A mature capability case becomes a reusable baseline only after a human approves it and the evidence gate admits it.
  6. Monitor and learn. Review canary failures and production incidents, then add suitable cases back to the offline dataset.

The import command diagnoses decisions already collected by another evaluator. The snapshot and check commands below provide the same admission policy when AgentVerity calls your agent directly.

Outcomes, trajectories, and decisions remain separate evidence:

Layer Example question
Outcome Was the refund recorded in the case system?
Trajectory Which tools or agents acted, and in what order?
Decision Did the workflow choose refund, review, or deny?

An evaluation framework can grade the first two. AgentVerity qualifies the repeatability and declared coverage of the bounded decision before that result becomes a regression reference.

What it checks

Check Developer question
Decision stability Does the same case keep reaching the same decision?
Observed spread Did all test inputs collapse onto one decision?
Declared coverage Were all required decisions and critical routes represented and returned?
Route evidence Which route moves, and which quiet routes still lack enough reruns?
Relation coverage Did an input transformation genuinely exercise each route, or was it a no-op?

It keeps three outcomes separate:

  • stable enough for the declared tolerance
  • unstable above that tolerance
  • undecided because the run did not collect enough evidence

Is it for my agent?

Use it when:

  • the component chooses from named routes, approvals, policies, tools, or hand-offs
  • repeated runs can start from equivalent isolated state
  • you can write deliberately varied cases for the decisions that matter

It fits the agent workflow patterns whose value is a named choice rather than open prose:

Pattern The decision under test
Routing Which specialist path or queue an input is classified into
Orchestrator-workers Which worker the orchestrator dispatches to next
Evaluator-optimiser Whether the evaluator accepts, revises, or rejects
Tool use Which tool is selected, and the ordered tool path taken
Multi-agent supervisor Which agent receives the handoff
Guardrail or policy gate Approve, review, escalate, or deny

Concrete examples: support and payment routing, fraud triage, incident dispatch, approval flows, and bounded tool selection.

Use another evaluator for open-ended chat, RAG quality, generated content, or coding-agent output. If such a system also emits a reviewed route or approval, AgentVerity can qualify that decision layer.

Check applicability and exact limits.

Why rerun counts are harder than they look

Picking three or five reruns by convention is guesswork:

  • 36 paired reruns with no changes only bound the change rate below 9.6%.
  • A claim below 5% needs 73 zero-change pairs.
  • A short quiet run is therefore undecided, not proven stable.

AgentVerity sizes the run from your tolerance, uses non-overlapping pairs, and keeps three answers: stable enough, unstable, or undecided. The default balanced setting uses a 5% tolerance.

A small pytest loop can collect calls. The library packages the harder policy: evidence sizing, route-specific targets, and one consistent decision across text, JSON, JUnit, telemetry, and snapshots.

Read the executable arithmetic and design.

The evidence gate

The evidence gate refuses to save a baseline until:

  • calls complete
  • decisions are stable enough
  • the cases reach the required decisions
  • a person approves the reference outputs as correct

The bundled payment-dispute example runs two test sets:

python examples/payment_dispute_gate.py
Probe set Exact-match Verdict stability Declared coverage Baseline
Narrow, 6 duplicate-charge cases ✅ 6/6 ✅ verdict-deterministic ❌ 1/6 required routes ❌ REFUSED
Repaired, 6 dispute categories ✅ 6/6 ✅ verdict-deterministic ✅ 6/6 required routes ✅ ADMITTED

Both score 6/6. The narrow set is a valid unit test for one route, but it is not a system-wide baseline. The repaired set reaches all six required routes and can be admitted.

Before spending model calls, inspect the zero-change evidence budget:

agentverity plan --suite examples/route_stability_plan.json

Then create the reviewed baseline:

agentverity snapshot \
  --agent examples/payment_dispute_gate.py:build_agent \
  --suite examples/payment_decisions.json \
  --output baseline.json \
  --accept-reference

The same checks run before agentverity check reports differences as regressions. Snapshot files retain SHA-256 input fingerprints rather than raw prompts.

The contract can also declare stricter stability targets for critical routes and minimum case counts. Repeats support a stability claim. Distinct reviewed cases support breadth. AgentVerity keeps those two claims separate.

Measured production example

The optional production example combines a Strands routing agent on Amazon Bedrock, DeepEval route-quality checks, AgentVerity, AgentCore Runtime, and CloudWatch.

At its declared 10% canary tolerance, the London run recorded 6/6 correct routes, no changes across 36 repeat pairs, all six routes reached, and 78 successful cloud calls with no errors or throttles. An earlier run was stable but only 5/6 correct. The release policy therefore requires both quality and evidence rather than treating either tool as sufficient.

This is deployment proof, not an AWS requirement. The zero-dependency callable works with any stack.

Run the production example · Read the measured result

What it does not prove

TRUSTWORTHY means the supplied cases produced stable, non-collapsed evidence at the declared tolerance and satisfied any declared decision contract.

It does not prove:

  • that each decision was correct or safe
  • that every code branch or behavioural boundary was tested
  • that several cases are semantically diverse
  • that an open-ended answer is high quality

It also does not store traces, host a dashboard, or monitor production traffic. Static coverage, Promptfoo or DeepEval quality checks, and production observability remain separate parts of the stack.

Go deeper

Read the design story: Introducing AgentVerity: What Does a Green Agent Test Prove?

Development

pip install -e ".[dev]"
python -m pytest -q
ruff check .

CI covers Python 3.10 through 3.14, lint, package construction, and the generated README evidence. A coverage job enforces at least 90% statement coverage, and the branch-protection CI gate requires that job to pass.

Status and licence

Alpha. Pin a minor series for production use, for example agentverity~=0.13.0. Patch releases preserve the public API.

Apache-2.0. Contributions are welcome through the pull-request workflow.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentverity-0.14.0.tar.gz (899.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentverity-0.14.0-py3-none-any.whl (83.9 kB view details)

Uploaded Python 3

File details

Details for the file agentverity-0.14.0.tar.gz.

File metadata

  • Download URL: agentverity-0.14.0.tar.gz
  • Upload date:
  • Size: 899.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.14.0.tar.gz
Algorithm Hash digest
SHA256 7a6ec001c5d67aa34e6b6f652e172f16c327654bc2cd33b7d67fd7d4db3d1e47
MD5 6f134eec4909119b61b1f5de1b6af01e
BLAKE2b-256 2a613c0e1f5505f3b79295e2b35c14d960a042cc39269815bf7ae5107760b7cb

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.14.0.tar.gz:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agentverity-0.14.0-py3-none-any.whl.

File metadata

  • Download URL: agentverity-0.14.0-py3-none-any.whl
  • Upload date:
  • Size: 83.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.14.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b37a46b7d6167ddbeb640b185341b10a5e6e952a0ce10b1484ca19d744afb8bc
MD5 04c85e2c2780d65a3a64085a79177631
BLAKE2b-256 68815b08b5d29234e06fa62a0bf8ad2ecfdb6c9c3d016e1dd260e59a79f88a60

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.14.0-py3-none-any.whl:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page