Skip to main content

AgentVerity

Your agent test passed. Would it pass again?

PyPI Python 3.10+ CI Coverage: 90%+ License: Apache-2.0

AgentVerity is an offline Python library and CLI that qualifies repeated, categorical AI-agent evidence before it becomes a regression reference for future releases. It finds routes whose repeatability is rejected, weak decision coverage, and runs too small to support a conclusion. It does not judge whether an answer is correct.

The 60-second problem

A payment router sends disputes to six specialist queues. Promptfoo runs six reviewed cases 26 times, and all 156/156 assertions pass. One ambiguous case allows either of two valid fraud queues.

AgentVerity reads that same export and finds:

route              cases  pairs  flips  95% CI            result
card_security          1     13      8  [0.355, 0.823]    stochastic
cash_withdrawal        1     13      0  [0.000, 0.228]    undecided
duplicate_charge       1     13      0  [0.000, 0.228]    undecided

flip pairs:
  card_security <-> merchant_dispute  x8

The quality policy accepts both answers, but a reference that switches queues will make later regression checks noisy. The changing route is stochastic; the five quiet routes are undecided because 13 pairs are too few to certify them separately. A flip means the two observations in a paired rerun differed. AgentVerity therefore refuses this evidence as a regression reference.

Try it without model calls

git clone --depth 1 https://github.com/mrwersa/agentverity.git
cd agentverity
python -m pip install .
agentverity assess \
  --promptfoo examples/promptfoo_bridge/results.json \
  --suite examples/payment_decisions.json

assess performs arithmetic over recorded decisions. It makes no model or provider calls. You can also reuse precomputed DeepEval LLMTestCase objects or any ordered JSONL log:

agentverity assess --jsonl runs.jsonl \
  --input-path probe.text --decision-path result.route

Order matters because observations are paired in collection order. See imported evidence before converting a log.

To call an agent directly, install only the framework adapter you need:

pip install "agentverity[strands]"
pip install "agentverity[langgraph]"

Plain Python callables need no extra dependency:

from agentverity import from_callable, run


def route(text):
    return "billing" if "charge" in text.lower() else "cash_withdrawal"


agent = from_callable(lambda text: {"verdict": route(text)})
result = run(agent, inputs=["duplicate charge", "cash withdrawal"])
print(result.summary())

What it decides

AgentVerity keeps three API outcomes separate and explains them in repeatability terms:

  • deterministic: repeatability qualified; evidence supports the tolerance
  • stochastic: repeatability rejected; changes exceed the tolerance
  • undecided: inconclusive; the evidence supports neither direction

These strings remain the public machine contract. deterministic does not claim that the underlying agent has zero randomness.

It then checks whether the probe set collapsed onto one decision and, when a decision contract is supplied, whether every required route was intended and observed. Per-route results show where changes concentrate. Optional relations check reviewed input transformations and report no-op transforms as untested, not passed.

Once you have two evidence windows, agentverity compare-evidence before.json after.json reports changed route conclusions, gained or lost decisions, changed flip pairs, isolation, and provenance.

Where it fits

Layer Question
Promptfoo, DeepEval, Ragas, or labelled assertions Was the answer acceptable?
AgentVerity Is the repeated categorical evidence strong enough to preserve as a regression reference?
LangSmith, Phoenix, AgentCore, or another trace system What happened during the run and in production?
Security and authority tests Was the agent allowed to take that action?

AgentVerity is a local test and release step, not serving-path middleware. Use it for named routes, approvals, policy outcomes, tool choices, hand-offs, or a reviewed finite tool path. Use another evaluator for open-ended chat, RAG quality, generated content, or coding-agent output. If those systems also emit a bounded route or approval, AgentVerity can qualify that decision layer.

Command Purpose
agentverity plan Price the best-case evidence budget without calling an agent
agentverity run Collect and assess isolated repeated decisions
agentverity assess Assess Promptfoo, DeepEval, JSONL, or native evidence
agentverity snapshot Admit a human-reviewed regression reference when evidence permits
agentverity check Re-run the admission policy and compare with a snapshot
agentverity compare-evidence Compare two independently collected evidence windows

Why rerun counts are harder than they look

Three or five reruns by convention do not state what variation they can rule out. With no observed changes:

  • 36 independent pairs bound the change rate below about 9.6%
  • a claim below 5% needs 73 pairs
  • a short quiet run is therefore undecided, not repeatability qualified

AgentVerity sizes calls from the tolerance, uses non-overlapping pairs, and places a Wilson interval around the flip rate. Optional sequential collection uses checkpoints declared before collection; it does not repeatedly inspect a fixed-sample interval and stop when the result looks favourable.

For evidence already collected, best_case_admission_pairs tests whether an all-agree continuation could admit within a predeclared pair budget. It may justify stopping an impossible run early; it never creates an early admission. For live fixed-endpoint collection, opt into the same bound with --curtail. It reports the stopping pair and avoided calls but no final repeatability class; a path that could admit still pays the full fixed budget and is classified only there.

Use agentverity plan --suite examples/route_stability_plan.json before spending remote calls. The decision repeatability method guide explains the arithmetic, and the validation artifact records exact-boundary checks and dependence stress tests.

The evidence gate

snapshot refuses a regression reference until calls complete, the evidence supports the declared repeatability and coverage policy, and a person approves the reference as acceptable. The bundled offline example shows why correctness alone is not enough:

python examples/payment_dispute_gate.py
Probe set Exact-match Verdict repeatability Declared coverage Reference
Narrow, 6 duplicate-charge cases ✅ 6/6 ✅ verdict-deterministic ❌ 1/6 required routes ❌ REFUSED
Repaired, 6 dispute categories ✅ 6/6 ✅ verdict-deterministic ✅ 6/6 required routes ✅ ADMITTED

Both sets score 6/6. Only the repaired set reaches all six required routes.

Real-system evidence is also committed and reproducible without new calls:

  • The AgentCore canary validates a production-shaped integration while explicitly stopping short of per-route certification.
  • The AgentKit study records 4,380 calls across three models and shows that the most repeatable model can be less correct.

Never repeat live customer requests. Use reviewed synthetic cases in CI, before release, or on a schedule.

What it does not prove

TRUSTWORTHY means the supplied cases produced repeatable, non-collapsed evidence at the declared tolerance and satisfied any declared decision contract. It does not prove correctness, safety, semantic diversity, complete behavioural coverage, provider independence, or production reliability. AgentVerity also does not store traces, host a dashboard, monitor traffic, or score open-ended answers.

Documentation

Development

pip install -e ".[dev]"
python -m pytest -q
ruff check .

CI covers Python 3.10–3.14, package construction, and at least 90% statement coverage. See the contributing guide above before opening a pull request.

Status and licence

Alpha. Pin the current minor series for production use: agentverity~=0.21.0. Patch releases preserve the public API.

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentverity-0.21.0.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentverity-0.21.0-py3-none-any.whl (112.1 kB view details)

Uploaded Python 3

File details

Details for the file agentverity-0.21.0.tar.gz.

File metadata

  • Download URL: agentverity-0.21.0.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.21.0.tar.gz
Algorithm Hash digest
SHA256 21f4f56aadcbb4e11f69ef7846c2b35d63fb870e064bcacd982b4d3db1deeea4
MD5 9b4be9593eafea70301fd6e5e612436f
BLAKE2b-256 62c750e524d99c2e1a804923532fea74988f1bd34905990da64e13afb23188ab

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.21.0.tar.gz:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agentverity-0.21.0-py3-none-any.whl.

File metadata

  • Download URL: agentverity-0.21.0-py3-none-any.whl
  • Upload date:
  • Size: 112.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentverity-0.21.0-py3-none-any.whl
Algorithm Hash digest
SHA256 fe8ad299dff12b3e8c00264f32525358b090cd1faac0cfc65bb99443d104b7f1
MD5 3bd130c54cff1b769f6dc8030181eef1
BLAKE2b-256 7d3eb8ffd72b1e8621ff29e8af1ab8b1ee380fc51b76a6ff200e16e1f70480dc

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentverity-0.21.0-py3-none-any.whl:

Publisher: release.yml on mrwersa/agentverity

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.23.1

2 files

0.23.0

2 files

0.22.0

2 files

This release

0.21.0 This release

2 files

0.20.0

2 files

0.19.0

2 files

0.18.3

2 files

0.18.2

2 files

0.18.1

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.2

2 files

0.13.1

2 files

0.13.0

2 files

0.12.2

2 files

0.12.1

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.1

2 files

0.9.0

2 files

0.8.6

2 files

0.8.5

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page