Skip to main content

AgentEval

CI for AI agents — turn flaky agent behavior and production failures into tests that fail the PR.

AgentEval regression gate PyPI version License: MIT Python 3.10+

Static demo · Streamlit dashboard · PyPI: nishanttyagi-agenteval · CLI: agenteval


The problem

You change a prompt, model, or tool. The agent still “answers.” Unit tests still pass. A week later someone finds a wrong refund, a hallucinated fact, or a tool the agent never should have called.

Normal tests prove the code ran. They do not prove the agent still behaves.

In simple terms

Imagine a support agent gets:

“Cancel order 4821 and refund the customer.”

A green unit test only means something came back. AgentEval checks the behavior:

Question What gets checked
Did it solve the request? Final answer vs expected outcome
Did it invent anything? Unsupported claims vs ground truth
Did it use the right tools? e.g. lookup_order + issue_refund, not a random tool
Did it take the right steps? Trajectory vs expected sequence
Is it stable? Optional repeats → stable / flaky / unstable
Did this change make it worse? Current run vs a versioned baseline (CI exit code)
Has this failed in production before? Failure Memory → approved golden case → CI again

If the agent says “refunded” without calling the refund tool, that can look fine to a human skim. AgentEval records it as a failure and can block the PR.

What I built

Four outcomes, not a platform pitch:

  1. A regression gate for agents — YAML golden suites, scored runs, compare to a git-trackable baseline, fail CI when quality drops.
  2. Evidence, not one score — correctness, hallucination, tool-call accuracy, latency/cost, trajectory, flakiness, plus optional RAG checks and a SQL safety scanner.
  3. Failure Memory — capture → redact → cluster → replay → minimize → human approve → golden YAML → CI, so the same production failure is harder to ship twice.
  4. One CLI across frameworks — adapters for CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, and custom agents; composite GitHub Action; templates (including an Indic-language pack).

A few numbers

Evidence from this repo on main — no invented percentages.

Fact Value
Automated tests (this checkout, optional extras absent) 1152 passed, 3 skipped (Docker, FastAPI, KarmaSakshi)
Failure Memory suite 59 passed
Test modules (tests/**/test_*.py) 105
Package version on main 0.5.0
Latest published PyPI release 0.3.0 (PyPI)
Top-level CLI commands 19 (run, compare, report, generate, generate-adversarial, import, generate-cases, init, compare-models, trace, diff, calibrate, audit-log, serve, gate, plugins, templates, sql, memory)
agenteval memory subcommands 19
Bundled templates 4 — coding-agent (7), customer-support (7), rag-assistant (7), indic-agent (34 cases: 28 offline / 6 opt-in LLM-judge)
Framework adapters CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, KarmaSakshi bridge (+ custom)
Optional extras dev, crewai, autogen, openai-agents, karmasakshi
Python 3.10+
License MIT
Classifier Alpha (Development Status :: 3 - Alpha)
CI workflows in-repo eval.yml, failure-memory.yml, karmasakshi-bridge.yml, action-smoke.yml, docker.yml, publish.yml, landing-page.yml
Offline demos examples/mock_agent, examples/failure_memory_demo(_v21), examples/indic_mock_agent, examples/karmasakshi_bridge (no API key)

Why this matters

If you ship agent changes, you want (1) a hard gate when known behavior regresses, and (2) a memory of real failures that does not reset every run. AgentEval is built for that workflow: golden YAML and baselines live in git; Failure Memory stays local-first with human approval before anything blocks CI; demos run without network or provider keys.

How it works

Simple flow

  1. Write expected behavior as YAML golden cases (or scaffold with agenteval init).
  2. Run the agent through an adapter: agenteval run.
  3. Compare to a baseline: agenteval compare (CI fails on regression).
  4. Optionally: ingest a production failure into Failure Memory → redact → replay → minimize → approve → export golden → run with --production-cases.
  5. Inspect evidence: JSON/HTML report, Streamlit dashboard, or agenteval trace / diff.

Optional architecture (for engineers)

Agent adapter + golden YAML  →  runner / evaluators  →  JSON run + HTML report
                                                      →  baseline compare / CI gate

Production traces  →  redact  →  Failure Memory (SQLite)
                              →  cluster / replay / minimize
                              →  human approve  →  golden YAML  →  same CI gate

Implemented in-repo: adapters/, core/, evaluators/, failure_memory/, CLI entry agenteval → agenteval.cli:main, optional Streamlit dashboard, composite Action (action.yml).


Install

Python 3.10+.

From PyPI (latest published release is 0.3.0):

pip install nishanttyagi-agenteval==0.3.0
agenteval --version
agenteval --help

From this repo (current main is 0.5.0):

git clone https://github.com/nishanttyagi28/agenteval.git
cd agenteval
python -m pip install -e ".[dev]"
agenteval --version   # expect 0.5.0 on main
python -m pytest -q

Framework extras (optional): pip install "nishanttyagi-agenteval[crewai]" (same for autogen, openai-agents, karmasakshi).

Quick start

Offline mock agent (no provider)

agenteval run --agent mock_agent --registry examples/mock_agent/agents.yaml
agenteval compare --agent mock_agent --registry examples/mock_agent/agents.yaml

Failure Memory demo (zero network, no API key)

python examples/failure_memory_demo_v21/run_demo.py

Capture → redact → cluster → replay → minimize → export golden → fail broken agent / pass fixed agent → coverage check. See examples/failure_memory_demo_v21/ and docs/failure-memory.md.

KarmaSakshi bridge demo (offline seal → witness → score)

pip install -e ".[dev,karmasakshi]"
python examples/karmasakshi_bridge/run_demo.py

Approved ₹1500→Priya: correct attempt passes; ₹1501 or wrong payee is blocked by KarmaSakshi and recorded as an AgentEval failure. See examples/karmasakshi_bridge/ and docs/karmasakshi-bridge.md.

Scaffold a project

agenteval init

Release Desk (v0.5.0)

Turn a production miss into a CI test without leaving the browser. No API key.

agenteval gate --local

Open the URL it prints (default http://127.0.0.1:8741). Load the refund sample, or POST an incident:

curl -s -X POST http://127.0.0.1:8741/api/gate/ingest \
  -H 'Content-Type: application/json' \
  -d '{"incident":{"agent":"refund-desk","prompt":"Refund order 4821","output":"Cancelled. No refund.","tools_called":["cancel_order"],"expected_tools":["lookup_order","issue_refund"]}}'

Ship writes .agenteval/production-regressions.yaml. The next run that still does this fails:

agenteval run --agent my_agent --production-cases .agenteval/production-regressions.yaml

The desk is local-only (pass --local). It redacts secret-shaped strings before storage and will not bind a public address unless you pass --allow-remote. See docs/release-desk.md.

Open-ended llm_judge cases no longer require the data-analyst sibling repo. With no key, the deterministic offline judge scores them. Set AGENTEVAL_JUDGE_PROVIDER to openai, groq, anthropic, or legacy.

Indic-language pack (v0.4.0 on main)

pip install -e examples/plugins/agenteval-indic-evaluators
agenteval run --agent indic_mock_agent \
  --registry examples/indic_mock_agent/agents.yaml \
  --tag core --no-llm-judge

34 cases (28 deterministic offline; 6 refusal/safety need an LLM judge). The mock demo expects deliberate failures so each checker is shown catching something — see examples/indic_mock_agent/README.md.

CLI

agenteval run | compare | report | generate | generate-adversarial | import | generate-cases
agenteval init | compare-models | trace | diff | calibrate | audit-log | serve
agenteval plugins | templates | sql | memory | gate
Command Role
run / compare Golden suite + baseline regression gate
report Self-contained HTML report
init Scaffold registry, sample cases, CI workflow
trace / diff Step evidence and trajectory diff
generate / generate-adversarial Reviewable adversarial / red-team candidates (not auto-blocking)
memory … Failure Memory loop (see below)
gate Local Release Desk: ingest a production failure, approve it, write the CI case
sql scan SQL agent structural safety scan
templates / plugins Bundled starters and evaluator entry points

Failure Memory CLI

agenteval memory init
agenteval memory ingest traces.jsonl
agenteval memory cluster
agenteval memory list
agenteval memory review <candidate_id> approve --correctness-type contains --ground-truth "…"
agenteval memory replay <candidate_id>
agenteval memory minimize <candidate_id>
agenteval memory approve-minimization <minimization_id>
agenteval memory export-minimized <minimization_id>
agenteval memory coverage
agenteval run --agent my_agent --production-cases .agenteval/production-regressions.yaml

Default DB: .agenteval/failure-memory.db (--db or AGENTEVAL_FAILURE_MEMORY_DB).

Production failure → redact → ingest → cluster → replay → minimize
  → human approve → golden YAML → CI gate → recurrence / coverage

What AgentEval evaluates

Capability What you get
YAML golden suites Versioned prompts, expectations, tools, tags
Correctness Exact, contains, numeric, numeric-table, optional LLM judge
Hallucination Unsupported claims vs ground truth
Tool-call accuracy Required tools: precision / recall / F1
Latency & cost p50/p95 and suite cost; opt-in budget gates
Trajectory Expected steps (LCS F1); agenteval diff
Flakiness Optional repeats; stable / flaky / unstable
Baseline regression Compare to versioned baseline; CI exit codes
RAG mode Context relevance, faithfulness, citation checks
SQL safety scanner Structural / policy-oriented checks (agenteval sql)
Failure Memory Production failure → approved golden regression
Adapters CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, KarmaSakshi bridge, custom
GitHub Action Composite action + in-repo regression workflow
Reports Streamlit, HTML, local read-only API (serve)

Framework registry example

version: 1
agents:
  mock_agent:
    adapter: examples.mock_agent.adapter:MockAgentAdapter
    cases: examples/mock_agent/cases.yaml
    enabled: true

Composite Action for consumer repos:

- uses: nishanttyagi28/agenteval@v0.3.0
  with:
    agent: my_agent
    config-file: agents.yaml
    cases-file: tests/golden/cases.yaml
    baseline-file: baselines/my_agent.json

Pin a release tag you trust. main moves; PyPI 0.5.0 is not published yet. The latest published release remains 0.3.0.

Security and privacy defaults

Default Meaning
Content capture off Prompts/outputs not stored unless you opt in
Redaction before disk Applied before SQLite / JSONL persistence
Best-effort DLP Common secret/PII patterns only — not a universal guarantee
Human approval Required before golden promotion; nothing auto-enters blocking CI
Local-first SQLite Default under .agenteval/; no hosted multi-tenant control plane
No telemetry by default Failure Memory does not phone home
Keep secrets out of git Do not commit DBs, raw traces, or sensitive generated suites

Limitations and non-goals

  • Local-first / single-user Failure Memory — not a hosted multi-tenant product
  • Not an OpenTelemetry collector (optional OTel-shaped JSON helpers only)
  • No mandatory vector DB or embeddings for clustering
  • No mandatory LLM judge for classification / clustering
  • Redaction will miss arbitrary unlabeled secrets
  • Live agent eval still needs your runtime and any provider keys you choose
  • Cost may fall back to estimates when providers omit usage
  • Adversarial / generated cases stay candidates until a human promotes them
  • Not claiming v1 stability — see docs/v1-readiness.md

Documentation

Resource Link
Release Desk docs/release-desk.md
Failure Memory docs/failure-memory.md
KarmaSakshi bridge docs/karmasakshi-bridge.md
Compatibility docs/compatibility.md
SQL scanner docs/sql-scanner.md
Templates / plugins docs/templates.md, docs/plugins.md
Multi-turn / tool efficiency / red-team docs/multi-turn-evaluation.md, docs/tool-efficiency.md, docs/redteam-generation.md
Changelog CHANGELOG.md
Contributing CONTRIBUTING.md
v0.3.0 release GitHub release

Status

WIP · Alpha · main at 0.5.0 · PyPI latest published 0.3.0. Useful today for local eval loops, golden suites, the Release Desk, and CI experiments. APIs and schemas can still move — pin a commit or release tag if you depend on behavior. Not a hosted observability replacement.

Why I built this

I kept hitting the same gap: agent quality lived in manual spot-checks and chat paste, while the rest of the stack had real CI. Plausible answers hid wrong tools, flaky trajectories, and regressions that only showed up after ship. Then the same production failure would return because nothing turned it into a permanent test.

I built AgentEval to make agent regression gates and failure memory as boring as unit tests — CLI-first, local-first, human approval before anything blocks merge — starting from a builder’s harness rather than a product pitch.

License

MIT. See LICENSE.

Metadata

Release files for nishanttyagi-agenteval 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nishanttyagi-agenteval 0.5.0
File Size Uploaded
nishanttyagi_agenteval-0.5.0.tar.gz 437.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nishanttyagi-agenteval 0.5.0
File Interpreter ABI Platform
nishanttyagi_agenteval-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 748.7 kB

Release files / nishanttyagi_agenteval-0.5.0.tar.gz

Download URL nishanttyagi_agenteval-0.5.0.tar.gz
Size 437.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0a38b21a9883dbca873ae6ba697ebdd9bde2c0a5e239d21c40e0c31983513dd0
BLAKE2b-256 checksum
How to use checksums
7107ecbfa0ebcdfa9b5bfe916bc7acebdf8781aa767ca17e8c92412c11b34508
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / nishanttyagi_agenteval-0.5.0-py3-none-any.whl

Download URL nishanttyagi_agenteval-0.5.0-py3-none-any.whl
Size 311.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dc6db82a35ebe2831e17c82425e99f848330cc2028d7d173e2cbe14f04c6f608
BLAKE2b-256 checksum
How to use checksums
32f707887935a7f5f538b65a6b814a1ed11f6ecccd28b98e39fb512bc025b419
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page