AgentEval
CI for AI agents — turn flaky agent behavior and production failures into tests that fail the PR.
Static demo · Streamlit dashboard · PyPI: nishanttyagi-agenteval · CLI: agenteval
The problem
You change a prompt, model, or tool. The agent still “answers.” Unit tests still pass. A week later someone finds a wrong refund, a hallucinated fact, or a tool the agent never should have called.
Normal tests prove the code ran. They do not prove the agent still behaves.
In simple terms
Imagine a support agent gets:
“Cancel order 4821 and refund the customer.”
A green unit test only means something came back. AgentEval checks the behavior:
| Question | What gets checked |
|---|---|
| Did it solve the request? | Final answer vs expected outcome |
| Did it invent anything? | Unsupported claims vs ground truth |
| Did it use the right tools? | e.g. lookup_order + issue_refund, not a random tool |
| Did it take the right steps? | Trajectory vs expected sequence |
| Is it stable? | Optional repeats → stable / flaky / unstable |
| Did this change make it worse? | Current run vs a versioned baseline (CI exit code) |
| Has this failed in production before? | Failure Memory → approved golden case → CI again |
If the agent says “refunded” without calling the refund tool, that can look fine to a human skim. AgentEval records it as a failure and can block the PR.
What I built
Four outcomes, not a platform pitch:
- A regression gate for agents — YAML golden suites, scored runs, compare to a git-trackable baseline, fail CI when quality drops.
- Evidence, not one score — correctness, hallucination, tool-call accuracy, latency/cost, trajectory, flakiness, plus optional RAG checks and a SQL safety scanner.
- Failure Memory — capture → redact → cluster → replay → minimize → human approve → golden YAML → CI, so the same production failure is harder to ship twice.
- One CLI across frameworks — adapters for CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, and custom agents; composite GitHub Action; templates (including an Indic-language pack).
A few numbers
Evidence from this repo on main — no invented percentages.
| Fact | Value |
|---|---|
| Automated tests (this checkout, optional extras absent) | 1152 passed, 3 skipped (Docker, FastAPI, KarmaSakshi) |
| Failure Memory suite | 59 passed |
Test modules (tests/**/test_*.py) |
105 |
Package version on main |
0.5.0 |
| Latest published PyPI release | 0.3.0 (PyPI) |
| Top-level CLI commands | 19 (run, compare, report, generate, generate-adversarial, import, generate-cases, init, compare-models, trace, diff, calibrate, audit-log, serve, gate, plugins, templates, sql, memory) |
agenteval memory subcommands |
19 |
| Bundled templates | 4 — coding-agent (7), customer-support (7), rag-assistant (7), indic-agent (34 cases: 28 offline / 6 opt-in LLM-judge) |
| Framework adapters | CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, KarmaSakshi bridge (+ custom) |
| Optional extras | dev, crewai, autogen, openai-agents, karmasakshi |
| Python | 3.10+ |
| License | MIT |
| Classifier | Alpha (Development Status :: 3 - Alpha) |
| CI workflows in-repo | eval.yml, failure-memory.yml, karmasakshi-bridge.yml, action-smoke.yml, docker.yml, publish.yml, landing-page.yml |
| Offline demos | examples/mock_agent, examples/failure_memory_demo(_v21), examples/indic_mock_agent, examples/karmasakshi_bridge (no API key) |
Why this matters
If you ship agent changes, you want (1) a hard gate when known behavior regresses, and (2) a memory of real failures that does not reset every run. AgentEval is built for that workflow: golden YAML and baselines live in git; Failure Memory stays local-first with human approval before anything blocks CI; demos run without network or provider keys.
How it works
Simple flow
- Write expected behavior as YAML golden cases (or scaffold with
agenteval init). - Run the agent through an adapter:
agenteval run. - Compare to a baseline:
agenteval compare(CI fails on regression). - Optionally: ingest a production failure into Failure Memory → redact → replay → minimize → approve → export golden → run with
--production-cases. - Inspect evidence: JSON/HTML report, Streamlit dashboard, or
agenteval trace/diff.
Optional architecture (for engineers)
Agent adapter + golden YAML → runner / evaluators → JSON run + HTML report
→ baseline compare / CI gate
Production traces → redact → Failure Memory (SQLite)
→ cluster / replay / minimize
→ human approve → golden YAML → same CI gate
Implemented in-repo: adapters/, core/, evaluators/, failure_memory/, CLI entry agenteval → agenteval.cli:main, optional Streamlit dashboard, composite Action (action.yml).
Install
Python 3.10+.
From PyPI (latest published release is 0.3.0):
pip install nishanttyagi-agenteval==0.3.0
agenteval --version
agenteval --help
From this repo (current main is 0.5.0):
git clone https://github.com/nishanttyagi28/agenteval.git
cd agenteval
python -m pip install -e ".[dev]"
agenteval --version # expect 0.5.0 on main
python -m pytest -q
Framework extras (optional): pip install "nishanttyagi-agenteval[crewai]" (same for autogen, openai-agents, karmasakshi).
Quick start
Offline mock agent (no provider)
agenteval run --agent mock_agent --registry examples/mock_agent/agents.yaml
agenteval compare --agent mock_agent --registry examples/mock_agent/agents.yaml
Failure Memory demo (zero network, no API key)
python examples/failure_memory_demo_v21/run_demo.py
Capture → redact → cluster → replay → minimize → export golden → fail broken agent / pass fixed agent → coverage check. See examples/failure_memory_demo_v21/ and docs/failure-memory.md.
KarmaSakshi bridge demo (offline seal → witness → score)
pip install -e ".[dev,karmasakshi]"
python examples/karmasakshi_bridge/run_demo.py
Approved ₹1500→Priya: correct attempt passes; ₹1501 or wrong payee is blocked by KarmaSakshi and recorded as an AgentEval failure. See examples/karmasakshi_bridge/ and docs/karmasakshi-bridge.md.
Scaffold a project
agenteval init
Release Desk (v0.5.0)
Turn a production miss into a CI test without leaving the browser. No API key.
agenteval gate --local
Open the URL it prints (default http://127.0.0.1:8741). Load the refund sample, or POST an incident:
curl -s -X POST http://127.0.0.1:8741/api/gate/ingest \
-H 'Content-Type: application/json' \
-d '{"incident":{"agent":"refund-desk","prompt":"Refund order 4821","output":"Cancelled. No refund.","tools_called":["cancel_order"],"expected_tools":["lookup_order","issue_refund"]}}'
Ship writes .agenteval/production-regressions.yaml. The next run that still
does this fails:
agenteval run --agent my_agent --production-cases .agenteval/production-regressions.yaml
The desk is local-only (pass --local). It redacts secret-shaped strings before
storage and will not bind a public address unless you pass --allow-remote.
See docs/release-desk.md.
Open-ended llm_judge cases no longer require the data-analyst sibling repo.
With no key, the deterministic offline judge scores them. Set
AGENTEVAL_JUDGE_PROVIDER to openai, groq, anthropic, or legacy.
Indic-language pack (v0.4.0 on main)
pip install -e examples/plugins/agenteval-indic-evaluators
agenteval run --agent indic_mock_agent \
--registry examples/indic_mock_agent/agents.yaml \
--tag core --no-llm-judge
34 cases (28 deterministic offline; 6 refusal/safety need an LLM judge). The mock demo expects deliberate failures so each checker is shown catching something — see examples/indic_mock_agent/README.md.
CLI
agenteval run | compare | report | generate | generate-adversarial | import | generate-cases
agenteval init | compare-models | trace | diff | calibrate | audit-log | serve
agenteval plugins | templates | sql | memory | gate
| Command | Role |
|---|---|
run / compare |
Golden suite + baseline regression gate |
report |
Self-contained HTML report |
init |
Scaffold registry, sample cases, CI workflow |
trace / diff |
Step evidence and trajectory diff |
generate / generate-adversarial |
Reviewable adversarial / red-team candidates (not auto-blocking) |
memory … |
Failure Memory loop (see below) |
gate |
Local Release Desk: ingest a production failure, approve it, write the CI case |
sql scan |
SQL agent structural safety scan |
templates / plugins |
Bundled starters and evaluator entry points |
Failure Memory CLI
agenteval memory init
agenteval memory ingest traces.jsonl
agenteval memory cluster
agenteval memory list
agenteval memory review <candidate_id> approve --correctness-type contains --ground-truth "…"
agenteval memory replay <candidate_id>
agenteval memory minimize <candidate_id>
agenteval memory approve-minimization <minimization_id>
agenteval memory export-minimized <minimization_id>
agenteval memory coverage
agenteval run --agent my_agent --production-cases .agenteval/production-regressions.yaml
Default DB: .agenteval/failure-memory.db (--db or AGENTEVAL_FAILURE_MEMORY_DB).
Production failure → redact → ingest → cluster → replay → minimize
→ human approve → golden YAML → CI gate → recurrence / coverage
What AgentEval evaluates
| Capability | What you get |
|---|---|
| YAML golden suites | Versioned prompts, expectations, tools, tags |
| Correctness | Exact, contains, numeric, numeric-table, optional LLM judge |
| Hallucination | Unsupported claims vs ground truth |
| Tool-call accuracy | Required tools: precision / recall / F1 |
| Latency & cost | p50/p95 and suite cost; opt-in budget gates |
| Trajectory | Expected steps (LCS F1); agenteval diff |
| Flakiness | Optional repeats; stable / flaky / unstable |
| Baseline regression | Compare to versioned baseline; CI exit codes |
| RAG mode | Context relevance, faithfulness, citation checks |
| SQL safety scanner | Structural / policy-oriented checks (agenteval sql) |
| Failure Memory | Production failure → approved golden regression |
| Adapters | CrewAI, AutoGen, OpenAI Agents SDK, LangGraph, KarmaSakshi bridge, custom |
| GitHub Action | Composite action + in-repo regression workflow |
| Reports | Streamlit, HTML, local read-only API (serve) |
Framework registry example
version: 1
agents:
mock_agent:
adapter: examples.mock_agent.adapter:MockAgentAdapter
cases: examples/mock_agent/cases.yaml
enabled: true
Composite Action for consumer repos:
- uses: nishanttyagi28/agenteval@v0.3.0
with:
agent: my_agent
config-file: agents.yaml
cases-file: tests/golden/cases.yaml
baseline-file: baselines/my_agent.json
Pin a release tag you trust. main moves; PyPI 0.5.0 is not published yet. The latest published release remains 0.3.0.
Security and privacy defaults
| Default | Meaning |
|---|---|
| Content capture off | Prompts/outputs not stored unless you opt in |
| Redaction before disk | Applied before SQLite / JSONL persistence |
| Best-effort DLP | Common secret/PII patterns only — not a universal guarantee |
| Human approval | Required before golden promotion; nothing auto-enters blocking CI |
| Local-first SQLite | Default under .agenteval/; no hosted multi-tenant control plane |
| No telemetry by default | Failure Memory does not phone home |
| Keep secrets out of git | Do not commit DBs, raw traces, or sensitive generated suites |
Limitations and non-goals
- Local-first / single-user Failure Memory — not a hosted multi-tenant product
- Not an OpenTelemetry collector (optional OTel-shaped JSON helpers only)
- No mandatory vector DB or embeddings for clustering
- No mandatory LLM judge for classification / clustering
- Redaction will miss arbitrary unlabeled secrets
- Live agent eval still needs your runtime and any provider keys you choose
- Cost may fall back to estimates when providers omit usage
- Adversarial / generated cases stay candidates until a human promotes them
- Not claiming v1 stability — see
docs/v1-readiness.md
Documentation
| Resource | Link |
|---|---|
| Release Desk | docs/release-desk.md |
| Failure Memory | docs/failure-memory.md |
| KarmaSakshi bridge | docs/karmasakshi-bridge.md |
| Compatibility | docs/compatibility.md |
| SQL scanner | docs/sql-scanner.md |
| Templates / plugins | docs/templates.md, docs/plugins.md |
| Multi-turn / tool efficiency / red-team | docs/multi-turn-evaluation.md, docs/tool-efficiency.md, docs/redteam-generation.md |
| Changelog | CHANGELOG.md |
| Contributing | CONTRIBUTING.md |
| v0.3.0 release | GitHub release |
Status
WIP · Alpha · main at 0.5.0 · PyPI latest published 0.3.0. Useful today for local eval loops, golden suites, the Release Desk, and CI experiments. APIs and schemas can still move — pin a commit or release tag if you depend on behavior. Not a hosted observability replacement.
Why I built this
I kept hitting the same gap: agent quality lived in manual spot-checks and chat paste, while the rest of the stack had real CI. Plausible answers hid wrong tools, flaky trajectories, and regressions that only showed up after ship. Then the same production failure would return because nothing turned it into a permanent test.
I built AgentEval to make agent regression gates and failure memory as boring as unit tests — CLI-first, local-first, human approval before anything blocks merge — starting from a builder’s harness rather than a product pitch.
License
MIT. See LICENSE.
Metadata
Release files for nishanttyagi-agenteval 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nishanttyagi_agenteval-0.5.0.tar.gz | 437.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nishanttyagi_agenteval-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 748.7 kB
Release files / nishanttyagi_agenteval-0.5.0.tar.gz
| Download URL | nishanttyagi_agenteval-0.5.0.tar.gz |
|---|---|
| Size | 437.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0a38b21a9883dbca873ae6ba697ebdd9bde2c0a5e239d21c40e0c31983513dd0
|
|
BLAKE2b-256 checksum How to use checksums |
7107ecbfa0ebcdfa9b5bfe916bc7acebdf8781aa767ca17e8c92412c11b34508
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / nishanttyagi_agenteval-0.5.0-py3-none-any.whl
| Download URL | nishanttyagi_agenteval-0.5.0-py3-none-any.whl |
|---|---|
| Size | 311.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
dc6db82a35ebe2831e17c82425e99f848330cc2028d7d173e2cbe14f04c6f608
|
|
BLAKE2b-256 checksum How to use checksums |
32f707887935a7f5f538b65a6b814a1ed11f6ecccd28b98e39fb512bc025b419
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log