Agent Evaluator
Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.
It asks not just "does the agent work well?" but "is the agent ready for production?" One decorator line auto-recognizes 24 frameworks (LangChain, CrewAI, AutoGen, …) and measures 58 metrics (25 Native Trackers + 33 Harness Config) without touching your agent code — then aggregates them into 7 Gate pass/warn/fail judgments, a root-cause diagnosis engine for regressions, and statistically valid A/B testing.
pip install agent-evaluator
from agent_evaluator import QuickEval
eval = QuickEval("results/")
@eval.qa
def my_agent(question: str, ground_truth: str = "") -> str:
return llm.invoke(question) # your agent code — unchanged
my_agent("What is the capital of South Korea?", ground_truth="Seoul")
eval.save() # results/quickeval.json + .html
eval.gate(tcr=85, accuracy=70, hallucination=5) # CI/CD gate — sys.exit(1) if unmet
The 7 Harness Gates
| Gate | Area | Judgment Criteria | Harness Config (count) |
|---|---|---|---|
| A 🟢 | Goal Achievement | Instruction compliance · goal alignment · plan consistency · context retention | InstructionConfig · GoalAlignmentConfig · PlanConfig · SubtaskConfig · ContextRetentionConfig · KnowledgeRetentionConfig (6) |
| B 🔵 | Behavioral Integrity | Loop detection · scope deviation · tool safety · state consistency · deadlock detection | LoopDetectionConfig · ScopeConfig · ToolParameterSafetyConfig · ContextWindowConfig · StateConsistencyConfig · DeadlockConfig (6) |
| C 🟡 | Reliability | Reproducibility · error recovery rate · hallucination faithfulness · quality floor · idempotency | ReproducibilityConfig · FaultToleranceConfig · GracefulDegradationConfig · RetryConsistencyConfig · IdempotencyConfig (5) |
| D 🔵 | Performance Contract | SLA compliance · token efficiency · TTFT variability · cost predictability | SLAConfig · EfficiencyConfig · ResourceBudgetConfig · TTFTVariabilityConfig · CostPredictabilityConfig (5) |
| E 🔴 | Security Boundary | Threat severity · compliance · threat response behavior | ThreatSeverityConfig · ComplianceConfig · ThreatResponseConfig (3) |
| F 🟣 | Multi-Agent Coordination | Inter-agent consensus · information propagation accuracy · role compliance · conflict resolution | ConsensusConfig · PropagationConfig · AgentRoleConfig · ConflictResolutionConfig (4) |
| G 🩵 | Observability | Reasoning explainability · internal state tracking · error diagnosis · latency attribution | ExplainabilityConfig · ObservabilityConfig · ErrorDiagnosisConfig · LatencyAttributionConfig (4) |
Pass any of the 33 Configs above as @agent_eval/@batch_eval/@conversation_eval parameters and
PerformanceMonitor auto-aggregates each Gate's pass/warn/fail from the underlying trackers — no
separate scoring pass needed.
@agent_eval(monitor, task_type="qa",
instructions=InstructionConfig(required_keywords=["Seoul"], fail_on_violation=True), # Gate A
loop_detection=LoopDetectionConfig(consecutive_repeat_threshold=6), # Gate B
sla=SLAConfig(p95_ms=3000), # Gate D
)
def my_agent(question: str, ground_truth: str = "") -> str: ...
Full Gate reference: Docs/05_QUALITY_GATE.md · Runnable walkthrough:
Evaluator_Examples/ch03_harness_basics.py
What's Inside
- 3 decorator types —
@agent_eval(1 call → 1 result),@batch_eval(1 call → N results),@conversation_eval(N calls → 1 multi-turn result). All non-invasive: your function's signature, return value, and exceptions are untouched. →Docs/01_GETTING_STARTED.md - 24 framework adapters —
framework="langchain"/"crewai"/"anthropic"/"openai"/… auto-extractstool_calls/chain_steps/tokens_usedfrom the framework's native response object (duck typing — works without agent-evaluator importing the framework itself). →Docs/03_INTEGRATION_GUIDE.md - 58 metrics — 25 Native Trackers (accuracy, hallucination, latency, tool efficiency, 5 security
trackers, …) + the 33 Harness Configs above. →
Docs/02_METRICS_GUIDE.mdor the in-app SDK Reference (agent-eval dashboard→/sdk-docs) - Self-contained HTML report —
save_to_file()(andagent-eval gate --html-out) write one.htmlnext to the JSON — no server, no external JS. It is laid out in 3 tiers: judgment (a one-line deployment-readiness verdict + a HIGH/MEDIUM/LOW confidence badge + Next actions 1·2·3 + a Path to Green — quantified gap to each failing gate, impact-ordered), iteration (a lifecycle-phase strip — analysis → operations — and a proof panel), and evidence (collapsible groups: per-gate Score Breakdown, worst failure cases each with a tool-call trajectory waterfall, the eval-set design contract, stats, versioning, governance). Pass a baseline and it adds the regressed/new/fixed failure-set diff plus prompt/config change attribution.agent-eval gate --html-summaryprints a short Markdown block (verdict + path + the one next command) for a PR body. →Docs/13_OUTPUTS.md - CI/CD quality gating —
agent-eval gate result.json --tcr 85 --accuracy 70, plus baseline regression detection, per-version baselines, golden-set regression gating,--require-spec-coverage(exit 4 if a declared requirement has no golden case), and--hold-on-undecided(exit 75 — "hold for human" when the verdict is statistically borderline; never overrides a real fail). →Docs/05_QUALITY_GATE.md - Root-cause diagnosis (RCA) —
agent-eval diagnose/agent_evaluator.rca.diagnose()automates detect → attribute → cross-reference for a Gate regression, and links Gate F findings to the MAST failure-mode taxonomy (Cemri et al., 2025). Candidates and evidence only — HOTL, never a verdict. →Evaluator_Examples/ch28_rca_diagnosis.py - Statistically valid A/B testing —
agent-eval abtestauto-selects Welch's t-test (2 files), mSPRT always-valid inference (--sequential, safe under repeated peeking), or N-way + FDR correction (3+ files). →Evaluator_Examples/ch29_sequential_ab_test.py - Machine-readable insight layer + closed improvement loop — every result JSON carries an
extra_metrics.insightsobject (~65 keys: deployment-readiness verdict +decision_ready, Path to Green, failure clustering, paste-ready fix snippets, per-(gate, change)track record,lifecycle_phase,nondeterminism_repeat,eval_set_delta,spec_coverage,deploy_decision).agent-eval targetpins your project SLOs,agent-eval benchmarkan external reference distribution, andagent-eval experiment/agent-eval improveregister a hypothesis → apply → re-verify loop (improve apply-verifyruns the apply in a throw-awaygit worktreeand never merges). Schema-validated, never raises. →Docs/13_OUTPUTS.md - Development-support layer (all opt-in, no default behavior change) —
create_taskresult(covers=[req_id])ties golden cases to requirements; a deploy-decision ledger (gate --decision-log+agent-eval decisions record) records who accepted / held / overrode a gate run and why;FaultInjectionConfiginjects seeded tool failures / latency into Gate C/D scoring without ever touching the blocking path;run_repeated()folds a K-run verdict-stability summary intoinsights.nondeterminism_repeat; aStreamingEvaluator(golden_candidate_sink=)queue routes production errors / low-confidence answers toagent-eval dataset review-candidatesfor human promotion. →Docs/05_QUALITY_GATE.md - Real-time guardrail — two reference stacks, plus a host-less mode — the same
LiveGuardrailengine blocks a single tool call before it runs (Gate B/E), wired into either AOO (Agent-Evaluator + Ollama + OpenCode — fully local, no cloud model) viaagent-eval opencode install, or AC (Agent-Evaluator + Claude Code — native CLI hooks) viaagent-eval claude install; identical verdict logic, the difference is the process model (a resident subprocess vs. per-call replay). No host at all?tool_guard()+live_guardrail_session()wrap any Python agent loop directly, and (1.1.0)audit_blocked=Trueis now the default,audit_log_path=flushes a durable record on exit even if the caller's own error handling is silent, andon_block=webhook_on_block(url)fires an out-of-band alert the instant a call is blocked. A blocked call keeps a redacted excerpt of its command, soagent-eval {claude,opencode} violations(with no query — 1.1.0 — browses recent history; a keyword narrows it) /blocked-detail <task_id>(and thelist_violations/show_violationMCP tools) surface what was blocked;doctornow also reports the audit DB's row count proactively, and the HTML report shows blocked attempts directly (insights.blocked_attempts_audit) even though they never move Gate B/E scores. →Docs/06_LIVEGUARDRAIL.md·Docs/07_CLAUDE_CODE_HOOKS.md·Docs/08_AOO_STACK.md·Docs/09_OPENCODE_VS_CLAUDE_CODE.md - Dashboard —
agent-eval dashboard(FastAPI): Harness Gate breakdown, File Compare with pairwise LLM Judge, anomaly/cost tracking, and a 🔧 Improve tab surfacing the RCA engine.
Positioning & Direction
Where it fits
Agent-Evaluator is a deployment-readiness judge and iteration harness — not a tracing platform, and not a metric library:
| Layer | Representative tools | Question they answer | Relation |
|---|---|---|---|
| Tracing / observability | LangSmith · Arize Phoenix · Langfuse | What happened in this run? | Consumes it — ships an OTEL exporter + a Phoenix monitor; not a competitor |
| Metric libraries | DeepEval · Ragas · promptfoo | What's the score on metric X? | Plugs them in as adapters; the 25 native trackers mean a base install needs none |
| Readiness + iteration | Agent-Evaluator | Is this ready to ship — and if not, what is the smallest set of fixes, in what order? | This layer |
Its center of gravity is where the ship / no-ship decision is made: the CI/CD gate and the analysis → design → development → verification loop before it. Streaming + OTEL cover production, but that is not the focus.
What makes it different
- Seven non-redundant Gates, not one score. Goal, behavior, reliability, performance, security, multi-agent coordination, observability — each graded pass / warn / fail. "83% accuracy" cannot hide a security regression, because they are scored separately.
- A verdict and a path, not just numbers. Every report and
insightsobject leads with ship / hold / not-ready + a HIGH/MEDIUM/LOW confidence badge, then an impact-ordered Path to Green with paste-ready@agent_evalsnippets — the smallest next step, not a dashboard to interpret. - Two closed loops, deliberately separated. A real-time guardrail blocks a dangerous tool call
before it runs (Gate B/E); a batch pass grades after. Different code, different timing — the
guardrail never scores and the scorer never blocks (
FaultInjectionConfig,--hold-on-undecided, everything respects that line). - Statistical honesty over a confident number. Wilson CI on the pass-rate;
decision_ready→ exit 75 ("hold for human") when the verdict is borderline; always-valid A/B inference (safe under repeated peeking); K-run verdict-stability; Benjamini–Hochberg correction across slices. It declines to over-claim on n = 12. - Human-on-the-loop by construction. RCA emits candidates + evidence, never a verdict.
improve apply-verifyruns the apply in a throw-awaygit worktreeand never merges or commits. A deploy-decision ledger records who accepted / held / overrode each gate run, and why. - Non-invasive and framework-agnostic. One decorator; your function's signature, return value, and exceptions are untouched. 24 framework adapters by pure duck typing — agent-evaluator never imports the framework itself.
- A local-first option. The AOO stack (Agent-Evaluator + Ollama + OpenCode) runs the entire loop with no cloud model and no token bill; AC (Agent-Evaluator + Claude Code) targets cloud models. Same verdict engine, different process model.
- Backed by a written methodology. Harness Methodology (a companion book) states the discipline — decision before code · prove, don't pass · block ≠ score · don't rebuild what already exists · size the model to the task · the human stays accountable — and this SDK is one reference implementation of it.
Direction
The trajectory from 1.0.0 has moved from "score the agent" toward "drive the whole build-and-improve loop, with the human's judgment on record" — and, most recently, toward making the real-time guardrail trustworthy even with no host process watching it:
- 1.0.0 — the machine-readable
insightslayer + thetarget/benchmark/experiment/improveCLI loop. - 1.0.5 — Harness Methodology alignment (
--hold-on-undecided,human_only_patterns, verdict-stability), a development-support framework (requirement → golden-case coverage, a deploy-decision ledger, thin fault injection, a production → golden candidate queue,improve apply-verify), and the HTML report re-cast as a methodology instrument — three tiers (judgment / iteration / evidence), a lifecycle-phase read, and--html-summaryfor a PR body. - 1.1.0 —
LiveGuardraildiscovery/durability hardening: a blocked call is durable and discoverable by default now, even for a host-less custom agent loop with no Claude Code/OpenCode bridge and no visible message (tool_guard(audit_blocked=True)default,audit_log_path=crash-safe flush,on_block=webhook_on_block(...)out-of-band alert) — plus keyword-freeviolations/list_violationsbrowsing and an HTML report section for everyone, since a blocked attempt never moves Gate B/E scores and a clean scorecard alone would hide it.
Explicit non-goals — the boundary is a deliberate design choice, not a missing feature: it does not author or parse specs (EARS), does not provide a sandbox (E2B / Firecracker — that is the team's infrastructure), does not enforce which model tier a task runs on (that is the runtime's), and never auto-merges or auto-deploys.
Installation
Extras are organized into 5 categories by intent — pick the one(s) that match what you're trying to do. Every category is additive and independent; combine as needed.
| # | Category | Install | What it adds |
|---|---|---|---|
| 1 | Base measurement + diagnosis | pip install agent-evaluator |
25 trackers · 33 Harness Config · 7 Gates · LLMJudge · RCA diagnosis engine (agent_evaluator.rca/ontology, no extra deps needed) · full CLI (gate/decisions/diagnose/abtest/trend/dataset/feedback/experiment/target/benchmark/improve/claims) |
| 2 | SDK — dashboard + monitoring | pip install "agent-evaluator[sdk]" |
FastAPI dashboard (serve), Phoenix/OTEL (otel), Korean RAG PDF processing (pdf+korean) — recommended for most users |
| 3 | Real-time guardrail — OpenCode/Claude Code + MCP | pip install "agent-evaluator[mcp]" |
search_violations + show_violation, recommend_fix, and ask_insights stdio MCP servers so OpenCode, Claude Code (or another MCP client) can call them as tools during a live session — the underlying functions already work without this (recommend_fix's knowledge is used directly by agent-eval diagnose); this only wires up the MCP protocol layer |
| 4 | Your agent's framework | pip install "agent-evaluator[langchain]" (or [crewai]/[autogen]/[dspy]/[pydanticai]/[eval]) |
Packages your agent code imports directly — agent-evaluator itself works without them via duck typing; install only what you actually use |
| 5 | Examples / full / dev | pip install "agent-evaluator[examples]" |
Everything needed to run Evaluator_Examples/ with real (non-mock) DeepEval/Ragas/dashboard/Phoenix output. [full] = category 4's frameworks all at once (⚠️ 10+ min install); [dev] = contributor tooling |
Single-feature extras that don't fit the 5 categories above: [export] (dashboard Parquet/Excel),
[wandb], [mlflow]. Full package-by-package breakdown: pyproject.toml.
CLI Commands
| Command | Description |
|---|---|
agent-eval init / check |
Interactive API key setup / configuration status |
agent-eval dashboard [dir] |
FastAPI dashboard web server |
agent-eval gate <result.json> |
CI/CD quality gating (--require-spec-coverage exit 4 · --hold-on-undecided exit 75 · --decision-log · --html-out / --html-summary) |
agent-eval decisions list|record |
Deploy-decision ledger — record who accepted / held / overrode / rejected a gate run, and why |
agent-eval diagnose <result.json> |
Root-cause diagnosis for a Gate regression |
agent-eval abtest <files...> |
Statistical A/B / N-way comparison |
agent-eval trend <dir> |
Regression detection across sequential results |
agent-eval dataset build|promote|health|review-candidates |
Golden-dataset extraction / HITL promotion / coverage health / production-candidate review queue |
agent-eval feedback export-preferences |
Export A/B preference rows (pairwise-judge / annotation / contrast-pair) to JSONL — export only |
agent-eval target set|show|clear |
Pin project SLOs (.aoo/targets.json) — used by gate and the report's "below target" lines |
agent-eval benchmark set|show|clear |
Pin an external reference distribution (.aoo/reference.json) for percentile + gap-to-frontier |
agent-eval experiment register|list|score |
Register a Gate/field hypothesis, score predicted vs actual |
agent-eval improve plan|start|verify|patch|apply-verify |
Closed loop: proposal → experiment → re-verify → outcome log (apply-verify applies in an isolated git worktree, never merges) |
agent-eval monitor |
Arize Phoenix + OTEL real-time monitoring |
agent-eval opencode / claude install|upgrade|doctor|test-config|uninstall |
Install & manage the LiveGuardrail OpenCode plugin / Claude Code CLI hooks (test-config asserts the resolved guardrail config against a case file) |
agent-eval opencode / claude violations|blocked-detail |
Search (or, with no query, browse most-recent-first) past Gate B/E blocks; show the exact blocked command for a session |
agent-eval claims add|list|release|audit |
Team scope-claim management (.aoo/claims.jsonl) |
Examples
32 standalone, book-chapter-based files in Evaluator_Examples/ (ch01–ch32),
covering everything from a first evaluation to the full RCA/A/B-testing improvement loop:
pip install "agent-evaluator[examples]"
cd Evaluator_Examples && python ch01_first_eval.py # ... through ch32_ollama_realtime.py
Project Structure
agent_evaluator/
├── decorators.py # agent_eval · batch_eval · conversation_eval · QuickEval facade (quick_eval.py)
├── gates/ # Gate A–G scoring (gate_a_goal/ … gate_g_observability/) + LiveGuardrail
├── core/trackers/ # 25 Native Trackers (Layer 1 foundation · Layer 2 agentic/security) + monitor.py
├── rca/ # diagnose() — Gate-regression root-cause diagnosis + improvement/experiment logs
├── ontology/ # GATE_GUIDANCE · NATIVE_METRIC_RULES · MAST + single-agent failure taxonomies
├── reporting/ # insights.py (build_insights) + comprehensive_report.py (self-contained HTML)
├── integrations/ # LLMJudge · DeepEval/Ragas adapters · MCP servers · live-guardrail bridges
├── serve/ # FastAPI dashboard ([sdk] extra)
└── cli/ # agent-eval CLI (init, check, gate, decisions, diagnose, abtest, trend, dataset,
# feedback, experiment, target, benchmark, improve, claims, monitor, opencode, claude)
Evaluator_Examples/ # 32 example files (ch01–ch32)
tests/ # 5,150+ test functions
Changelog
- v1.1.0 (2026-09-11) — LiveGuardrail discovery/durability hardening.
tool_guard(audit_blocked=True)now default;on_block=/webhook_on_block()out-of-band block alerts; keyword-freelist_violations()/violationsbrowsing;blocked_attempt_capture.max_chars240→500;insights.blocked_attempts_auditHTML section. Additive/opt-in beyond the default change. - v1.0.6 (2026-09-10) — Maintenance:
decorators.pysplit intoframework_adapters.py+_eval_shared.py(re-exported, no API change);--helpnow lists all 18 subcommands. - v1.0.5 (2026-09-09) — Harness Methodology alignment + dev-support framework + 3-tier HTML report.
--hold-on-undecided,--require-spec-coverage, deploy-decision ledger,run_repeated(),FaultInjectionConfig,dataset review-candidates,improve apply-verify. All opt-in. - v1.0.4 (2026-09-08) — Blocked calls keep a redacted command excerpt; new
show_violationMCP tool +{claude,opencode} violations/blocked-detailCLI. - v1.0.3 (2026-09-08) — Fixed
search_violationsMCP to point at the Claude Code DB (was defaulting to OpenCode's). - v1.0.2 (2026-09-04) — Phoenix/OTEL: version-scoped pin, span/trace/session annotations, ASCII console output.
- v1.0.1 (2026-09-03) — Report-generation hardening (malformed JSON no longer crashes), dashboard/report value parity, English-only runtime output.
- v1.0.0 (2026-08-31) — General Availability: machine-readable
insightslayer (~62 keys) +target/benchmark/experiment/improveCLI loop.
Full history (incl. the 1.0.0-rc.1–rc4 series): CHANGELOG.md.
Documentation
Docs/01_GETTING_STARTED.md |
Decorators, QuickEval, first evaluation |
Docs/02_METRICS_GUIDE.md |
All 58 metrics — formulas, activation conditions |
Docs/03_INTEGRATION_GUIDE.md |
24 framework adapters, auto-detection |
Docs/04_DATA_GUIDE.md |
Golden datasets, evaluation data design |
Docs/05_QUALITY_GATE.md |
Harness Gates, CI/CD gating, RCA diagnosis |
Docs/06_LIVEGUARDRAIL.md |
LiveGuardrail subsystem reference — all usage modes + v1.1.0 discovery/durability hardening |
Docs/07_CLAUDE_CODE_HOOKS.md |
AC stack (Agent-Evaluator + Claude Code) — the same guardrail via native Claude Code CLI hooks |
Docs/08_AOO_STACK.md |
AOO stack (Agent-Evaluator + Ollama + OpenCode) — the fully-local real-time-guardrail reference integration |
Docs/09_OPENCODE_VS_CLAUDE_CODE.md |
AOO vs AC — detailed side-by-side comparison |
Docs/10_OBSERVABILITY.md |
Dashboard, alerts, anomaly detection |
Docs/11_OTEL_DATA_REFERENCE.md |
Every span, attribute, metric & Phoenix annotation sent over OpenTelemetry |
Docs/12_OPERATIONS.md |
Install variants, Docker, per-environment config, performance tuning, troubleshooting |
Docs/13_OUTPUTS.md |
Result JSON · HTML reports · CLI · dashboard · AI-runtime output system |
Docs/14_API_REFERENCE.md |
Full public API reference |
Docs/15_CTX_SESSION_SEARCH.md |
Optional cross-session search workflows (ctx) — no Agent-Evaluator dependency |
CHANGELOG.md |
Version history |
Also available in-app once the dashboard is running: agent-eval dashboard → SDK Reference
(/sdk-docs) and REST API (/api/docs).
Development
git clone https://github.com/bullpeng72/Agent-Evaluator.git
cd Agent-Evaluator
pip install -e ".[dev]"
pytest # run tests
ruff check agent_evaluator/ # lint
mypy agent_evaluator/ # type check
License
MIT — see LICENSE.
Release files for agent-evaluator 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_evaluator-1.1.0.tar.gz | 1.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_evaluator-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.8 MB
Release files / agent_evaluator-1.1.0.tar.gz
| Download URL | agent_evaluator-1.1.0.tar.gz |
|---|---|
| Size | 1.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
416bba55b96bd9d56cfa2016414551bd7cca3cbe8175a3abf65a5a8c5994a4ad
|
|
BLAKE2b-256 checksum How to use checksums |
e56ccb4b140e8cc472cbe0ab70b14925a713cbaf97f92e715435fc08ef207866
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.15
|
Release files / agent_evaluator-1.1.0-py3-none-any.whl
| Download URL | agent_evaluator-1.1.0-py3-none-any.whl |
|---|---|
| Size | 1.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
800f535eb86535477d242e9a55ab67f22818f7a599eef8ca3e06b35e57771de1
|
|
BLAKE2b-256 checksum How to use checksums |
f70f1ef482f948bceb7c432fcbe27c257927830d01fde0c681f5adafa14db7c7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.15
|