A framework-agnostic evaluation harness for tool-using AI agents. Define scenario packs, run agents consistently, score outcomes and trajectories, and gate releases with evidence.
Project description
Agent Eval Forge
Stop unsafe agent changes from shipping.
pytest for agents — a framework-agnostic evaluation harness that catches
regressions before you ship: correctness, tool discipline, safety boundaries,
cost, and latency.
Why
You change a prompt, swap a model, or add a new tool. The demo looks great. Three days later, a customer hits a regression the demo never covered.
Agent teams change prompts, models, tools, and orchestration logic constantly — but most still judge progress by eyeballing a handful of examples. That works for demos. It fails for production. The final answer isn't the only thing that matters: the path taken, the tools selected, the arguments passed, the cost incurred, and the safety boundaries respected all matter.
EvalForge turns release decisions from anecdotes into evidence.
"Did the agent actually get better, or did it just change?"
What It Is
A release-discipline product for tool-using AI agents. Local-first, CI-second. Strong default rubrics out of the box, all overridable. Built for one job: deciding whether an agent change is safe to ship.
| Capability | Description |
|---|---|
| Scenario packs | YAML-defined evaluation scenarios with inputs, tools, expected behaviors, and scoring rules |
| Trajectory scoring | Score the agent's path — not just the final answer. Tool selection, argument quality, step efficiency |
| Safety gating | Catch disallowed tool use, policy violations, and data boundary breaches before they ship |
| Regression detection | Compare candidate versions against explicit golden baselines at scenario, family, and pack level |
| Framework-agnostic | Adapters for subprocess, Python import, and HTTP. Proven targets: LangGraph, PydanticAI |
| CI-native | Runs locally for developer decisions, in CI for enforcement. Structured exit codes for safety violations |
| Deterministic + LLM judge | Cheap deterministic checks first, semantic LLM-as-judge only when needed |
| Security model | Sandbox mode, trust policies, audit trail, API key sanitization |
| 20 launch scenarios | Production-grade scenarios across 10 families, plus 8 security scenarios |
| External benchmarks | SWE-bench and WebArena connectors |
Product philosophy
- Safety > Correctness > Efficiency — safety regressions fail by default
- Explicit golden baselines — compare against what you accepted, not what happened to run last
- Local-first, CI-second — catch issues before merge, enforce in CI
- Strong default rubrics — credible pass/fail behavior out of the box
- Framework-agnostic by adapter — not claimed, proven with tested integrations
What It Is Not
| Not this | Why |
|---|---|
| A generic LLM eval framework | EvalForge goes deeper on scenario packs + trajectory/regression for agents |
| A hosted observability platform | EvalForge is for evaluation and regression discipline, not production tracing ownership |
| An auto-prompt optimizer | EvalForge diagnoses and measures — it does not mutate systems automatically in v0.1 |
| A benchmark leaderboard | The focus is practical agent evaluation, not infinite leaderboard collection |
Quickstart
# Install
pip install agent-eval-forge
# With framework adapters and judge backends
pip install agent-eval-forge[langgraph,pydanticai,judge]
# Run a scenario pack against your agent
evalforge run \
--pack scenarios/core-launch.yaml \
--agent python:my_agent.py \
--baseline v1.0.0 \
--judge openai:gpt-4o-mini
# Save a baseline once you're happy
evalforge baseline save --name v1.0.0 --run run-20260728-001
# Gate in CI
evalforge run --pack scenarios/core-launch.yaml --agent python:my_agent.py --ci
See the User Guide for a complete walkthrough — first eval in 5 minutes, all commands, and 12 real-world gotchas.
Launch Pack (v0.1)
20 scenarios across 10 families, plus 8 security scenarios:
| # | Family | What it tests |
|---|---|---|
| 1 | Single-Tool Factual Retrieval | Bounded lookup with the right tool |
| 2 | Multi-Tool Retrieval Synthesis | Combining evidence from multiple sources |
| 3 | Structured JSON Extraction | Transforming messy input into valid JSON |
| 4 | Tool Argument Precision | Right tool, wrong arguments |
| 5 | Tool Avoidance | Not using tools when none are needed |
| 6 | Disallowed Tool Refusal | Respecting tool policy boundaries |
| 7 | Ambiguous Request Clarification | Asking before acting under ambiguity |
| 8 | Budget-Constrained Completion | Trading off completeness and efficiency |
| 9 | Graceful Timeout / Failure Recovery | Recovering cleanly from failing dependencies |
| 10 | Coding-Agent Regression | Diff review, test failure classification |
| + | Security Scenarios | Prompt injection, exfiltration, SSRF, sandbox escape |
All scenarios ship in scenarios/core-launch.yaml and scenarios/security-launch.yaml.
Framework Adapters
| Adapter | Status | Agent spec |
|---|---|---|
| Subprocess | v0.1 | subprocess:./agent.py |
| Python Import | v0.1 | python:my_pkg.agent:run |
| HTTP | v0.1 | http:http://localhost:8000/run |
| LangGraph | v0.1 | langgraph:my_pkg.graph:build_agent |
| PydanticAI | v0.1 | pydanticai:my_pkg.agent:build_agent |
| CrewAI | v0.2+ | — |
Scoring
Two layers, cheapest first:
-
Deterministic scorers (always run, free, reproducible) — tool correctness, argument precision, schema validity, budget adherence, safety gates, grounding checks. 17 scorers built-in.
-
LLM-as-Judge scorers (run only when configured) — output correctness, task completion, synthesis quality, hallucination detection, refusal quality. 11 judge metrics built-in. Backends: OpenAI, Anthropic, Ollama, MLX (Apple Silicon).
Evaluation hierarchy: Safety failures trump everything. Correctness regressions warn by default. Efficiency regressions are informational.
See the Scoring Guide for the full metric catalog, custom scorers, and failure taxonomy.
Security
⚠️ Running without
--sandboxexposes your API keys and environment variables to the agent process. Always use--sandboxin CI. For untrusted scenario packs, also use--trust externalwhich restricts adapters to sandboxed subprocess only.
The security model includes:
- Sandbox mode — env stripping, filesystem isolation, optional Docker network isolation
- Trust policies —
trustfield on packs, adapter/tool matrix enforced at validate/run - Audit trail — per-run append-only log of policy decisions
- API key sanitization — deep-redaction in artifacts and logs
See Security Review for the full model.
Field Testing
EvalForge was validated against 19 real-world open-source agents (11 LangGraph + 8 PydanticAI) sourced from GitHub across local (MLX), cheap (gpt-4o-mini), and better (gpt-4o) tiers. The field harness exposed integration failures that mock-based testing completely missed — import-time side effects, hardcoded model names, Python version skew, and nested repo structures.
| Report | Scope |
|---|---|
| Field Test Report — 08.01.2026 | Roster scaled from 8 to 20 agents |
| Field Test Report — 08.02.2026 | 19-agent compatibility sweep, cloud tier comparison |
| Field Test Plan | Agent selection, sourcing, config schema, acceptance criteria |
| Learnings | Real-world gotchas from running 20+ agents |
Project Status
| Milestone | Status |
|---|---|
| M0 — Scaffold | ✅ |
| M1 — Core Runner | ✅ |
| M2 — Scoring Engine | ✅ |
| M3 — Comparison & Baselines | ✅ |
| M4 — Launch Scenarios 1–5 | ✅ |
| M5 — Launch Scenarios 6–10 | ✅ |
| M6 — Framework Adapters | ✅ |
| M7 — CLI & pytest | ✅ |
| M8 — CI & Polish | ✅ |
| M9 — Hardening & Security | ✅ |
| M9.5 — Integration & Field Tests | ✅ (FT5 CI job pending) |
| M10 — Example Agents & DX | ✅ |
| M11 — Scale-Up, Docker & CI | 🔄 In progress |
| M12 — Ship v0.1.0 & Launch | ⏳ |
See the WBS for the full milestone plan with GitHub issue tracking.
Documentation
| Document | Description |
|---|---|
| User Guide | Start here — installation, first eval, all commands, gotchas |
| PRD | Product requirements — the what and why, 20 canonical user journeys |
| Spec | Technical specification — architecture, data model, scoring, all scenarios |
| WBS | Work breakdown — 13 milestones, GitHub issues linked |
| Scenarios | Scenario authoring — pack anatomy, metric reference |
| Scoring | Scoring — metric categories, custom scorers, failure taxonomy |
| CI Integration | GitHub Actions, GitLab CI, Docker sandbox |
| Adapters — LangGraph | LangGraph adapter usage |
| Adapters — PydanticAI | PydanticAI adapter usage |
| Adapters — Custom | How to write a custom adapter |
| External Benchmarks | SWE-bench, WebArena connectors |
| Security Review | Security model and hardening |
| Learnings | Real-world gotchas from 20+ agents |
| Field Test Reports | Compatibility sweep across tiers |
| CHANGELOG | Release history |
| CONTRIBUTING | How to contribute |
| CODE OF CONDUCT | Community standards |
| GOVERNANCE | Project governance |
| SECURITY | Security policy and reporting |
Community
License
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agent_eval_forge-0.1.0.tar.gz.
File metadata
- Download URL: agent_eval_forge-0.1.0.tar.gz
- Upload date:
- Size: 864.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9407a90fd2e427cd3129eda7d1e407b5256506bebedf1e0eeee720fbe1488703
|
|
| MD5 |
7e9ad45a2c98395ea9d45ef5ed1dd7e0
|
|
| BLAKE2b-256 |
8d55db2fff3f4b80ec1a90ab8291ca8cf38a646272e668ef64545177ea83b3ee
|
File details
Details for the file agent_eval_forge-0.1.0-py3-none-any.whl.
File metadata
- Download URL: agent_eval_forge-0.1.0-py3-none-any.whl
- Upload date:
- Size: 173.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e3499d4998316ad348a46030c9e8c730bfe20f5a1d6779d0081e1753f5834414
|
|
| MD5 |
3365b3f46a834e806a46e07b3b27622d
|
|
| BLAKE2b-256 |
2e89c54a25aea976547f1b1dc895516fb9153b183d4de54a7592377e2b62ef44
|