Agent Eval Flow
Turn agent runs into evidence you can use to improve the system.
Evaluate complete agent setups: models, instructions, skills, tools, loops, memory and environment. Run through an existing runtime or import saved logs, apply your checks, and compare changes in a human-readable report.
Browse the reports · Try the offline example · Read the design · Report an issue
This is a developer preview. Missing evidence stays unknown; configuration
inspection and behavioral evaluation can run in parallel. The Python package
is agent_eval_flow.
Try it
git clone https://github.com/guybass/agent-eval-flow.git
cd agent-eval-flow
python -m pip install -e ".[test,cli]"
python -m pytest
python examples/archive_review.py --output demo-output
The example imports a real, downloaded SWE-agent repair history containing 12 actions, failed edits, corrections and a submitted patch. It produces:
demo-output/report.html: task results, diagnostics and linked native evidence;demo-output/result/: the typed saved evaluation;demo-output/regraded/: a changed evaluation over the same saved execution.
It makes no model calls and never executes commands from the archive. The example source shows an application-owned importer and custom metrics using the production library.
Requires Python 3.11 or later. The default tests and archive example need no model credentials. Native agent runs require the integration-specific setup below.
Example reports
Two real agent workflows, controlled synthetic tasks, and specific changes measured from retained runs. Click a report to open it in your browser.
| OpenSRE: recovery and stopping | OpenKritt: executed proof |
|---|---|
| Safe recovery stayed 2/2; post-report calls fell 1 → 0 after a terminal-tool binding fix. | The demo grounding contract passed 0/1 → 1/1 after requiring execution and clarifying the reporting contract. |
| Read the report · Task and fix walkthrough | Read the report · Task and fix walkthrough |
These are small development experiments with synthetic tasks, not general reliability benchmarks. The gallery contains presentation reports; raw local run captures and unpublished social drafts are excluded. The offline example above is the reproducible starting point for trying the evaluation API. See the case details, run identifiers and limitations.
Configure once, then evaluate
from agent_eval_flow import EvaluationPipeline, EvaluationResult
pipeline = EvaluationPipeline(
study=study, # data, candidates, execution policy and suite
backends=backends, # existing agent runtimes
evaluators=evaluators, # your metric implementations
)
result = pipeline.eval()
result.save("results/experiment")
result.report("results/experiment.html")
saved = EvaluationResult.load("results/experiment")
explanation = saved.explain(saved.runs.runs[0].id)
comparison = saved.compare("baseline", "challenger", metrics=("success_rate",))
selection = saved.select(policy) # change cost/latency/quality priorities
# Regrade captured runs without binding or invoking an agent backend.
regraded = EvaluationPipeline(study=changed_study, evaluators=evaluators).eval(runs=saved.runs)
In an async application, use await pipeline.aeval(). Construction and planning
do not start an agent. Changing the selection policy does not rerun agents or
metrics. Comparisons currently provide descriptive differences; they do not
invent confidence intervals.
The six objects
| Object | Responsibility |
|---|---|
Study |
Question, candidates, execution conditions and evaluation suite |
EvalDataset |
Keyed task tables, public inputs and private evaluator references |
Candidate |
Model, skill, tool, flow and harness configuration |
RunSet |
Every assignment, output, execution, event, resource observation and failure |
EvalSuite |
Versioned metrics, dependencies, score rules and summaries |
EvaluationResult |
Saved measurements, explanations, comparisons and selection |
Pydantic validates the shared records. AnyIO coordinates runtime calls. Jinja2 renders self-contained reports. Native agent frameworks keep their loops and schedulers. Our code owns the common evidence and comparison contracts.
Missing usage remains unknown, failed work remains in the assignment inventory,
and shared grading activities are counted once. A captured JSON null is distinct
from absent output. Detail rows explain tasks; they do not inflate the sample size.
Runtime evidence checks
Use the same optional observation contract to check instructions, tools, model selection, loops, memory and environment state. Versioned collectors retain declared and observed values with phase, boundary, coverage and evidence. Reusable checks run through the existing evaluation pipeline, including saved run regrading; reports show the expected and observed values. Missing evidence stays unknown. See the runtime evidence guide.
Runtime integrations
Concrete adapter modules cover Codex, Claude Code, OpenSRE, OpenKritt, Harbor, SkillEvaluator imports and NeMo Agent Toolkit batch grading. Their native capture/lifecycle code is separate from the core. Prepared service connections, version-specific configuration and deployment credentials remain runtime bindings. See implementation status and setup for the supported boundaries and the live checks still required.
OpenSRE and OpenKritt are separate studies: compare candidate versions within each project. Their GCP showcase tests require prepared native runtimes and rich capture observers. Default offline test success does not establish live compatibility.
A live local OpenKritt + Codex example has now completed against 22 real Flaskr files: eight workflow executions, four post-processing executions, 20 tool results and two candidate findings. It saves native evidence, verifies the result round trip, and regrades without new model calls. Its score checks integration behavior, not security accuracy.
The separate local OpenSRE + Codex example has completed an investigation of the downloaded HDFS sample: six ReAct iterations, eight real tool calls and a cited incident report. It retains native events and CLI receipts, then saves, reloads and regrades the captured result.
Verification and design
Release checks cover the offline suite, archived-run import and regrading, package contents, and installed-wheel reporting. See the developer-preview release notes for the verified release scope.
CI runs on Windows and Ubuntu with Python 3.11, 3.12 and 3.13. Checks cover native-format fixtures, evidence transport, missing outcomes, persistence, regrading and package contents. All 90 frozen acceptance/fixture files are verified by hash. Live model tests are opt-in; offline success does not establish compatibility with every deployed agent runtime. Historical checkpoints remain in implementation status.
- Data structures and configuration
- Unified assessment flow: configuration and execution in parallel
- Module tree and communication graph
- Real-workflow E2E scenarios
- Visual pipeline and score examples
- Library reuse decisions
- What this adds to existing work
The older design and test-writing checkpoints are retained as historical records.
For changes and bug reports, see CONTRIBUTING.md.
Metadata
Release files for agent-eval-flow 0.5.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_eval_flow-0.5.1.tar.gz | 3.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_eval_flow-0.5.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.2 MB
Release files / agent_eval_flow-0.5.1.tar.gz
| Download URL | agent_eval_flow-0.5.1.tar.gz |
|---|---|
| Size | 3.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
02c70e270d4f294a512e7302a58be1682c6434b9ed4ef4538409fbaf12a831b5
|
|
BLAKE2b-256 checksum How to use checksums |
3189121019f2cc0fec0f8d2fad2969d840e777278b90defeef513aadd3ba4c8e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency logRelease files / agent_eval_flow-0.5.1-py3-none-any.whl
| Download URL | agent_eval_flow-0.5.1-py3-none-any.whl |
|---|---|
| Size | 191.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
709d6b7c3799dfe824b3c766a5f2372d882525a4c67b446a8a5f9bf75f9a00e3
|
|
BLAKE2b-256 checksum How to use checksums |
e8eaa9d8468292116cef780f6be2a917c057a7f27dbf522d714bf46b34b0a4dc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency log