Skip to main content

Agent Eval Flow — capture evidence, apply checks, compare a change

Agent Eval Flow

Release Tests and package Python 3.11+ Example reports

Turn agent runs into evidence you can use to improve the system.

Evaluate complete agent setups: models, instructions, skills, tools, loops, memory and environment. Run through an existing runtime or import saved logs, apply your checks, and compare changes in a human-readable report.

Browse the reports · Try the offline example · Read the design · Report an issue

This is a developer preview. Missing evidence stays unknown; configuration inspection and behavioral evaluation can run in parallel. The Python package is agent_eval_flow.

Try it

git clone https://github.com/guybass/agent-eval-flow.git
cd agent-eval-flow
python -m pip install -e ".[test,cli]"
python -m pytest
python examples/archive_review.py --output demo-output

The example imports a real, downloaded SWE-agent repair history containing 12 actions, failed edits, corrections and a submitted patch. It produces:

  • demo-output/report.html: task results, diagnostics and linked native evidence;
  • demo-output/result/: the typed saved evaluation;
  • demo-output/regraded/: a changed evaluation over the same saved execution.

It makes no model calls and never executes commands from the archive. The example source shows an application-owned importer and custom metrics using the production library.

Requires Python 3.11 or later. The default tests and archive example need no model credentials. Native agent runs require the integration-specific setup below.

Example reports

Two real agent workflows, controlled synthetic tasks, and specific changes measured from retained runs. Click a report to open it in your browser.

OpenSRE: recovery and stopping OpenKritt: executed proof
OpenSRE before-and-after evaluation report OpenKritt before-and-after evaluation report
Safe recovery stayed 2/2; post-report calls fell 1 → 0 after a terminal-tool binding fix. The demo grounding contract passed 0/1 → 1/1 after requiring execution and clarifying the reporting contract.
Read the report · Task and fix walkthrough Read the report · Task and fix walkthrough

These are small development experiments with synthetic tasks, not general reliability benchmarks. The gallery contains presentation reports; raw local run captures and unpublished social drafts are excluded. The offline example above is the reproducible starting point for trying the evaluation API. See the case details, run identifiers and limitations.

Configure once, then evaluate

from agent_eval_flow import EvaluationPipeline, EvaluationResult

pipeline = EvaluationPipeline(
    study=study,                 # data, candidates, execution policy and suite
    backends=backends,            # existing agent runtimes
    evaluators=evaluators,        # your metric implementations
)
result = pipeline.eval()
result.save("results/experiment")
result.report("results/experiment.html")

saved = EvaluationResult.load("results/experiment")
explanation = saved.explain(saved.runs.runs[0].id)
comparison = saved.compare("baseline", "challenger", metrics=("success_rate",))
selection = saved.select(policy)  # change cost/latency/quality priorities

# Regrade captured runs without binding or invoking an agent backend.
regraded = EvaluationPipeline(study=changed_study, evaluators=evaluators).eval(runs=saved.runs)

In an async application, use await pipeline.aeval(). Construction and planning do not start an agent. Changing the selection policy does not rerun agents or metrics. Comparisons currently provide descriptive differences; they do not invent confidence intervals.

The six objects

Object Responsibility
Study Question, candidates, execution conditions and evaluation suite
EvalDataset Keyed task tables, public inputs and private evaluator references
Candidate Model, skill, tool, flow and harness configuration
RunSet Every assignment, output, execution, event, resource observation and failure
EvalSuite Versioned metrics, dependencies, score rules and summaries
EvaluationResult Saved measurements, explanations, comparisons and selection

Pydantic validates the shared records. AnyIO coordinates runtime calls. Jinja2 renders self-contained reports. Native agent frameworks keep their loops and schedulers. Our code owns the common evidence and comparison contracts.

Missing usage remains unknown, failed work remains in the assignment inventory, and shared grading activities are counted once. A captured JSON null is distinct from absent output. Detail rows explain tasks; they do not inflate the sample size.

Runtime evidence checks

Use the same optional observation contract to check instructions, tools, model selection, loops, memory and environment state. Versioned collectors retain declared and observed values with phase, boundary, coverage and evidence. Reusable checks run through the existing evaluation pipeline, including saved run regrading; reports show the expected and observed values. Missing evidence stays unknown. See the runtime evidence guide.

Runtime integrations

Concrete adapter modules cover Codex, Claude Code, OpenSRE, OpenKritt, Harbor, SkillEvaluator imports and NeMo Agent Toolkit batch grading. Their native capture/lifecycle code is separate from the core. Prepared service connections, version-specific configuration and deployment credentials remain runtime bindings. See implementation status and setup for the supported boundaries and the live checks still required.

OpenSRE and OpenKritt are separate studies: compare candidate versions within each project. Their GCP showcase tests require prepared native runtimes and rich capture observers. Default offline test success does not establish live compatibility.

A live local OpenKritt + Codex example has now completed against 22 real Flaskr files: eight workflow executions, four post-processing executions, 20 tool results and two candidate findings. It saves native evidence, verifies the result round trip, and regrades without new model calls. Its score checks integration behavior, not security accuracy.

The separate local OpenSRE + Codex example has completed an investigation of the downloaded HDFS sample: six ReAct iterations, eight real tool calls and a cited incident report. It retains native events and CLI receipts, then saves, reloads and regrades the captured result.

Verification and design

Release checks cover the offline suite, archived-run import and regrading, package contents, and installed-wheel reporting. See the developer-preview release notes for the verified release scope.

CI runs on Windows and Ubuntu with Python 3.11, 3.12 and 3.13. Checks cover native-format fixtures, evidence transport, missing outcomes, persistence, regrading and package contents. All 90 frozen acceptance/fixture files are verified by hash. Live model tests are opt-in; offline success does not establish compatibility with every deployed agent runtime. Historical checkpoints remain in implementation status.

The older design and test-writing checkpoints are retained as historical records.

For changes and bug reports, see CONTRIBUTING.md.

Metadata

Release files for agent-eval-flow 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-eval-flow 0.5.1
File Size Uploaded
agent_eval_flow-0.5.1.tar.gz 3.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-eval-flow 0.5.1
File Interpreter ABI Platform
agent_eval_flow-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 3.2 MB

Release files / agent_eval_flow-0.5.1.tar.gz

Download URL agent_eval_flow-0.5.1.tar.gz
Size 3.0 MB
Tags Source
SHA-256 checksum
How to use checksums
02c70e270d4f294a512e7302a58be1682c6434b9ed4ef4538409fbaf12a831b5
BLAKE2b-256 checksum
How to use checksums
3189121019f2cc0fec0f8d2fad2969d840e777278b90defeef513aadd3ba4c8e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / agent_eval_flow-0.5.1-py3-none-any.whl

Download URL agent_eval_flow-0.5.1-py3-none-any.whl
Size 191.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
709d6b7c3799dfe824b3c766a5f2372d882525a4c67b446a8a5f9bf75f9a00e3
BLAKE2b-256 checksum
How to use checksums
e8eaa9d8468292116cef780f6be2a917c057a7f27dbf522d714bf46b34b0a4dc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

This release

0.5.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page