Skip to main content

Agent Eval Flow

Agent Eval Flow

PyPI Tests Python 3.11+

Evaluate agent runs, understand failures, and compare changes.

Use your existing agent runtime or import retained runs into a shared evidence format. Apply your checks, inspect an HTML report, change the system, and compare another run. The target can include models, instructions, skills, tools, loops, memory and environment.

Example reports · Releases · Report an issue · Contribute

Install

Requires Python 3.11 or later.

python -m pip install agent-eval-flow

Try an offline example

git clone https://github.com/guybass/agent-eval-flow.git
cd agent-eval-flow
python -m pip install -e ".[cli]"
python examples/archive_review.py --output demo-output/archive
python examples/assessment_review.py --output demo-output/assessment
python examples/gpt_researcher_review.py --output demo-output/gpt-researcher

Open report.html inside an output directory. These examples make no model calls.

  • Archived-run evaluation imports a published SWE-agent trajectory, checks its evidence, saves the result, and regrades the same capture without running the agent again.
  • Configuration assessment compares two local instruction files against an explicit requirement.

What you can evaluate

  • Outcomes and execution evidence with application-defined metrics.
  • Configuration and observed runtime behavior, with explicit coverage and expected versus observed values.
  • Candidate changes using saved results, comparisons and selection policies.

Missing evidence stays unknown. Runtime-specific importers and collectors map native logs into the shared records; arbitrary logs are not interpreted automatically. Native adapters keep the agent's own execution loop.

Tool calling report

Add deterministic Toolscore diagnostics to captured runs and generate a separate tools report alongside the general report:

python -m pip install -e ".[toolscore]"
python examples/toolscore_review.py --output demo-output/toolscore

The offline example writes linked report.html and tools.html, tools.json, and verified evidence files. It uses synthetic retained streams and makes no agent or model calls. Toolscore evaluates requested tools and arguments; task outcomes remain separate checks. Missing capture stays unknown.

Use result.report(path, tools=True) after configuring the optional evaluator. See the integration guide for trace coverage, expected calls, scoring rules, and supported formats.

Example reports

OpenSRE OpenKritt
OpenSRE report OpenKritt report
Safe recovery stayed 2/2; late tool calls fell 1 → 0 after a stopping-contract fix. The demo grounding check passed 0/1 → 1/1 after requiring execution and clarifying the reporting contract.

These are small native-agent experiments on controlled synthetic tasks. They demonstrate specific changes, not general reliability or security accuracy. Read the reports, walkthroughs and limits.

Contributing

Contributions are welcome, including documentation fixes, offline examples, regression tests, adapters, and report improvements. You can get started without model credentials. Small fixes can go straight to a pull request; for larger features, open an issue to discuss the approach.

Start with the contribution guide for setup on Windows, macOS, or Linux, a code map, local checks, and the steps to your first pull request.

This is a developer preview. Native integrations require their own runtime setup and credentials; offline test success does not establish live compatibility.

See the test guide for live profiles, and third-party notices for fixture provenance and licenses.

Metadata

Release files for agent-eval-flow 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-eval-flow 0.6.0
File Size Uploaded
agent_eval_flow-0.6.0.tar.gz 1.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-eval-flow 0.6.0
File Interpreter ABI Platform
agent_eval_flow-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.7 MB

Release files / agent_eval_flow-0.6.0.tar.gz

Download URL agent_eval_flow-0.6.0.tar.gz
Size 1.5 MB
Tags Source
SHA-256 checksum
How to use checksums
e91d0a2d3608b818745eee63644ebaa5d426989743d9693592fc1e0b3591e7a3
BLAKE2b-256 checksum
How to use checksums
012d761a61e71ad5f78302e265dc70b8625461832a8ac07f3033bfeaefa553d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / agent_eval_flow-0.6.0-py3-none-any.whl

Download URL agent_eval_flow-0.6.0-py3-none-any.whl
Size 203.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
276c967a09433efd608277c5917bdf635aa123faac66e83a8cc96b9a008a0629
BLAKE2b-256 checksum
How to use checksums
1fde6ff2c59a04fbd791ea948a171c44c51ffab81ea1d8f306d9081272f4f9c1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page