Agent Eval Flow
Evaluate agent runs, understand failures, and compare changes.
Use your existing agent runtime or import retained runs into a shared evidence format. Apply your checks, inspect an HTML report, change the system, and compare another run. The target can include models, instructions, skills, tools, loops, memory and environment.
Example reports · Releases · Report an issue · Contribute
Install
Requires Python 3.11 or later.
python -m pip install agent-eval-flow
Try an offline example
git clone https://github.com/guybass/agent-eval-flow.git
cd agent-eval-flow
python -m pip install -e ".[cli]"
python examples/archive_review.py --output demo-output/archive
python examples/assessment_review.py --output demo-output/assessment
python examples/gpt_researcher_review.py --output demo-output/gpt-researcher
Open report.html inside an output directory. These examples make no model
calls.
- Archived-run evaluation imports a published SWE-agent trajectory, checks its evidence, saves the result, and regrades the same capture without running the agent again.
- Configuration assessment compares two local instruction files against an explicit requirement.
What you can evaluate
- Outcomes and execution evidence with application-defined metrics.
- Configuration and observed runtime behavior, with explicit coverage and expected versus observed values.
- Candidate changes using saved results, comparisons and selection policies.
Missing evidence stays unknown. Runtime-specific importers and collectors map native logs into the shared records; arbitrary logs are not interpreted automatically. Native adapters keep the agent's own execution loop.
Tool calling report
Add deterministic Toolscore diagnostics to captured runs and generate a separate tools report alongside the general report:
python -m pip install -e ".[toolscore]"
python examples/toolscore_review.py --output demo-output/toolscore
The offline example writes linked report.html and tools.html, tools.json,
and verified evidence files. It uses synthetic retained streams and makes no
agent or model calls. Toolscore evaluates requested tools and arguments;
task outcomes remain separate checks. Missing capture stays unknown.
Use result.report(path, tools=True) after configuring the optional evaluator.
See the integration guide for trace coverage, expected
calls, scoring rules, and supported formats.
Example reports
These are small native-agent experiments on controlled synthetic tasks. They demonstrate specific changes, not general reliability or security accuracy. Read the reports, walkthroughs and limits.
Contributing
Contributions are welcome, including documentation fixes, offline examples, regression tests, adapters, and report improvements. You can get started without model credentials. Small fixes can go straight to a pull request; for larger features, open an issue to discuss the approach.
Start with the contribution guide for setup on Windows, macOS, or Linux, a code map, local checks, and the steps to your first pull request.
This is a developer preview. Native integrations require their own runtime setup and credentials; offline test success does not establish live compatibility.
See the test guide for live profiles, and third-party notices for fixture provenance and licenses.
Metadata
Release files for agent-eval-flow 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_eval_flow-0.6.0.tar.gz | 1.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_eval_flow-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.7 MB
Release files / agent_eval_flow-0.6.0.tar.gz
| Download URL | agent_eval_flow-0.6.0.tar.gz |
|---|---|
| Size | 1.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e91d0a2d3608b818745eee63644ebaa5d426989743d9693592fc1e0b3591e7a3
|
|
BLAKE2b-256 checksum How to use checksums |
012d761a61e71ad5f78302e265dc70b8625461832a8ac07f3033bfeaefa553d3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / agent_eval_flow-0.6.0-py3-none-any.whl
| Download URL | agent_eval_flow-0.6.0-py3-none-any.whl |
|---|---|
| Size | 203.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
276c967a09433efd608277c5917bdf635aa123faac66e83a8cc96b9a008a0629
|
|
BLAKE2b-256 checksum How to use checksums |
1fde6ff2c59a04fbd791ea948a171c44c51ffab81ea1d8f306d9081272f4f9c1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log