Skip to main content

agentbench

agentbench is an evaluation framework for AI agents. It provides core abstractions for tasks, scenarios, agents, and judges, with an async runner that supports retries, rate limiting, and comprehensive reporting.

Features

  • Core abstractions: Task, Scenario, AgentAdapter, Judge, and CompositeJudge
  • Async execution: Built-in retry logic, fail-fast behavior, and provider-based rate limiting
  • Lifecycle hooks: Setup, reset, and teardown methods for agent management
  • Comprehensive reporting: HTML reports, JSONL results, traces, and metadata
  • Run comparison: Compare multiple runs to identify regressions and improvements
  • Presets: Ready-to-use scenarios and judges for quick evaluation

Quick start

Install in editable mode:

pip install -e .

Run a simple example using presets:

python examples/math_with_presets.py

Example

Here's a minimal example that evaluates a simple math agent:

from agentbench import (
    RunConfig,
    run,
    build_math_basic_scenario,
    ExactMatchJudge,
)


class SimpleMathAgent:
    name = "simple_math_agent"
    version = "0.0.1"
    provider_key = None

    async def setup(self) -> None:
        return None

    async def reset(self) -> None:
        return None

    async def teardown(self) -> None:
        return None

    async def run_task(self, task, context=None):
        prompt = task.input["prompt"]
        if "2 + 3" in prompt:
            response = "5"
        elif "4 * 7" in prompt:
            response = "28"
        else:
            response = "I do not know yet"
        return {"response": response}


scenario = build_math_basic_scenario()
agent = SimpleMathAgent()
judge = ExactMatchJudge()

config = RunConfig(
    name="math_example",
    agents=[agent],
    scenarios=[scenario],
    judges=[judge],
)

results = run(config)
print(f"Got {sum(1 for r in results if r.passed)} / {len(results)} passing results.")

Run artifacts

Running an evaluation creates a directory under runs/<run_id>/ with:

  • results.jsonl: All evaluation results in JSONL format
  • traces.jsonl: Event traces for debugging and analysis
  • run_metadata.json: Run configuration, environment, and version information
  • report.html: Interactive HTML report with statistics, per-agent performance, and failing tasks

Open report.html in your browser to view the detailed evaluation report.

Testing

agentbench includes a comprehensive test suite with 43+ tests covering all major components:

  • Core types (Task, Cost, EvaluationResult, Trace)
  • Scenarios and presets
  • Judges (ExactMatchJudge, CompositeJudge)
  • Runner and lifecycle hooks
  • Storage and reporting functions
  • Integration tests

Run tests:

# Install dev dependencies
pip install -e ".[dev]"

# Run all tests
pytest

# Generate HTML test report
pytest --html=test-results/report.html --self-contained-html

# Generate coverage report
pytest --cov=src/agentbench --cov-report=html:htmlcov

Test reports are generated in test-results/ and coverage reports in htmlcov/.

Project status

agentbench is in active development. The core framework is functional and ready for evaluation use cases.

Current version: 0.0.1

Release files for agentflowtest 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentflowtest 0.0.1
File Size Uploaded
agentflowtest-0.0.1.tar.gz 20.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentflowtest 0.0.1
File Interpreter ABI Platform
agentflowtest-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 38.7 kB

Release files / agentflowtest-0.0.1.tar.gz

Download URL agentflowtest-0.0.1.tar.gz
Size 20.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a1cc19d30caf30f32f4b4ccf5b3b81df3bb1bf0e0368fc2da54b3a70f646fd5f
BLAKE2b-256 checksum
How to use checksums
b025b2a1162692c79de1a393194415eb4580c974807567311ae4b0f762f2baa6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.14

Release files / agentflowtest-0.0.1-py3-none-any.whl

Download URL agentflowtest-0.0.1-py3-none-any.whl
Size 18.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4ab86755ca8ac29860e28d9186b3c340181528283599e2411360e8fffbb23249
BLAKE2b-256 checksum
How to use checksums
b783a4987075bb307d7d47faeafde795e585819f5ebcc815b0fb11a51bbb9ecd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.14

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page