Multi-turn agent benchmarking with ACP — run any agent, any model, any provider.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

bingran-you xdotli

These details have not been verified by PyPI

Project links

Discord

Project description

BenchFlow

The universal environment framework — a benchmark is just a frozen environment.

What

BenchFlow is a universal environment framework: it runs AI agents against task environments and scores them through one hardened contract. A benchmark is just a frozen environment — point BenchFlow at any of them, drive it with any ACP agent, and run single-agent, multi-agent, or multi-round patterns over the same Scene-based lifecycle.

Run any benchmark — three-layer routing runs supported frameworks natively, translates unknown formats and proves equivalence with a parity gate, or runs a bespoke harness as-is; every layer emits one scored-trajectory contract. See Run any benchmark
Any ACP agent — Gemini CLI, Claude Code, Codex, OpenCode, OpenHands, Pi, or your own
Single + multi + progressive — single-agent / multi-agent (coder + reviewer, simulated user) / multi-round with a Python BaseUser callback
Loop strategies — wrap any agent in a --loop-strategy (verify-retry, self-review); every rollout captures a per-iteration reward + token trajectory, so you can plot capability against cost (can a cheap model + loops match an expensive one at equal token spend?)
task.md tasks — one file (YAML frontmatter + prompt body) replaces the split task.toml + instruction.md layout; author with bench tasks init / check / migrate / export
Hosted environments — run external PrimeIntellect / Verifiers environments through --source-env, without converting them to BenchFlow tasks
Sandboxes — Docker locally, Daytona for parallel cloud runs (orphaned sandboxes auto-reaped at eval start), Modal for serverless/GPU-backed task environments
Hardened verifier — defaults block BenchJack/Meerkat-style reward-hacking; tasks opt out per-feature
Training-ready output — every scored rollout emits ATIF (trainer/atif.json) and ADP (trainer/adp.jsonl) trajectory records next to the Verifiers/ORS (OpenReward) reward record

Quickstart

# Install or upgrade to the latest stable BenchFlow CLI
uv tool install --python 3.12 --upgrade benchflow

# Run a benchmark: any task source, any ACP agent, any sandbox
export GEMINI_API_KEY=...            # or claude login / codex --login for subscription auth
bench eval run \
    --source-repo benchflow-ai/skillsbench --source-path tasks \
    --agent gemini --model gemini-3.1-flash-lite-preview \
    --sandbox docker

Each run writes a per-task result.json (rewards + trajectory + token usage) and a job summary.json (pass-rate, cost, and — for looped runs — a pass@iteration convergence curve). New here? Start with Getting started, or paste the agent quickstart prompt into Claude Code / Codex / Gemini CLI and let it drive the whole thing.

Install

Install or upgrade to the latest stable release from PyPI with uv:

uv tool install --python 3.12 --upgrade benchflow

Confirm with bench --version.
BenchFlow CLI releases require Python 3.12 or newer. Keep --python 3.12 in the install command so uv does not resolve an older Python-compatible package that lacks the CLI entrypoints.
If you see Executables already exist: bench, benchflow, re-run with uv tool install --python 3.12 --upgrade --force benchflow to replace stale entrypoints from an older install.
For Daytona or Modal extras, install the relevant optional package, for example uv tool install --python 3.12 --upgrade 'benchflow[sandbox-daytona]'.

Internal users wanting the newest preview from main install the internal preview channel (uv tool install --python 3.12 --prerelease allow --upgrade benchflow).

Requirements & auth. Install uv; the --python 3.12 flag lets it provision a compatible interpreter for the tool install. Set DAYTONA_API_KEY for Daytona or configure Modal auth for Modal; export an agent API key (GEMINI_API_KEY, ANTHROPIC_API_KEY, …) or use subscription auth (claude login / codex --login). Provider-prefixed models may need provider-specific credentials; Azure Foundry uses AZURE_API_KEY + AZURE_API_ENDPOINT.

Documentation

Start with Getting started, then Concepts for the mental model. Prefer to have an AI coding agent run the whole quickstart for you? Paste the agent quickstart prompt into Claude Code, Codex CLI, or Gemini CLI. Then by goal:

If you want to…	Read
Run an eval on an existing task	Getting started
Understand how BenchFlow runs any benchmark (the three-layer model)	Run any benchmark
Have an AI agent install + run the quickstart end to end	Agent quickstart prompt
Understand Rollout / Scene / Role / Verifier	Concepts
Author a new task	Task authoring
Author a task in the native `task.md` format	Native task.md authoring
Run a hosted PrimeIntellect / Verifiers environment	CLI reference
Multi-agent: coder + reviewer, simulated user, BYOS, stateful envs	Use cases
Multi-round single-agent (progressive disclosure, oracle access)	Progressive disclosure
Skill evaluation (when the artifact is a skill, not a workspace)	Skill eval
Understand the security model	Sandbox hardening
Use public vs internal preview SDK releases	Release channels
CLI flags + commands	CLI reference
Python API surface	Python API reference

Notebooks and runnable example scripts live under docs/examples/ so examples stay versioned with the docs that explain them.

bench agent vs bench eval adopt. bench agent list / bench agent show inspect registered AI agents (the solver programs like Claude Code or Gemini CLI). Onboarding a third-party benchmark into benchmarks/<name>/ is a separate workflow — bench eval adopt <source> scaffolds and drives the conversion, and bench eval adopt <name> --verify parity-gates it. (The legacy bench agent create|run|verify commands still work as deprecated aliases.) See the CLI reference for details.

Benchmark task sources

Benchmark datasets live in external Git repos and are referenced with two fields:

# benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yaml
source:
  repo: benchflow-ai/benchmarks    # GitHub org/repo
  path: datasets/harvey-lab/tasks  # optional subpath within repo
  ref: main                         # optional branch/tag
agent: gemini
model: gemini/gemini-3.1-flash-lite-preview

Run any benchmark via the CLI:

# From a YAML config (shipped with the repo)
bench eval run --config benchmarks/harvey-lab/harvey-lab-gemini-flash-lite.yaml

# Inline — mirrors the YAML source fields
bench eval run \
    --source-repo benchflow-ai/skillsbench --source-path tasks \
    --agent gemini --model gemini-3.1-flash-lite-preview --sandbox daytona --concurrency 64

Repos are cloned and cached locally under .cache/datasets/ on first use.

Hosted environments are another source type. Instead of a repo, pass --source-env with the environment's pinned source version to run an external PrimeIntellect / Verifiers environment on its own native harness — BenchFlow preserves the hosted identity (env_uid, hub_url) and still writes the shared rollout output contract. See the CLI reference for the full hosted-environment command shape.

Downstream projects should depend on the public PyPI release by default. For internal validation before the next public release, install or lock the internal preview channel with prereleases enabled; see Release channels.

Authoring tasks

A task is one task.md (YAML frontmatter for config + a markdown prompt body) plus environment/ and verifier/ sidecars. The bench tasks commands cover the authoring lifecycle:

bench tasks init my-task                 # scaffold a task.md package under tasks/
bench tasks check tasks/my-task          # validate (default --level structural)
bench tasks migrate legacy-task/ --remove-legacy  # convert old split packages to task.md
bench tasks export tasks/my-task out/             # write a compatibility export + loss report

See Native task.md authoring and the task standard.

Featured

Progressive disclosure on SWE-bench Pro — the BaseUser abstraction drives a multi-round rollout: terse round-0 prompt → failing-test hints → full spec. 5/5 oracle on Daytona, runnable demo at docs/examples/swebench_pro_progressive_disclosure.ipynb. See Progressive disclosure.

Audience

Eval researchers / paper writers → Getting started → Concepts → Use cases
Task authors → Task authoring → Sandbox hardening
Agent builders integrating with benchflow → Concepts → Python API reference → benchflow.agents.registry
External benchmark adapters → Task authoring → Progressive disclosure

Contributing

PRs welcome. Open against main. CI runs ruff + tests on every PR; please run ruff check . and pytest tests/ locally first.

Release channels are documented in Release channels. In short: merges to main publish an internal preview after CI passes, while a matching release tag publishes the public release.

License

Apache-2.0.

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

bingran-you xdotli

These details have not been verified by PyPI

Project links

Discord

Release history Release notifications | RSS feed

This version

0.6.5

Jul 11, 2026

0.6.4

Jun 27, 2026

0.6.3

Jun 16, 2026

0.6.2

Jun 14, 2026

0.6.1

Jun 14, 2026

0.6.0

Jun 13, 2026

0.5.3.dev994 pre-release

Jun 12, 2026

0.5.3.dev992 pre-release

Jun 12, 2026

0.5.3.dev983 pre-release

Jun 12, 2026

0.5.3.dev980 pre-release

Jun 12, 2026

0.5.3.dev978 pre-release

Jun 12, 2026

0.5.3.dev976 pre-release

Jun 12, 2026

0.5.3.dev957 pre-release

Jun 11, 2026

0.5.3.dev956 pre-release

Jun 11, 2026

0.5.3.dev908 pre-release

Jun 9, 2026

0.5.3.dev906 pre-release

Jun 8, 2026

0.5.3.dev902 pre-release

Jun 8, 2026

0.5.3.dev899 pre-release

Jun 8, 2026

0.5.3.dev894 pre-release

Jun 8, 2026

0.5.3.dev885 pre-release

Jun 5, 2026

0.5.3.dev883 pre-release

Jun 5, 2026

0.5.3.dev881 pre-release

Jun 5, 2026

0.5.3.dev879 pre-release

Jun 5, 2026

0.5.2

Jun 5, 2026

0.5.2.dev875 pre-release

Jun 5, 2026

0.5.1

Jun 5, 2026

0.5.1.dev871 pre-release

Jun 5, 2026

0.5.1.dev869 pre-release

Jun 5, 2026

0.5.0

Jun 5, 2026

0.4.0

May 20, 2026

0.3.4

May 17, 2026

0.3.3

May 16, 2026

0.3.2

Apr 23, 2026

0.3.1

Apr 22, 2026

0.3.0

Apr 21, 2026

0.3.0a10 pre-release

Apr 20, 2026

0.3.0a9 pre-release

Apr 20, 2026

0.3.0a8 pre-release

Apr 20, 2026

0.3.0a7 pre-release

Apr 20, 2026

0.3.0a6 pre-release

Apr 20, 2026

0.3.0a5 pre-release

Apr 20, 2026

0.3.0a4 pre-release

Apr 20, 2026

0.3.0a3 pre-release

Apr 20, 2026

0.3.0a2 pre-release

Apr 20, 2026

0.3.0a1 pre-release

Apr 20, 2026

0.2.3

Apr 16, 2026

0.2.2

Apr 14, 2026

0.2.1

Apr 13, 2026

0.2.0

Apr 9, 2026

0.1.13

Mar 10, 2025

0.1.12

Mar 6, 2025

0.1.11

Mar 6, 2025

0.1.10

Mar 6, 2025

0.1.9

Feb 28, 2025

0.1.8

Feb 27, 2025

0.1.7

Feb 19, 2025

0.1.6

Feb 17, 2025

0.1.5

Feb 7, 2025

0.1.4

Feb 4, 2025

0.1.3

Jan 31, 2025

0.1.2

Jan 25, 2025

0.1.1

Jan 24, 2025

0.1.0

Jan 22, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

benchflow-0.6.5.tar.gz (1.6 MB view details)

Uploaded Jul 11, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

benchflow-0.6.5-py3-none-any.whl (913.3 kB view details)

Uploaded Jul 11, 2026 Python 3

File details

Details for the file benchflow-0.6.5.tar.gz.

File metadata

Download URL: benchflow-0.6.5.tar.gz
Upload date: Jul 11, 2026
Size: 1.6 MB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for benchflow-0.6.5.tar.gz
Algorithm	Hash digest
SHA256	`b6f34b29f4ec39661eff888dc33d546c08ba0ab85348ebad7d658c7041fb64aa`
MD5	`9821ea929a5a0d8fe50192a3b97f13c7`
BLAKE2b-256	`9896b0045f460cb0f4b7f32c8fde1a4fff3681400f7d03474abb82fee4c88cb4`

See more details on using hashes here.

File details

Details for the file benchflow-0.6.5-py3-none-any.whl.

File metadata

Download URL: benchflow-0.6.5-py3-none-any.whl
Upload date: Jul 11, 2026
Size: 913.3 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for benchflow-0.6.5-py3-none-any.whl
Algorithm	Hash digest
SHA256	`6ca34010ab3c47627214280c26b2bf487113534454b1c5a374bb6c3c34f7a87a`
MD5	`8f2b4532d1ecfe42d5d9d9bf339d10d1`
BLAKE2b-256	`c5300caddec2ef504dc6a6cff10bae98a3936f6d9bfcac76bf4414f629b16482`

See more details on using hashes here.

benchflow 0.6.5

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

BenchFlow

What

Quickstart

Install

Documentation

Benchmark task sources

Authoring tasks

Featured

Audience

Contributing

License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes