Benchmark what your agent does, not what it says.
benchspec runs an agent (Claude Code, Codex, or OpenCode) against a task in a fresh sandbox, checks what it actually did in the workspace, and reports the result as a comparison: with your skill versus without, one model versus another, one harness versus another.
End-of-run summary for a two-arm run of the in-repo hello
suite (illustrative numbers). Rows are evals, columns are arms (baseline ran
the agent bare, trial installed the skill), and every non-baseline cell shows
its assertion pass rate plus the delta against the baseline in percentage
points. The same matrix lands in benchmark.md, with machine-readable artifacts
alongside.
Teams pick harnesses, models, and prompts by anecdote: run it once, eyeball the transcript, trust the vibe. benchspec turns that guess into a measurement. Write the goal once, run it across the configurations you care about, and read off — in percentage points — how good each one actually is at accomplishing it.
What you need
| Platform | Any OS with a Docker daemon (default). Apple Silicon or Linux with /dev/kvm for the microsandbox opt-in. Python 3.11+. |
| Agent CLI | claude, codex, or opencode on PATH, with its credential (for Claude Code, CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY). |
| Binder credential | The binder is a fixed model call that classifies each assertion, required for every analyze and run: GEMINI_API_KEY by default, or OPENROUTER_API_KEY with [tool.benchspec.binder] provider = "openrouter". |
| Judge credential | The judge runs on the host through an agent CLI, using either its env credential or the CLI's own login (claude login, codex login); the default is claude-code with sonnet. Prefer a different vendor from the arms (this repo's own suite judges Claude arms with Codex). |
Credentials can live in a repo-root .env. A graded run can touch up to three
vendors: the agent's, the binder's, and the judge's. Or exactly one: set
provider = "openrouter" on the binder, the judge, and the arms, and the whole
run needs only OPENROUTER_API_KEY (see
configuration.md). lint is free;
analyze and run spend API calls, and run also boots sandboxes. Preflight
lists every missing piece and exits before anything is spent.
Getting Started
pip install benchspec
For the microsandbox isolation opt-in, add its extra:
pip install "benchspec[microsandbox]"
Pre-1.0. The eval format and the artifact schemas are the surfaces most likely to change.
An eval is one Markdown file: a prompt, then a checklist of plain-prose claims
about the workspace after the agent is done. There is no checker syntax to learn;
the wording is the spec. evals/hello/greets-by-name.eval.md:
---
---
## Prompt
You are working in a workspace rooted at your current working directory.
Greet Alice by name.
## Assertions
- [ ] ./Greetings/Alice.md contains the exact line 'Hello, Alice!'
- [ ] Skill `hello` invoked
- [ ] The greeting feels warm and personable, not curt or robotic
The benchmark is a block in pyproject.toml. Arms are the report columns; the
baseline is what the others are measured against:
[tool.benchspec]
default-set = "default"
[tool.benchspec.sets.default]
harness = "claude-code"
model = "sonnet"
baseline = "baseline"
arms = [
{ name = "baseline" }, # installs nothing
{ name = "trial" }, # setup.sh installs the skill
]
Then:
benchspec lint # static checks on the assertions
benchspec analyze # which assertions grade deterministically, which go to the judge
benchspec run # every (eval × arm) in its own sandbox, graded, reported
The Skill `hello` invoked line needs a skill for the trial arm to install
and a setup.sh that installs it; the quickstart writes
both and takes an empty directory to that first graded report.
How a run works
flowchart LR
E["greets-by-name.eval.md<br/>prompt + assertions"] --> A1["arm: baseline<br/>fresh sandbox,<br/>setup.sh installs nothing"]
E --> A2["arm: trial<br/>fresh sandbox,<br/>setup.sh installs the hello skill"]
A1 --> F1["facts: files, SHAs,<br/>final message, tool calls"]
A2 --> F2["facts"]
F1 --> G["binder: deterministic checkers<br/>everything else: LLM judge"]
F2 --> G
G --> R["benchmark.md + benchmark.json<br/>meta.json + index.jsonl"]
Each (eval × arm) pair is one parametrized pytest test. A cell:
- Boots a sandbox — a Docker container by default, or a microsandbox
microVM for the opt-in — from a cached snapshot with the agent CLI already
installed. The first run builds the snapshot (about a minute); later runs
reuse it.
benchspec sandbox:buildpays that cost up front,benchspec sandbox:cleanreclaims the disk. - Seeds the clean room — the eval's optional
workspace/files land in a fresh directory mounted at/workspace, the agent's working directory. - Runs
setup.sh, where arms diverge: it sees$BENCHSPEC_ARM, so the baseline branch exits early and the trial branch copies the skill into place. - Invokes the agent on the eval's prompt.
- Collects the facts — file tree, contents, SHA-256s, the final message, the tool calls.
- Grades — the binder maps each assertion to a deterministic checker where it can do so without risk; the judge grades everything else from the collected evidence alone.
On either backend, nothing in the guest can write back to your checkout:
setup.sh reaches the skill under test through a read-only staged copy of your
repo at /project (what a git clone would contain — never .env, .git, or
earlier runs' artifacts). Credential exposure differs — microsandbox injects each
credential at the network boundary, Docker as a plain container environment
variable the agent can read. sandbox.md has the tradeoff.
benchspec run is pytest underneath, and everything after -- goes to
pytest verbatim: benchspec run -- -k greets-by-name (equivalently
pytest -k greets-by-name) runs one eval, -n 8 fans cells across eight
sandboxes, and --count 5 samples each cell five times so the report can flag a
delta that sits within noise. The repo's own make e2e defaults to six
workers through the WORKERS variable; make e2e WORKERS=1 runs the cells
sequentially.
Why benchspec
- Comparison is first-class. A single pass rate is a number without a reference point. Arms and a baseline make the headline a delta; skip the baseline when absolute rates are what you want.
- Deterministic where possible, judged where necessary. The binder is tuned so a false positive, a surface check passing on wrong output, is the one unacceptable error; anything doubtful goes to the judge, which sees the collected evidence and never grades from recall.
- Self-describing artifacts. Every run writes
meta.json(planned config plus observed provenance, down to the agent version inside the guest),index.jsonl(one row per sample), andbenchmark.json, so other tools can aggregate runs without knowing the directory layout.
Documentation
docs/quickstart.md |
Empty directory to a graded two-arm run. |
docs/concepts.md |
The vocabulary: eval, arm, set, baseline, binder, judge. |
docs/writing-evals.md |
The eval format, workspaces, setup.sh, and how grading decides what binds. |
docs/configuration.md |
Sets, arms, the judge, every CLI flag, exit codes. |
docs/sandbox.md |
Snapshots, mounts, credentials, host requirements. |
docs/results.md |
Reading benchmark.md and the machine-readable artifacts. |
docs/harnesses.md |
claude-code, codex, opencode, and adding your own. |
Development
Clone to a graded run of the in-repo suite in four steps. Steps 1 and 2 cost nothing; step 3 is the first paid call.
git clone https://github.com/theycallmeswift/benchspec && cd benchspec
-
Install.
uvmanages the venv;make installrunsuv sync. Then the free checks:curl -LsSf https://astral.sh/uv/install.sh | sh # if you don't have uv make install make test # unit suite: no credentials, no sandbox
-
Confirm the suite collects. Needs no credentials and no Docker:
make e2e EVAL_ARGS="--collect-only -q" # 6 cells: 2 evals × 3 arms, then 8 through OpenRouter
-
Set up the credentials.
make e2erunsevals/e2e/hello/twice: two evals across three Claude Code arms, judged by Codex; then the same evals across Claude Code, Codex, and OpenCode arms with the binder, the judge, and every arm on OpenRouter. It needsclaudeandcodexonPATH, a running Docker daemon, and four credentials in.env:cp .env.example .env
Variable For CLAUDE_CODE_OAUTH_TOKEN(fromclaude setup-token) orANTHROPIC_API_KEYthe arms GEMINI_API_KEYthe binder OPENAI_API_KEYthe Codex judge OPENROUTER_API_KEYthe second run: binder, judge, and arms through OpenRouter Preflight lists every missing piece in one message and exits before anything is spent.
-
Run it. The cheapest real run is one eval, one sandbox:
make e2e WORKERS=1 EVAL_ARGS="-k greets-by-name" # one eval, sequential make e2e # the whole suite, six sandboxes
The first run builds the sandbox snapshot (about a minute); with the snapshot cached, the whole suite takes about a minute on six workers. The report lands in
tmp/evals/iteration_01/benchmark.md.
Then make lint (ruff, ty, houserules; needs GEMINI_API_KEY) before a pull
request. On Claude Code on the web, the environment's setup script does steps 1
and 3 once for every session; see
sandbox.md.
Issues and pull requests are welcome at
github.com/theycallmeswift/benchspec.
Until CONTRIBUTING.md lands, docs/style/development.md
is the code style, and make test plus make lint are the bar.
License
MIT.
Release files for benchspec 0.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchspec-0.0.3.tar.gz | 347.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchspec-0.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 523.5 kB
Release files / benchspec-0.0.3.tar.gz
| Download URL | benchspec-0.0.3.tar.gz |
|---|---|
| Size | 347.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f437a79fc5b0c5af2326af754663971df46df0bf0839b863ad841177289a2cd5
|
|
BLAKE2b-256 checksum How to use checksums |
bccce5e0babfa9c998748084fa2c382dee8ee8b3900859135017230b2a1244dc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / benchspec-0.0.3-py3-none-any.whl
| Download URL | benchspec-0.0.3-py3-none-any.whl |
|---|---|
| Size | 176.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
40894731d2c78bee1ab05d566c89ad59afc933645123891d8ffd7a4114be60f3
|
|
BLAKE2b-256 checksum How to use checksums |
eddcf1d65f9bec85f9bd3e1c653679c59b1a0eb143d1a2f95c47fbd10baa3648
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency log