pytest-xharness-eval 🧪🤖
pytest plugin for cross AI agent harness evaluation.
Write the eval once. Run it against every harness and every model.
What it does
A pytest plugin that runs the claude and codex CLIs headlessly against a fixture
workspace, captures each run's own session log, prices it, and grades what the agent
left behind. A skill opts in by adding an evals/ directory; pytest does the rest.
Every eval cell is live and costs money. There is no replay mode. Preview the spend
with --dry-run before a sweep. The design rationale lives in
ARCHITECTURE.md and the decision log in
docs/adrs/; agents start at AGENTS.md.
Quickstart
-
Install the plugin into the repository that holds your skills. The
pytest11entry point registers it; noconftest.pywiring is needed:uv add --dev pytest-xharness-eval
-
Pin pytest's rootdir to the repository root, so the plugin finds
skills/from any argument path. An empty[tool.pytest.ini_options]table is enough:[tool.pytest.ini_options] -
Add an eval beside the skill. Files and functions both carry the
eval_prefix, astest_does for pytest. Fixtures are seed workspaces copied fresh for every cell; every run output lands under one git-ignored cache root, never in the skills tree (ADR 0032):skills/<skill>/ SKILL.md evals/ eval_<suite>.py fixtures/<name>/ goldens/<name>/ # optional known-good output (ADR 0046) .xharness_eval_cache/ build/ # per-cell workspaces results/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/ # log.jsonl, result.json, history.json report/ # report.json + the aggregated micrositefrom pytest_xharness_eval import CaseOutput, evalcase from pytest_xharness_eval.verify import check_files_written, check_rollout @evalcase(task="...", skill="<skill>", fixture="<name>") def eval_<case>(output: CaseOutput) -> None: check_rollout(output) # real session, billed, priced check_files_written(output, "OUTPUT.md") # this run is what produced it assert "the thing" in output.read("OUTPUT.md")
taskis what a user types after naming the skill, it never names the skill, a CLI, or where aSKILL.mdlives. Each harness renders its own invocation around it:/<skill> <task>forclaude,$<skill> <task>forcodex(ADR 0044). The full grader surface, every field you can assert on, the bundled verifiers, and the goldens convention, is docs/rollout.md. -
Preview the matrix. Nothing is invoked:
uv run pytest skills/<skill>/evals --dry-run
xharness-eval: skills root = /repo/skills, cache = /repo/.xharness_eval_cache xharness-eval: matrix = plugin default (2 entries); a case's models= overrides it collected 2 items skills/<skill>/evals/eval_<case>.py ss ============================ agent eval report ============================ dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-sonnet-5] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-haiku-4-5-20251001] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-sol] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-luna] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-terra] total spend: $0.0000 across 6 cell(s) report: /repo/.xharness_eval_cache/report/report.json
-
Run it live, with
-vso every cell reports its verdict, USD, context, wall clock, turns, and tool calls as it lands. Add-n 2to run cells in parallel. This spends money:uv run pytest skills/<skill>/evals -v
skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] PASSED est $0.5762 (harness $0.5773) 352,451 accumulative_billed_tokens 23,898 baseline_tokens 76.0s 9 turns 8 tools
Read the status word as: this plugin's estimate from its price table (and the harness CLI's own figure, where it reports one), every billed token summed over all turns (the cached prefix is re-read each turn), the harness's own prompt on turn 1, wall clock, model calls, tool calls. Every estimate records the rates it used and where they came from (
rates_applied).Each cell leaves its verbatim session log (
log.jsonl), a normalisedresult.jsonwith a per-turn ledger, and onehistory.jsonmetrics record in its ownresults/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/directory, no two cells share a file, so parallel workers never contend (ADR 0032). At session end the one combine step aggregates everything underresults/, every skill, every run, intoreport/:report.json, the accumulatedhistory.jsonl, and a browsablereport.htmlwith its glossary (XHARNESS-REPORT-GLOSSARY.md) beside it. Serve it withpython3 -m http.server --directory .xharness_eval_cacheand open/report/report.html; it fetches the JSON beside it.
run is a RunResult: session id, log path, token usage by tier, tool calls, files
written, and USD cost. The reference case, with its assertions written as a tutorial,
is
eval_palette_mandate.py, the reference case kept beside the skill it grades in the consuming repository.
Narrow a run
The matrix is the spend dial. These options are the plugin's own; everything else is
stock pytest (-k, -x, -m eval, node ids).
| Option | Effect | Example |
|---|---|---|
| path | One skill or all of them | pytest skills/x/evals, pytest skills/*/evals |
--harness <name> |
Only cells for that harness (claude or codex), repeatable |
pytest skills/x/evals --harness codex |
--model <substring> |
Only cells whose model id contains the string, or one exact harness/model, repeatable |
pytest skills/x/evals --model opus |
--effort <rung> |
Only cells at that reasoning rung, repeatable. Matches the resolved rung, so --effort max selects claude's max and codex's xhigh alike |
pytest skills/x/evals --effort max |
-k <expr> |
Boolean slices over cell ids and case names (stock pytest) | -k "opus or sol", -k "codex and not sol" |
--xharness-timeout <s> |
Seconds one cell's CLI may run before it is killed (default 600). Raise it for the top effort rungs, which think for longer by design | pytest skills/x/evals --xharness-timeout 1800 |
--dry-run |
Enumerate cells and validate pricing, invoke nothing | pytest skills/x/evals --dry-run |
--collect-only -q |
List cell node ids (stock pytest) | pytest --collect-only -q skills/x/evals |
Do not run pytest skills from the root: it walks into every skill's scripts/
directory and collects their unit tests too. skills/*/evals is the full matrix.
A case overrides the matrix with @evalcase(..., models=["codex/gpt-5.6-sol"]).
The effort axis
A matrix entry may name a third component, the reasoning budget the CLI is asked for:
[tool.pytest.ini_options]
xharness_matrix = """
claude/claude-opus-5/low
claude/claude-opus-5/max
codex/gpt-5.6-sol/mid
"""
Each line is a separately graded cell, so one model at two rungs is the cost-versus-quality comparison the axis exists for. An entry with no third component is unchanged: it leaves the CLI on whatever default its own configuration gives it.
Each harness declares its own ladder, and three portable aliases name a position on it rather than a level:
| You write | Position | On claude |
On codex |
|---|---|---|---|
min |
first rung | low |
low |
mid |
middle rung | high |
high |
max |
last rung | max |
max |
Both shipped CLIs happen to declare the same five rungs (low, medium, high, xhigh, max),
so the aliases resolve identically on each today. That is a fact about these two CLIs and
not a rule: the ladder lives on the harness class, so a third CLI may declare any rungs it
likes and the aliases keep working by position.
Resolution happens once, at collection, so a node id, an evidence directory and a report
row all carry the rung that was actually sent. A rung no harness has
(claude/claude-opus-5/minimal) stops the sweep at collection, before anything is spent.
That check is not theoretical: minimal appears in codex-cli's own local enum, and a paid
sweep found that no gpt-5.6 model accepts it — the CLI forwards it, the API answers 400,
and the run exits having produced nothing (ADR 0049).
Configuration
The matrix has three scopes, highest precedence first: a case's models=, the
project's xharness_matrix ini key, and the plugin's bundled default. The report
header names which one applied.
The bundled default is every model the bundled price table carries: three per harness,
claude/{claude-opus-5, claude-sonnet-5, claude-haiku-4-5-20251001} and
codex/{gpt-5.6-sol, gpt-5.6-luna, gpt-5.6-terra}. An axis nobody narrowed means the
whole axis, so this is deliberately the widest default that cannot abort at collection —
a model with no price row would stop the sweep before it spent anything (ADR 0007).
Preview it with --dry-run and narrow it with xharness_matrix before a first paid run.
Four ini keys, paths relative to pytest's rootdir:
| Key | Default | Purpose |
|---|---|---|
xharness_matrix |
(plugin default) | Project matrix: harness/model or harness/model/effort entries every case sweeps unless it sets models= |
xharness_skills_dir |
skills |
Directory holding <skill>/evals/ trees |
xharness_cache_dir |
.xharness_eval_cache |
The git-ignored root for build workspaces, results and the report (ADR 0032) |
xharness_skill_ignore |
(none) | gitignore-style patterns for skill files that are not decision surface; a bare pattern applies to every skill, <skill>: <pattern> to the skills matching the selector (ADR 0026) |
xharness_report_design_tokens |
bundled | design tokens JSON that themes report/report.html (flag: --xharness-report-design-tokens FILE) |
xharness_report_inline |
false |
embed every result, log and the tokens into report/report.html so it opens over file:// (flag: --xharness-report-inline) |
xharness_timeout_s |
600 |
Seconds one cell's CLI may run before it is killed (flag: --xharness-timeout SECONDS) |
xharness_prices |
(none) | Price rows that add to or override the bundled table: <model>: input=<usd/MTok> output=<usd/MTok> [cache_read=..] [cache_write=..] [cache_write_1h=..] (ADR 0030) |
[tool.pytest.ini_options]
xharness_matrix = [
"claude/claude-opus-5",
"claude/claude-sonnet-5",
"claude/claude-haiku-4-5-20251001",
"codex/gpt-5.6-luna",
"codex/gpt-5.6-terra",
"codex/gpt-5.6-sol",
]
An unpriced model stops the sweep at collection, before any spend. Add a price row to the same ini block, in USD per million tokens (ADR 0030):
xharness_prices = [
"gpt-5.6-luna: input=1.25 output=10.00 cache_read=0.125 cache_write=1.25",
]
How it works
flowchart LR
CASE["eval_*.py case"]
PLUG["plugin/<br/>collect, expand matrix"]
WS["model/workspace.py<br/>pristine copy"]
RUN["harness/<br/>ClaudeHarness | CodexHarness"]
LOG["session log<br/>this run's own"]
NORM["SessionLog.to_result<br/>RunResult"]
PRICE["runtime/pipeline.derive<br/>price, coverage, case"]
GRADE["case assertions"]
REP["report.json"]
CASE --> PLUG --> WS --> RUN --> LOG --> NORM --> PRICE --> GRADE --> REP
classDef new fill:#7c3aed,color:#fff
classDef data fill:#0f766e,color:#fff
classDef good fill:#047857,color:#fff
class CASE,PLUG,WS,RUN new
class LOG,NORM data
class PRICE,GRADE,REP good
One cell flows left to right: a case is expanded into cells, each cell gets a fresh workspace, the CLI runs, its own log is located and normalised, priced, graded, and reported. The hard part is the middle: the two CLIs need different contracts to tie a verdict to the right log. ARCHITECTURE.md explains both.
What it does not do
- It does not replay recorded runs. Every cell invokes the real CLI (ADR 0002).
- It does not give the agent a git repository. The workspace is a plain copy of the fixture, so skills that read git history are out of scope (ADR 0004).
- It does not mock either CLI, in tests or in evals.
- It does not price an unknown model as zero. It refuses to run (ADR 0007).
- It does not throttle providers.
-n Nruns N cells at once; each cell is isolated (own workspace, ownCODEX_HOME, own Claude session). If a provider rate-limits you,-n 2 --dist loadgroupkeeps each harness's cells on one worker (parallel across harnesses, serial within one).
Development
Build, test and release instructions live in CONTRIBUTING.md.
Read next
- docs/rollout.md: what a rollout leaves you, the
CaseOutputa grader is handed, everyRunResultfield it can assert on, the bundledcheck_*verifiers, and the goldens convention - docs/token-accounting.md: how
accumulative_billed_tokens(billed across turns) andpeak_context_tokens(the largest prompt) are derived from what each provider reports, with a worked session - ARCHITECTURE.md: why the two CLIs need different capture contracts, how pricing works, and the vocabulary the code uses.
- AGENTS.md: operating instructions and hard boundaries for agents.
- docs/adrs/index.md: the decision index, generated from the records (ADR 0047).
Metadata
Release files for pytest-xharness-eval 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pytest_xharness_eval-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Release files / pytest_xharness_eval-0.6.0-py3-none-any.whl
| Download URL | pytest_xharness_eval-0.6.0-py3-none-any.whl |
|---|---|
| Size | 732.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e3e44cd70d4d0a2695f921b0a689d6bd13298831b932dc54d3140334e2b3e1dd
|
|
BLAKE2b-256 checksum How to use checksums |
80c88aa1896231c7f8817dd8f6f3aa6604cd56833d0c8e8471396ebe835fd83e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|