pytest-xharness-eval 🧪🤖
pytest plugin for cross AI agent harness evaluation.
Write the eval once. Run it against every harness and every model.
What it does
A pytest plugin that runs the claude and codex CLIs headlessly against a fixture
workspace, captures each run's own session log, prices it, and grades what the agent
left behind. A skill opts in by adding an evals/ directory; pytest does the rest.
Every eval cell is live and costs money. There is no replay mode. Preview the spend
with --dry-run before a sweep. The design rationale lives in
ARCHITECTURE.md and the decision log in
docs/adrs/; agents start at AGENTS.md.
Quickstart
-
Install the plugin into the repository that holds your skills. The
pytest11entry point registers it; noconftest.pywiring is needed:uv add --dev pytest-xharness-eval
-
Pin pytest's rootdir to the repository root, so the plugin finds
skills/from any argument path. An empty[tool.pytest.ini_options]table is enough:[tool.pytest.ini_options] -
Add an eval beside the skill. Files and functions both carry the
eval_prefix, astest_does for pytest. Fixtures are seed workspaces copied fresh for every cell; every run output lands under one git-ignored cache root, never in the skills tree (ADR 0032):skills/<skill>/ SKILL.md evals/ eval_<suite>.py fixtures/<name>/ goldens/<name>/ # optional known-good output (ADR 0046) treatments/<name>[__<harness>]/ # optional overlays swept beside the control (ADR 0055) .xharness_eval_cache/ build/ # per-cell workspaces pricing/prices-YYYYMMDD.toml # rates priced live at collection (ADR 0060) results/{skill}/{harness}/{model}[--{effort}][+{treatment}]/{run}/{session}/ # log.jsonl, result.json, history.json report/ # report.json + the aggregated micrositefrom pytest_xharness_eval import CaseOutput, evalcase from pytest_xharness_eval.verify import check_files_written, check_rollout @evalcase(task="...", skill="<skill>", fixture="<name>") def eval_<case>(output: CaseOutput) -> None: check_rollout(output) # real session, billed, priced check_files_written(output, "OUTPUT.md") # this run is what produced it assert "the thing" in output.read("OUTPUT.md")
taskis what a user types after naming the skill, it never names the skill, a CLI, or where aSKILL.mdlives. Each harness renders its own invocation around it:/<skill> <task>forclaude,$<skill> <task>forcodex(ADR 0044). The full grader surface, every field you can assert on, the bundled verifiers, and the goldens convention, is docs/rollout.md. -
Preview the matrix. Nothing is invoked:
uv run pytest skills/<skill>/evals --dry-run
xharness-eval: skills root = /repo/skills, cache = /repo/.xharness_eval_cache xharness-eval: matrix = plugin default (11 of 14 catalogued models, output rate below $50/MTok); a case's models= overrides it collected 11 items skills/<skill>/evals/eval_<case>.py sssssssssss ============================ agent eval report ============================ dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-haiku-4-5-20251001] dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-sonnet-5] ... dry-run - skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-6.1-sol] total spend: $0.0000 across 11 cell(s) report: /repo/.xharness_eval_cache/report/report.json
-
Run it live, with
-vso every cell reports its verdict, USD, context, wall clock, turns, and tool calls as it lands. Add-n 2to run cells in parallel. This spends money:uv run pytest skills/<skill>/evals -v
skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] PASSED est $0.5762 (harness $0.5773) 352,451 accumulative_billed_tokens 23,898 baseline_tokens 76.0s 9 turns 8 tools
Read the status word as: this plugin's estimate from its price table (and the harness CLI's own figure, where it reports one), every billed token summed over all turns (the cached prefix is re-read each turn), the harness's own prompt on turn 1, wall clock, model calls, tool calls. Every estimate records the rates it used and where they came from (
rates_applied).Each cell leaves its verbatim session log (
log.jsonl), a normalisedresult.jsonwith a per-turn ledger, and onehistory.jsonmetrics record in its ownresults/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/directory, no two cells share a file, so parallel workers never contend (ADR 0032). At session end the one combine step aggregates everything underresults/, every skill, every run, intoreport/:report.json, the accumulatedhistory.jsonl, and a browsablereport.htmlwith its glossary (XHARNESS-REPORT-GLOSSARY.md) beside it. Serve it withpython3 -m http.server --directory .xharness_eval_cacheand open/report/report.html; it fetches the JSON beside it.
run is a RunResult: session id, log path, token usage by tier, tool calls, files
written, and USD cost. The reference case, with its assertions written as a tutorial,
is
eval_palette_mandate.py, the reference case kept beside the skill it grades in the consuming repository.
Narrow a run
The matrix is the spend dial. These options are the plugin's own; everything else is
stock pytest (-k, -x, -m eval, node ids).
| Option | Effect | Example |
|---|---|---|
| path | One skill or all of them | pytest skills/x/evals, pytest skills/*/evals |
--harness <name> |
Only cells for that harness (claude or codex), repeatable |
pytest skills/x/evals --harness codex |
--model <substring> |
Only cells whose model id contains the string, or one exact harness/model, repeatable |
pytest skills/x/evals --model opus |
--effort <rung> |
Only cells at that reasoning rung, repeatable. Matches the resolved rung, so --effort max selects claude's max and codex's xhigh alike |
pytest skills/x/evals --effort max |
--treatment <name> |
Only cells under that treatment, repeatable. control names the untreated cell (ADR 0055) |
pytest skills/x/evals --treatment control |
-k <expr> |
Boolean slices over cell ids and case names (stock pytest) | -k "opus or sol", -k "codex and not sol" |
--xharness-timeout <s> |
Seconds one cell's CLI may run before it is killed (default 600). Raise it for the top effort rungs, which think for longer by design | pytest skills/x/evals --xharness-timeout 1800 |
--xharness-keep-workspaces |
Leave each finished cell's build workspace in .xharness_eval_cache/build/ for inspection. By default it is removed once its evidence is captured and graded |
pytest skills/x/evals --xharness-keep-workspaces |
--dry-run |
Enumerate cells and validate pricing, invoke nothing | pytest skills/x/evals --dry-run |
--collect-only -q |
List cell node ids (stock pytest) | pytest --collect-only -q skills/x/evals |
Do not run pytest skills from the root: it walks into every skill's scripts/
directory and collects their unit tests too. skills/*/evals is the full matrix.
A case overrides the matrix with @evalcase(..., models=["codex/gpt-5.6-sol"]).
The effort axis
A matrix entry may name a third component, the reasoning budget the CLI is asked for:
[tool.pytest.ini_options]
xharness_matrix = """
claude/claude-opus-5/low
claude/claude-opus-5/max
codex/gpt-5.6-sol/mid
"""
Each line is a separately graded cell, so one model at two rungs is the cost-versus-quality comparison the axis exists for. An entry with no third component is unchanged: it leaves the CLI on whatever default its own configuration gives it.
Each harness declares its own ladder, and three portable aliases name a position on it rather than a level:
| You write | Position | On claude |
On codex |
|---|---|---|---|
min |
first rung | low |
low |
mid |
middle rung | high |
high |
max |
last rung | max |
max |
Both shipped CLIs happen to declare the same five rungs (low, medium, high, xhigh, max),
so the aliases resolve identically on each today. That is a fact about these two CLIs and
not a rule: the ladder lives on the harness class, so a third CLI may declare any rungs it
likes and the aliases keep working by position.
Resolution happens once, at collection, so a node id, an evidence directory and a report
row all carry the rung that was actually sent. A rung no harness has
(claude/claude-opus-5/minimal) stops the sweep at collection, before anything is spent.
That check is not theoretical: minimal appears in codex-cli's own local enum, and a paid
sweep found that no gpt-5.6 model accepts it — the CLI forwards it, the API answers 400,
and the run exits having produced nothing (ADR 0049).
The treatment axis
A treatment is a directory of files copied over a case's fixture: an AGENTS.md, a
CLAUDE.md, anything the agent should find in its workspace. It answers "does this change
to the agent's standing instructions change what the same cell costs and how well it does?"
evals/treatments/cheap_eval_subagents/AGENTS.md # every harness gets this
evals/treatments/cheap_eval_subagents__claude/CLAUDE.md # claude also gets this: "@AGENTS.md"
Name it on the case, or for every case with the xharness_treatments ini key:
@evalcase(task=TASK, skill=SKILL, fixture=FIXTURE, treatments=["cheap_eval_subagents"])
The axis is opt-in, and every treatment is swept beside its control, the same cell with
no treatment. So the case above collects claude/claude-sonnet-5 and
claude/claude-sonnet-5+cheap_eval_subagents as two separately graded cells.
--treatment control or --treatment cheap_eval_subagents narrows to one arm.
The __<harness> directory exists because the CLIs read different files: codex reads
AGENTS.md, claude reads CLAUDE.md. Each cell reads its workspace's own file and nothing
above it. A treatment with no files for a harness the case sweeps stops collection, because
that cell would be its control billed twice under a second name (ADR 0055).
Configuration
The matrix has three scopes, highest precedence first: a case's models=, the
project's xharness_matrix ini key, and the plugin's bundled default. The report
header names which one applied.
The bundled default is every model in the model catalogue whose output rate is below
xharness_output_rate_limit (default 50 USD per million tokens). Today that keeps every
catalogued model except the apex ones, Fable and Astra. Raise the limit to opt in to their
cost, or name them in xharness_matrix or a case's models=, which the limit never filters
(ADR 0058). Preview the default with --dry-run before a first paid run.
The model catalogue
Every supported model is listed in one config file,
src/pytest_xharness_eval/derive/prices/models.toml,
beside the dated price records. Each model carries three facts that every record stores
(ADR 0057, ADR 0059):
| Fact | Example | Meaning |
|---|---|---|
line |
opus, sol |
The provider's own product line |
family_tier |
3 |
The model's role in its lineup, 1 the smallest, curated by hand. A number rather than a name, so it survives a lineup change, and frozen at release, so history stays comparable |
released |
2026-07-24 |
The release date |
Today's tiers: 1 is Haiku and Luna, 2 is Sonnet and Terra, 3 is Opus and Sol, 4 is Fable and Astra. A tier is a role, never a price, which is what lets a report compare every tier 3 model across providers.
A matrix entry naming a model the catalogue does not list stops collection before anything is spent. Add a new model before a plugin release with one ini line:
xharness_models =
codex/gpt-6.2-sol: line=sol tier=3 released=2026-10-20
Live pricing
A catalogued model with no bundled price row is priced live at collection. Its rates are
looked up in LiteLLM's price feed by exact first-party id, and saved as a dated record under
.xharness_eval_cache/pricing/. The sweep and every later replay read that record, so a run
priced live is re-priced identically. A model the feed does not price either still stops
collection, and nothing is guessed (ADR 0060). With every model priced, nothing is fetched.
The bundled records are the offline default. In this repository, make test first curates a
new bundled snapshot (make prices) whenever models.toml lists a model they do not price.
The ini keys, paths relative to pytest's rootdir:
| Key | Default | Purpose |
|---|---|---|
xharness_matrix |
(plugin default) | Project matrix: harness/model or harness/model/effort entries every case sweeps unless it sets models= |
xharness_output_rate_limit |
50 |
The plugin default matrix sweeps only catalogued models whose output rate, in USD per million tokens, is below this. Raise it to opt in to apex models (ADR 0058) |
xharness_price_feed |
LiteLLM's feed | Where a model with no price row is priced live from at collection: a URL or a local path (ADR 0060) |
xharness_models |
(none) | Model catalogue rows that add or correct a model before a plugin release: <harness>/<model>: line=<line> tier=<n> released=YYYY-MM-DD [effort=false]; effort=false marks a model whose CLI ignores a reasoning rung, so a matrix entry naming one is refused (ADR 0057, ADR 0063) |
xharness_treatments |
(none) | Treatment names under each suite's evals/treatments/, swept beside the untreated control unless a case sets treatments= (ADR 0055) |
xharness_skills_dir |
skills |
Directory holding <skill>/evals/ trees |
xharness_cache_dir |
.xharness_eval_cache |
The git-ignored root for build workspaces, results and the report (ADR 0032) |
xharness_skill_ignore |
(none) | gitignore-style patterns for skill files that are not decision surface; a bare pattern applies to every skill, <skill>: <pattern> to the skills matching the selector (ADR 0026) |
xharness_report_design_tokens |
bundled | design tokens JSON that themes report/report.html (flag: --xharness-report-design-tokens FILE) |
xharness_report_inline |
false |
embed every result, log and the tokens into report/report.html so it opens over file:// (flag: --xharness-report-inline) |
xharness_timeout_s |
600 |
Seconds one cell's CLI may run before it is killed (flag: --xharness-timeout SECONDS). A killed run is captured and priced: it fails if its session was still active in the 5 minutes before the limit (it ran out of time), and errors if it had been silent longer (it stalled) (ADR 0064) |
xharness_keep_workspaces |
false |
Leave each finished cell's build workspace in place for inspection instead of removing it once its evidence is captured and graded (flag: --xharness-keep-workspaces, ADR 0062) |
xharness_prices |
(none) | Price rows that add to or override the bundled price records: <harness>/<model>: input=<usd/MTok> output=<usd/MTok> [cache_read=..] [cache_write=..] [cache_write_1h=..] [long_context_above=<prompt tokens> long_input=.. long_output=.. [long_cache_read=..] [long_cache_write=..] [long_cache_write_1h=..]] [from=YYYY-MM-DD] [to=YYYY-MM-DD] (ADR 0030, ADR 0050, ADR 0051) |
[tool.pytest.ini_options]
xharness_matrix = [
"claude/claude-opus-5",
"claude/claude-sonnet-5",
"claude/claude-haiku-4-5-20251001",
"codex/gpt-5.6-luna",
"codex/gpt-5.6-terra",
"codex/gpt-5.6-sol",
]
An unpriced model stops the sweep at collection, before any spend. Add a price row
to the same ini block, naming the harness that runs the model, in USD per million
tokens (ADR 0030, ADR 0050). A row with from=/to= applies only to runs stamped
inside [from, to):
xharness_prices = [
"codex/gpt-5.6-luna: input=0.20 output=1.20 cache_read=0.02 long_context_above=272000 long_input=0.40 long_output=1.80",
"claude/claude-sonnet-5: input=3.00 output=15.00 from=2026-10-01",
]
Some providers bill a long prompt at higher rates: OpenAI prices a call whose prompt
exceeds 272K tokens (cached input included) at the long-context rates for the whole
call. State that tier with long_context_above and the long_* keys. Every call is
priced on its own prompt, from the per-call ledger, so one long call in a run is billed
correctly beside many short ones (ADR 0051).
The bundled rates are dated records, one derive/prices/prices-YYYYMMDD.toml per
interval, curated from LiteLLM's feed with make prices. Every run is priced from the
record in effect on the day it ran, so a replay reproduces the bill it had then. Each
estimate's rates_applied names the record, its interval and any long-context tier,
and long_context_calls counts the calls billed at that tier.
How it works
flowchart LR
CASE["eval_*.py case"]
PLUG["plugin/<br/>collect, expand matrix"]
WS["model/workspace.py<br/>pristine copy"]
RUN["harness/<br/>ClaudeHarness | CodexHarness"]
LOG["session log<br/>this run's own"]
NORM["SessionLog.to_result<br/>RunResult"]
PRICE["runtime/pipeline.derive<br/>price, coverage, case"]
GRADE["case assertions"]
REP["report.json"]
CASE --> PLUG --> WS --> RUN --> LOG --> NORM --> PRICE --> GRADE --> REP
classDef new fill:#7c3aed,color:#fff
classDef data fill:#0f766e,color:#fff
classDef good fill:#047857,color:#fff
class CASE,PLUG,WS,RUN new
class LOG,NORM data
class PRICE,GRADE,REP good
One cell flows left to right: a case is expanded into cells, each cell gets a fresh workspace, the CLI runs, its own log is located and normalised, priced, graded, and reported. The hard part is the middle: the two CLIs need different contracts to tie a verdict to the right log. ARCHITECTURE.md explains both.
What it does not do
- It does not replay recorded runs. Every cell invokes the real CLI (ADR 0002).
- It does not give the agent a git repository. The workspace is a plain copy of the fixture, so skills that read git history are out of scope (ADR 0004).
- It does not mock either CLI, in tests or in evals.
- It does not price an unknown model as zero. It refuses to run (ADR 0007).
- It does not throttle providers.
-n Nruns N cells at once; each cell is isolated (own workspace, ownCODEX_HOME, own Claude session). If a provider rate-limits you,-n 2 --dist loadgroupkeeps each harness's cells on one worker (parallel across harnesses, serial within one).
Development
Build, test and release instructions live in CONTRIBUTING.md.
Read next
- docs/rollout.md: what a rollout leaves you, the
CaseOutputa grader is handed, everyRunResultfield it can assert on, the bundledcheck_*verifiers, and the goldens convention - docs/token-accounting.md: how
accumulative_billed_tokens(billed across turns) andpeak_context_tokens(the largest prompt) are derived from what each provider reports, with a worked session - ARCHITECTURE.md: why the two CLIs need different capture contracts, how pricing works, and the vocabulary the code uses.
- AGENTS.md: operating instructions and hard boundaries for agents.
- docs/adrs/index.md: the decision index, generated from the records (ADR 0047).
Metadata
Release files for pytest-xharness-eval 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pytest_xharness_eval-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Release files / pytest_xharness_eval-0.10.0-py3-none-any.whl
| Download URL | pytest_xharness_eval-0.10.0-py3-none-any.whl |
|---|---|
| Size | 777.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
63e0be0a328a1d2ceabab612a8fde5ff6481d9d950da1d09a98d00830cc0dbe1
|
|
BLAKE2b-256 checksum How to use checksums |
7ac0a408a41fd1361de5ddf17e7a8a4936e79387b529d6311a37abff76a0150f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|