Skip to main content

pytest-xharness-eval 🧪🤖

CICD Checks Build Status Coverage

GitHub commit activity GitHub open issues GitHub open pull requests

License Latest Release PyPI

pytest plugin for cross AI agent harness evaluation.

Write the eval once. Run it against every harness and every model.

What it does

A pytest plugin that runs the claude and codex CLIs headlessly against a fixture workspace, captures each run's own session log, prices it, and grades what the agent left behind. A skill opts in by adding an evals/ directory; pytest does the rest.

Every eval cell is live and costs money. There is no replay mode. Preview the spend with --dry-run before a sweep. The design rationale lives in ARCHITECTURE.md and the decision log in docs/adrs/; agents start at AGENTS.md.


Quickstart

  1. Install the plugin into the repository that holds your skills. The pytest11 entry point registers it; no conftest.py wiring is needed:

    uv add --dev pytest-xharness-eval
    
  2. Pin pytest's rootdir to the repository root, so the plugin finds skills/ from any argument path. An empty [tool.pytest.ini_options] table is enough:

    [tool.pytest.ini_options]
    
  3. Add an eval beside the skill. Files and functions both carry the eval_ prefix, as test_ does for pytest. Fixtures are seed workspaces copied fresh for every cell; every run output lands under one git-ignored cache root, never in the skills tree (ADR 0032):

    skills/<skill>/
      SKILL.md
      evals/
        eval_<suite>.py
        fixtures/<name>/
        goldens/<name>/                                  # optional known-good output (ADR 0046)
        treatments/<name>[__<harness>]/                  # optional overlays swept beside the control (ADR 0055)
    .xharness_eval_cache/
      build/                                             # per-cell workspaces
      pricing/prices-YYYYMMDD.toml                       # rates priced live at collection (ADR 0060)
      results/{skill}/{harness}/{model}[--{effort}][+{treatment}]/{run}/{session}/  # log.jsonl, result.json, history.json
      report/                                            # report.json + the aggregated microsite
    
    from pytest_xharness_eval import CaseOutput, evalcase
    from pytest_xharness_eval.verify import check_files_written, check_rollout
    
    @evalcase(task="...", skill="<skill>", fixture="<name>")
    def eval_<case>(output: CaseOutput) -> None:
        check_rollout(output)                     # real session, billed, priced
        check_files_written(output, "OUTPUT.md")  # this run is what produced it
        assert "the thing" in output.read("OUTPUT.md")
    

    task is what a user types after naming the skill, it never names the skill, a CLI, or where a SKILL.md lives. Each harness renders its own invocation around it: /<skill> <task> for claude, $<skill> <task> for codex (ADR 0044). The full grader surface, every field you can assert on, the bundled verifiers, and the goldens convention, is docs/rollout.md.

  4. Preview the matrix. Nothing is invoked:

    uv run pytest skills/<skill>/evals --dry-run
    
    xharness-eval: skills root = /repo/skills, cache = /repo/.xharness_eval_cache
    xharness-eval: matrix = plugin default (11 of 14 catalogued models, output rate below $50/MTok); a case's models= overrides it
    collected 11 items
    skills/<skill>/evals/eval_<case>.py sssssssssss
    
    ============================ agent eval report ============================
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-haiku-4-5-20251001]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-sonnet-5]
      ...
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-6.1-sol]
      total spend: $0.0000 across 11 cell(s)
      report: /repo/.xharness_eval_cache/report/report.json
    
  5. Run it live, with -v so every cell reports its verdict, USD, context, wall clock, turns, and tool calls as it lands. Add -n 2 to run cells in parallel. This spends money:

    uv run pytest skills/<skill>/evals -v
    
    skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] PASSED  est $0.5762 (harness $0.5773)  352,451 accumulative_billed_tokens  23,898 baseline_tokens  76.0s  9 turns  8 tools
    

    Read the status word as: this plugin's estimate from its price table (and the harness CLI's own figure, where it reports one), every billed token summed over all turns (the cached prefix is re-read each turn), the harness's own prompt on turn 1, wall clock, model calls, tool calls. Every estimate records the rates it used and where they came from (rates_applied).

    Each cell leaves its verbatim session log (log.jsonl), a normalised result.json with a per-turn ledger, and one history.json metrics record in its own results/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/ directory, no two cells share a file, so parallel workers never contend (ADR 0032). At session end the one combine step aggregates everything under results/, every skill, every run, into report/: report.json, the accumulated history.jsonl, and a browsable report.html with its glossary (XHARNESS-REPORT-GLOSSARY.md) beside it. Serve it with python3 -m http.server --directory .xharness_eval_cache and open /report/report.html; it fetches the JSON beside it.

run is a RunResult: session id, log path, token usage by tier, tool calls, files written, and USD cost. The reference case, with its assertions written as a tutorial, is eval_palette_mandate.py, the reference case kept beside the skill it grades in the consuming repository.


Narrow a run

The matrix is the spend dial. These options are the plugin's own; everything else is stock pytest (-k, -x, -m eval, node ids).

Option Effect Example
path One skill or all of them pytest skills/x/evals, pytest skills/*/evals
--harness <name> Only cells for that harness (claude or codex), repeatable pytest skills/x/evals --harness codex
--model <substring> Only cells whose model id contains the string, or one exact harness/model, repeatable pytest skills/x/evals --model opus
--effort <rung> Only cells at that reasoning rung, repeatable. Matches the resolved rung, so --effort max selects claude's max and codex's xhigh alike pytest skills/x/evals --effort max
--treatment <name> Only cells under that treatment, repeatable. control names the untreated cell (ADR 0055) pytest skills/x/evals --treatment control
-k <expr> Boolean slices over cell ids and case names (stock pytest) -k "opus or sol", -k "codex and not sol"
--xharness-timeout <s> Seconds one cell's CLI may run before it is killed (default 600). Raise it for the top effort rungs, which think for longer by design pytest skills/x/evals --xharness-timeout 1800
--xharness-keep-workspaces Leave each finished cell's build workspace in .xharness_eval_cache/build/ for inspection. By default it is removed once its evidence is captured and graded pytest skills/x/evals --xharness-keep-workspaces
--dry-run Enumerate cells and validate pricing, invoke nothing pytest skills/x/evals --dry-run
--collect-only -q List cell node ids (stock pytest) pytest --collect-only -q skills/x/evals

Do not run pytest skills from the root: it walks into every skill's scripts/ directory and collects their unit tests too. skills/*/evals is the full matrix.

A case overrides the matrix with @evalcase(..., models=["codex/gpt-5.6-sol"]).

The effort axis

A matrix entry may name a third component, the reasoning budget the CLI is asked for:

[tool.pytest.ini_options]
xharness_matrix = """
claude/claude-opus-5/low
claude/claude-opus-5/max
codex/gpt-5.6-sol/mid
"""

Each line is a separately graded cell, so one model at two rungs is the cost-versus-quality comparison the axis exists for. An entry with no third component is unchanged: it leaves the CLI on whatever default its own configuration gives it.

Each harness declares its own ladder, and three portable aliases name a position on it rather than a level:

You write Position On claude On codex
min first rung low low
mid middle rung high high
max last rung max max

Both shipped CLIs happen to declare the same five rungs (low, medium, high, xhigh, max), so the aliases resolve identically on each today. That is a fact about these two CLIs and not a rule: the ladder lives on the harness class, so a third CLI may declare any rungs it likes and the aliases keep working by position.

Resolution happens once, at collection, so a node id, an evidence directory and a report row all carry the rung that was actually sent. A rung no harness has (claude/claude-opus-5/minimal) stops the sweep at collection, before anything is spent. That check is not theoretical: minimal appears in codex-cli's own local enum, and a paid sweep found that no gpt-5.6 model accepts it — the CLI forwards it, the API answers 400, and the run exits having produced nothing (ADR 0049).

The treatment axis

A treatment is a directory of files copied over a case's fixture: an AGENTS.md, a CLAUDE.md, anything the agent should find in its workspace. It answers "does this change to the agent's standing instructions change what the same cell costs and how well it does?"

evals/treatments/cheap_eval_subagents/AGENTS.md          # every harness gets this
evals/treatments/cheap_eval_subagents__claude/CLAUDE.md  # claude also gets this: "@AGENTS.md"

Name it on the case, or for every case with the xharness_treatments ini key:

@evalcase(task=TASK, skill=SKILL, fixture=FIXTURE, treatments=["cheap_eval_subagents"])

The axis is opt-in, and every treatment is swept beside its control, the same cell with no treatment. So the case above collects claude/claude-sonnet-5 and claude/claude-sonnet-5+cheap_eval_subagents as two separately graded cells. --treatment control or --treatment cheap_eval_subagents narrows to one arm.

The __<harness> directory exists because the CLIs read different files: codex reads AGENTS.md, claude reads CLAUDE.md. Each cell reads its workspace's own file and nothing above it. A treatment with no files for a harness the case sweeps stops collection, because that cell would be its control billed twice under a second name (ADR 0055).


Configuration

The matrix has three scopes, highest precedence first: a case's models=, the project's xharness_matrix ini key, and the plugin's bundled default. The report header names which one applied.

The bundled default is every model in the model catalogue whose output rate is below xharness_output_rate_limit (default 50 USD per million tokens). Today that keeps every catalogued model except the apex ones, Fable and Astra. Raise the limit to opt in to their cost, or name them in xharness_matrix or a case's models=, which the limit never filters (ADR 0058). Preview the default with --dry-run before a first paid run.

The model catalogue

Every supported model is listed in one config file, src/pytest_xharness_eval/derive/prices/models.toml, beside the dated price records. Each model carries three facts that every record stores (ADR 0057, ADR 0059):

Fact Example Meaning
line opus, sol The provider's own product line
family_tier 3 The model's role in its lineup, 1 the smallest, curated by hand. A number rather than a name, so it survives a lineup change, and frozen at release, so history stays comparable
released 2026-07-24 The release date

Today's tiers: 1 is Haiku and Luna, 2 is Sonnet and Terra, 3 is Opus and Sol, 4 is Fable and Astra. A tier is a role, never a price, which is what lets a report compare every tier 3 model across providers.

A matrix entry naming a model the catalogue does not list stops collection before anything is spent. Add a new model before a plugin release with one ini line:

xharness_models =
    codex/gpt-6.2-sol: line=sol tier=3 released=2026-10-20

Live pricing

A catalogued model with no bundled price row is priced live at collection. Its rates are looked up in LiteLLM's price feed by exact first-party id, and saved as a dated record under .xharness_eval_cache/pricing/. The sweep and every later replay read that record, so a run priced live is re-priced identically. A model the feed does not price either still stops collection, and nothing is guessed (ADR 0060). With every model priced, nothing is fetched.

The bundled records are the offline default. In this repository, make test first curates a new bundled snapshot (make prices) whenever models.toml lists a model they do not price.

The ini keys, paths relative to pytest's rootdir:

Key Default Purpose
xharness_matrix (plugin default) Project matrix: harness/model or harness/model/effort entries every case sweeps unless it sets models=
xharness_output_rate_limit 50 The plugin default matrix sweeps only catalogued models whose output rate, in USD per million tokens, is below this. Raise it to opt in to apex models (ADR 0058)
xharness_price_feed LiteLLM's feed Where a model with no price row is priced live from at collection: a URL or a local path (ADR 0060)
xharness_models (none) Model catalogue rows that add or correct a model before a plugin release: <harness>/<model>: line=<line> tier=<n> released=YYYY-MM-DD [effort=false]; effort=false marks a model whose CLI ignores a reasoning rung, so a matrix entry naming one is refused (ADR 0057, ADR 0063)
xharness_treatments (none) Treatment names under each suite's evals/treatments/, swept beside the untreated control unless a case sets treatments= (ADR 0055)
xharness_skills_dir skills Directory holding <skill>/evals/ trees
xharness_cache_dir .xharness_eval_cache The git-ignored root for build workspaces, results and the report (ADR 0032)
xharness_skill_ignore (none) gitignore-style patterns for skill files that are not decision surface; a bare pattern applies to every skill, <skill>: <pattern> to the skills matching the selector (ADR 0026)
xharness_report_design_tokens bundled design tokens JSON that themes report/report.html (flag: --xharness-report-design-tokens FILE)
xharness_report_inline false embed every result, log and the tokens into report/report.html so it opens over file:// (flag: --xharness-report-inline)
xharness_timeout_s 600 Seconds one cell's CLI may run before it is killed (flag: --xharness-timeout SECONDS). A killed run is captured and priced: it fails if its session was still active in the 5 minutes before the limit (it ran out of time), and errors if it had been silent longer (it stalled) (ADR 0064)
xharness_keep_workspaces false Leave each finished cell's build workspace in place for inspection instead of removing it once its evidence is captured and graded (flag: --xharness-keep-workspaces, ADR 0062)
xharness_prices (none) Price rows that add to or override the bundled price records: <harness>/<model>: input=<usd/MTok> output=<usd/MTok> [cache_read=..] [cache_write=..] [cache_write_1h=..] [long_context_above=<prompt tokens> long_input=.. long_output=.. [long_cache_read=..] [long_cache_write=..] [long_cache_write_1h=..]] [from=YYYY-MM-DD] [to=YYYY-MM-DD] (ADR 0030, ADR 0050, ADR 0051)
[tool.pytest.ini_options]
xharness_matrix = [
    "claude/claude-opus-5",
    "claude/claude-sonnet-5",
    "claude/claude-haiku-4-5-20251001",
    "codex/gpt-5.6-luna",
    "codex/gpt-5.6-terra",
    "codex/gpt-5.6-sol",
]

An unpriced model stops the sweep at collection, before any spend. Add a price row to the same ini block, naming the harness that runs the model, in USD per million tokens (ADR 0030, ADR 0050). A row with from=/to= applies only to runs stamped inside [from, to):

xharness_prices = [
    "codex/gpt-5.6-luna: input=0.20 output=1.20 cache_read=0.02 long_context_above=272000 long_input=0.40 long_output=1.80",
    "claude/claude-sonnet-5: input=3.00 output=15.00 from=2026-10-01",
]

Some providers bill a long prompt at higher rates: OpenAI prices a call whose prompt exceeds 272K tokens (cached input included) at the long-context rates for the whole call. State that tier with long_context_above and the long_* keys. Every call is priced on its own prompt, from the per-call ledger, so one long call in a run is billed correctly beside many short ones (ADR 0051).

The bundled rates are dated records, one derive/prices/prices-YYYYMMDD.toml per interval, curated from LiteLLM's feed with make prices. Every run is priced from the record in effect on the day it ran, so a replay reproduces the bill it had then. Each estimate's rates_applied names the record, its interval and any long-context tier, and long_context_calls counts the calls billed at that tier.


How it works

flowchart LR
    CASE["eval_*.py case"]
    PLUG["plugin/<br/>collect, expand matrix"]
    WS["model/workspace.py<br/>pristine copy"]
    RUN["harness/<br/>ClaudeHarness | CodexHarness"]
    LOG["session log<br/>this run's own"]
    NORM["SessionLog.to_result<br/>RunResult"]
    PRICE["runtime/pipeline.derive<br/>price, coverage, case"]
    GRADE["case assertions"]
    REP["report.json"]

    CASE --> PLUG --> WS --> RUN --> LOG --> NORM --> PRICE --> GRADE --> REP

    classDef new fill:#7c3aed,color:#fff
    classDef data fill:#0f766e,color:#fff
    classDef good fill:#047857,color:#fff
    class CASE,PLUG,WS,RUN new
    class LOG,NORM data
    class PRICE,GRADE,REP good

One cell flows left to right: a case is expanded into cells, each cell gets a fresh workspace, the CLI runs, its own log is located and normalised, priced, graded, and reported. The hard part is the middle: the two CLIs need different contracts to tie a verdict to the right log. ARCHITECTURE.md explains both.


What it does not do

  • It does not replay recorded runs. Every cell invokes the real CLI (ADR 0002).
  • It does not give the agent a git repository. The workspace is a plain copy of the fixture, so skills that read git history are out of scope (ADR 0004).
  • It does not mock either CLI, in tests or in evals.
  • It does not price an unknown model as zero. It refuses to run (ADR 0007).
  • It does not throttle providers. -n N runs N cells at once; each cell is isolated (own workspace, own CODEX_HOME, own Claude session). If a provider rate-limits you, -n 2 --dist loadgroup keeps each harness's cells on one worker (parallel across harnesses, serial within one).

Development

Build, test and release instructions live in CONTRIBUTING.md.


  • docs/rollout.md: what a rollout leaves you, the CaseOutput a grader is handed, every RunResult field it can assert on, the bundled check_* verifiers, and the goldens convention
  • docs/token-accounting.md: how accumulative_billed_tokens (billed across turns) and peak_context_tokens (the largest prompt) are derived from what each provider reports, with a worked session
  • ARCHITECTURE.md: why the two CLIs need different capture contracts, how pricing works, and the vocabulary the code uses.
  • AGENTS.md: operating instructions and hard boundaries for agents.
  • docs/adrs/index.md: the decision index, generated from the records (ADR 0047).

Metadata

Release files for pytest-xharness-eval 0.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for pytest-xharness-eval 0.10.0
File Interpreter ABI Platform
pytest_xharness_eval-0.10.0-py3-none-any.whl Python 3 none any Details

Release files / pytest_xharness_eval-0.10.0-py3-none-any.whl

Download URL pytest_xharness_eval-0.10.0-py3-none-any.whl
Size 777.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
63e0be0a328a1d2ceabab612a8fde5ff6481d9d950da1d09a98d00830cc0dbe1
BLAKE2b-256 checksum
How to use checksums
7ac0a408a41fd1361de5ddf17e7a8a4936e79387b529d6311a37abff76a0150f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.10.0 This release

1 release file

0.9.0

1 release file

0.8.0

1 release file

0.7.0

1 release file

0.6.0

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.0

1 release file

0.3.0

1 release file

0.2.0

1 release file

0.1.1

1 release file

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page