Skip to main content

pytest-xharness-eval 🧪🤖

CICD Checks Build Status Coverage

GitHub commit activity GitHub open issues GitHub open pull requests

License Latest Release PyPI

pytest plugin for cross AI agent harness evaluation.

Write the eval once. Run it against every harness and every model.

What it does

A pytest plugin that runs the claude and codex CLIs headlessly against a fixture workspace, captures each run's own session log, prices it, and grades what the agent left behind. A skill opts in by adding an evals/ directory; pytest does the rest.

Every eval cell is live and costs money. There is no replay mode. Preview the spend with --dry-run before a sweep. The design rationale lives in ARCHITECTURE.md and the decision log in docs/adrs/; agents start at AGENTS.md.


Quickstart

  1. Install the plugin into the repository that holds your skills. The pytest11 entry point registers it; no conftest.py wiring is needed:

    uv add --dev pytest-xharness-eval
    
  2. Pin pytest's rootdir to the repository root, so the plugin finds skills/ from any argument path. An empty [tool.pytest.ini_options] table is enough:

    [tool.pytest.ini_options]
    
  3. Add an eval beside the skill. Files and functions both carry the eval_ prefix, as test_ does for pytest. Fixtures are seed workspaces copied fresh for every cell; every run output lands under one git-ignored cache root, never in the skills tree (ADR 0032):

    skills/<skill>/
      SKILL.md
      evals/
        eval_<suite>.py
        fixtures/<name>/
        goldens/<name>/                                  # optional known-good output (ADR 0046)
    .xharness_eval_cache/
      build/                                             # per-cell workspaces
      results/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/  # log.jsonl, result.json, history.json
      report/                                            # report.json + the aggregated microsite
    
    from pytest_xharness_eval import CaseOutput, evalcase
    from pytest_xharness_eval.verify import check_files_written, check_rollout
    
    @evalcase(task="...", skill="<skill>", fixture="<name>")
    def eval_<case>(output: CaseOutput) -> None:
        check_rollout(output)                     # real session, billed, priced
        check_files_written(output, "OUTPUT.md")  # this run is what produced it
        assert "the thing" in output.read("OUTPUT.md")
    

    task is what a user types after naming the skill, it never names the skill, a CLI, or where a SKILL.md lives. Each harness renders its own invocation around it: /<skill> <task> for claude, $<skill> <task> for codex (ADR 0044). The full grader surface, every field you can assert on, the bundled verifiers, and the goldens convention, is docs/rollout.md.

  4. Preview the matrix. Nothing is invoked:

    uv run pytest skills/<skill>/evals --dry-run
    
    xharness-eval: skills root = /repo/skills, cache = /repo/.xharness_eval_cache
    xharness-eval: matrix = plugin default (2 entries); a case's models= overrides it
    collected 2 items
    skills/<skill>/evals/eval_<case>.py ss
    
    ============================ agent eval report ============================
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-sonnet-5]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-haiku-4-5-20251001]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-sol]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-luna]
      dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-5.6-terra]
      total spend: $0.0000 across 6 cell(s)
      report: /repo/.xharness_eval_cache/report/report.json
    
  5. Run it live, with -v so every cell reports its verdict, USD, context, wall clock, turns, and tool calls as it lands. Add -n 2 to run cells in parallel. This spends money:

    uv run pytest skills/<skill>/evals -v
    
    skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] PASSED  est $0.5762 (harness $0.5773)  352,451 accumulative_billed_tokens  23,898 baseline_tokens  76.0s  9 turns  8 tools
    

    Read the status word as: this plugin's estimate from its price table (and the harness CLI's own figure, where it reports one), every billed token summed over all turns (the cached prefix is re-read each turn), the harness's own prompt on turn 1, wall clock, model calls, tool calls. Every estimate records the rates it used and where they came from (rates_applied).

    Each cell leaves its verbatim session log (log.jsonl), a normalised result.json with a per-turn ledger, and one history.json metrics record in its own results/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/ directory, no two cells share a file, so parallel workers never contend (ADR 0032). At session end the one combine step aggregates everything under results/, every skill, every run, into report/: report.json, the accumulated history.jsonl, and a browsable report.html with its glossary (XHARNESS-REPORT-GLOSSARY.md) beside it. Serve it with python3 -m http.server --directory .xharness_eval_cache and open /report/report.html; it fetches the JSON beside it.

run is a RunResult: session id, log path, token usage by tier, tool calls, files written, and USD cost. The reference case, with its assertions written as a tutorial, is eval_palette_mandate.py, the reference case kept beside the skill it grades in the consuming repository.


Narrow a run

The matrix is the spend dial. These options are the plugin's own; everything else is stock pytest (-k, -x, -m eval, node ids).

Option Effect Example
path One skill or all of them pytest skills/x/evals, pytest skills/*/evals
--harness <name> Only cells for that harness (claude or codex), repeatable pytest skills/x/evals --harness codex
--model <substring> Only cells whose model id contains the string, or one exact harness/model, repeatable pytest skills/x/evals --model opus
--effort <rung> Only cells at that reasoning rung, repeatable. Matches the resolved rung, so --effort max selects claude's max and codex's xhigh alike pytest skills/x/evals --effort max
-k <expr> Boolean slices over cell ids and case names (stock pytest) -k "opus or sol", -k "codex and not sol"
--xharness-timeout <s> Seconds one cell's CLI may run before it is killed (default 600). Raise it for the top effort rungs, which think for longer by design pytest skills/x/evals --xharness-timeout 1800
--dry-run Enumerate cells and validate pricing, invoke nothing pytest skills/x/evals --dry-run
--collect-only -q List cell node ids (stock pytest) pytest --collect-only -q skills/x/evals

Do not run pytest skills from the root: it walks into every skill's scripts/ directory and collects their unit tests too. skills/*/evals is the full matrix.

A case overrides the matrix with @evalcase(..., models=["codex/gpt-5.6-sol"]).

The effort axis

A matrix entry may name a third component, the reasoning budget the CLI is asked for:

[tool.pytest.ini_options]
xharness_matrix = """
claude/claude-opus-5/low
claude/claude-opus-5/max
codex/gpt-5.6-sol/mid
"""

Each line is a separately graded cell, so one model at two rungs is the cost-versus-quality comparison the axis exists for. An entry with no third component is unchanged: it leaves the CLI on whatever default its own configuration gives it.

Each harness declares its own ladder, and three portable aliases name a position on it rather than a level:

You write Position On claude On codex
min first rung low low
mid middle rung high high
max last rung max max

Both shipped CLIs happen to declare the same five rungs (low, medium, high, xhigh, max), so the aliases resolve identically on each today. That is a fact about these two CLIs and not a rule: the ladder lives on the harness class, so a third CLI may declare any rungs it likes and the aliases keep working by position.

Resolution happens once, at collection, so a node id, an evidence directory and a report row all carry the rung that was actually sent. A rung no harness has (claude/claude-opus-5/minimal) stops the sweep at collection, before anything is spent. That check is not theoretical: minimal appears in codex-cli's own local enum, and a paid sweep found that no gpt-5.6 model accepts it — the CLI forwards it, the API answers 400, and the run exits having produced nothing (ADR 0049).


Configuration

The matrix has three scopes, highest precedence first: a case's models=, the project's xharness_matrix ini key, and the plugin's bundled default. The report header names which one applied.

The bundled default is every model the bundled price table carries: three per harness, claude/{claude-opus-5, claude-sonnet-5, claude-haiku-4-5-20251001} and codex/{gpt-5.6-sol, gpt-5.6-luna, gpt-5.6-terra}. An axis nobody narrowed means the whole axis, so this is deliberately the widest default that cannot abort at collection — a model with no price row would stop the sweep before it spent anything (ADR 0007). Preview it with --dry-run and narrow it with xharness_matrix before a first paid run.

Four ini keys, paths relative to pytest's rootdir:

Key Default Purpose
xharness_matrix (plugin default) Project matrix: harness/model or harness/model/effort entries every case sweeps unless it sets models=
xharness_skills_dir skills Directory holding <skill>/evals/ trees
xharness_cache_dir .xharness_eval_cache The git-ignored root for build workspaces, results and the report (ADR 0032)
xharness_skill_ignore (none) gitignore-style patterns for skill files that are not decision surface; a bare pattern applies to every skill, <skill>: <pattern> to the skills matching the selector (ADR 0026)
xharness_report_design_tokens bundled design tokens JSON that themes report/report.html (flag: --xharness-report-design-tokens FILE)
xharness_report_inline false embed every result, log and the tokens into report/report.html so it opens over file:// (flag: --xharness-report-inline)
xharness_timeout_s 600 Seconds one cell's CLI may run before it is killed (flag: --xharness-timeout SECONDS)
xharness_prices (none) Price rows that add to or override the bundled price records: <harness>/<model>: input=<usd/MTok> output=<usd/MTok> [cache_read=..] [cache_write=..] [cache_write_1h=..] [long_context_above=<prompt tokens> long_input=.. long_output=.. [long_cache_read=..] [long_cache_write=..] [long_cache_write_1h=..]] [from=YYYY-MM-DD] [to=YYYY-MM-DD] (ADR 0030, ADR 0050, ADR 0051)
[tool.pytest.ini_options]
xharness_matrix = [
    "claude/claude-opus-5",
    "claude/claude-sonnet-5",
    "claude/claude-haiku-4-5-20251001",
    "codex/gpt-5.6-luna",
    "codex/gpt-5.6-terra",
    "codex/gpt-5.6-sol",
]

An unpriced model stops the sweep at collection, before any spend. Add a price row to the same ini block, naming the harness that runs the model, in USD per million tokens (ADR 0030, ADR 0050). A row with from=/to= applies only to runs stamped inside [from, to):

xharness_prices = [
    "codex/gpt-5.6-luna: input=0.20 output=1.20 cache_read=0.02 long_context_above=272000 long_input=0.40 long_output=1.80",
    "claude/claude-sonnet-5: input=3.00 output=15.00 from=2026-10-01",
]

Some providers bill a long prompt at higher rates: OpenAI prices a call whose prompt exceeds 272K tokens (cached input included) at the long-context rates for the whole call. State that tier with long_context_above and the long_* keys. Every call is priced on its own prompt, from the per-call ledger, so one long call in a run is billed correctly beside many short ones (ADR 0051).

The bundled rates are dated records, one derive/prices/prices-YYYYMMDD.toml per interval, curated from LiteLLM's feed with make prices. Every run is priced from the record in effect on the day it ran, so a replay reproduces the bill it had then. Each estimate's rates_applied names the record, its interval and any long-context tier, and long_context_calls counts the calls billed at that tier.


How it works

flowchart LR
    CASE["eval_*.py case"]
    PLUG["plugin/<br/>collect, expand matrix"]
    WS["model/workspace.py<br/>pristine copy"]
    RUN["harness/<br/>ClaudeHarness | CodexHarness"]
    LOG["session log<br/>this run's own"]
    NORM["SessionLog.to_result<br/>RunResult"]
    PRICE["runtime/pipeline.derive<br/>price, coverage, case"]
    GRADE["case assertions"]
    REP["report.json"]

    CASE --> PLUG --> WS --> RUN --> LOG --> NORM --> PRICE --> GRADE --> REP

    classDef new fill:#7c3aed,color:#fff
    classDef data fill:#0f766e,color:#fff
    classDef good fill:#047857,color:#fff
    class CASE,PLUG,WS,RUN new
    class LOG,NORM data
    class PRICE,GRADE,REP good

One cell flows left to right: a case is expanded into cells, each cell gets a fresh workspace, the CLI runs, its own log is located and normalised, priced, graded, and reported. The hard part is the middle: the two CLIs need different contracts to tie a verdict to the right log. ARCHITECTURE.md explains both.


What it does not do

  • It does not replay recorded runs. Every cell invokes the real CLI (ADR 0002).
  • It does not give the agent a git repository. The workspace is a plain copy of the fixture, so skills that read git history are out of scope (ADR 0004).
  • It does not mock either CLI, in tests or in evals.
  • It does not price an unknown model as zero. It refuses to run (ADR 0007).
  • It does not throttle providers. -n N runs N cells at once; each cell is isolated (own workspace, own CODEX_HOME, own Claude session). If a provider rate-limits you, -n 2 --dist loadgroup keeps each harness's cells on one worker (parallel across harnesses, serial within one).

Development

Build, test and release instructions live in CONTRIBUTING.md.


  • docs/rollout.md: what a rollout leaves you, the CaseOutput a grader is handed, every RunResult field it can assert on, the bundled check_* verifiers, and the goldens convention
  • docs/token-accounting.md: how accumulative_billed_tokens (billed across turns) and peak_context_tokens (the largest prompt) are derived from what each provider reports, with a worked session
  • ARCHITECTURE.md: why the two CLIs need different capture contracts, how pricing works, and the vocabulary the code uses.
  • AGENTS.md: operating instructions and hard boundaries for agents.
  • docs/adrs/index.md: the decision index, generated from the records (ADR 0047).

Metadata

Release files for pytest-xharness-eval 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for pytest-xharness-eval 0.9.0
File Interpreter ABI Platform
pytest_xharness_eval-0.9.0-py3-none-any.whl Python 3 none any Details

Release files / pytest_xharness_eval-0.9.0-py3-none-any.whl

Download URL pytest_xharness_eval-0.9.0-py3-none-any.whl
Size 745.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a606157994fc693226f4bfcf3405ec1396dab8ee0932e3f42caa12ce4765da7a
BLAKE2b-256 checksum
How to use checksums
5d85d287759df996b1d6ef04ab5bac51b95f968c5efd3538a8d49255d74d3a87
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.10.0

1 release file

This release

0.9.0 This release

1 release file

0.8.0

1 release file

0.7.0

1 release file

0.6.0

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.0

1 release file

0.3.0

1 release file

0.2.0

1 release file

0.1.1

1 release file

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page