Skip to main content

Harness Lab

A local-first experimentation platform for evaluating AI coding-agent harnesses.

Given the same software task and the same starting repository state, how do different agent harnesses and configurations compare in verified success, token usage, tool usage, latency, cost and behaviour? Harness Lab answers that question reproducibly:

                 TASK SUITE
                     │
            identical starting repo (one git commit)
                     │
          ┌──────────┼──────────┐
          ▼          ▼          ▼
       Harness A  Harness B  Harness C        codex · claude · fake · generic · yours
          │          │          │
      isolated    isolated    isolated        one git worktree per run
      worktree    worktree    worktree
          │          │          │
          └──────────┼──────────┘
                     ▼
              independent verifier            hidden tests + exit code / partial score
                     │
                     ▼
            normalized experiment             provider-neutral trace + metrics + diff
                     │
                     ▼
              Harness Lab UI                   matrix · run detail · compare · grow

Harness Lab is not an observability product. It is a scientific instrument for controlled comparisons: the agent never decides whether it succeeded, every run starts from the same commit, and every trace is stored in one provider-neutral schema.

Documentation: https://bilgin-kocak.github.io/harness-lab/

Install

Python 3.12+, git 2.20+, Linux or macOS (Windows through WSL).

pip install harnesslab            # or: pipx install harnesslab  /  uv tool install harnesslab
harnesslab doctor

Five minutes, no API keys

harnesslab run demo --variants fake-reference,fake-noop   # bundled suite, two fake agents
harnesslab serve                                          # http://127.0.0.1:8000
harnesslab init my-lab && cd my-lab                       # editable copy of the demo, sweeps, harnesses
harnesslab sweep run sweeps/demo-fake.yaml                # cheapest verified configuration
harnesslab grow run grow/demo-fake.yaml                   # grow a harness from failures
harnesslab run suites/demo/suite.yaml --variants claude-default   # a real harness (needs the claude CLI)

What it does

  • Runs the same task through different harnesses: Claude Code, the OpenAI Codex CLI, a deterministic fake agent, any CLI you wrap, or a Python adapter you write.
  • Decides success independently: hidden tests are injected after the agent exits and a verifier command produces the verdict.
  • Normalizes every trace: tool calls, commands, file changes, tokens, cost and LLM calls in one schema. Hidden reasoning is never persisted.
  • Records what reproduction needs: base commit, task hash, harness bundle hash, config hash, CLI version, model and environment on every run.
  • Compares configurations: sweeps search model × effort × toolset × compaction × action policy × harness bundle for the cheapest configuration that still passes.
  • Measures improvement, not just completion: improvement tasks give the agent several rounds to make a working repository better on a measured objective, with an optional in-loop evaluator, and report the improvement curve.
  • Measures safety next to success: every run gets risky-action findings, executed or blocked, and tasks can plant canary secrets and prompt-injection lures; a bundled sentinel hook can be A/B tested.
  • Grows a harness from its failures: the grow loop edits a harness bundle with an optimizer, keeps only edits that fix failures without regressing a held-out gate, and records the lineage (after Grow the Harness, Not the Context, arXiv 2609.26760).
  • Mines tasks from your repository: suite mine turns commits that changed code and tests into verifier-backed tasks (fail at the parent, pass at the commit), so comparisons run on dozens of tasks from your own code.
  • Proves each component earns its place: ablate tests every part of a harness bundle against its own absence, and every comparison carries paired, task-level confidence intervals.

Read more

Quickstart and concepts Start here
Claude Code, Codex, your own harness Guides
Task suites, sweeps, growing the harness Guides
CLI, YAML formats, metrics, Python API Reference
Security model and honest limits Explanations
Contributing and releasing Project

The same pages live in docs/ and can be read on GitHub.

Security warning

Harness Lab executes coding agents and agent-written code on your machine, isolated with git worktrees, not containers. Use trusted benchmark repositories and throwaway credentials. Read the security model first.

Developing

git clone https://github.com/bilgin-kocak/harness-lab && cd harness-lab
uv sync && uv run pytest && uv run ruff check src tests && uv run mkdocs build --strict

Changelog: CHANGELOG.md. License: MIT.

Metadata

Release files for harnesslab 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harnesslab 0.2.0
File Size Uploaded
harnesslab-0.2.0.tar.gz 270.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harnesslab 0.2.0
File Interpreter ABI Platform
harnesslab-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 540.5 kB

Release files / harnesslab-0.2.0.tar.gz

Download URL harnesslab-0.2.0.tar.gz
Size 270.6 kB
Tags Source
SHA-256 checksum
How to use checksums
85c89cd2d3cc59c2869a1ef337fbf094d41581fab89594930da5e2f51d29868f
BLAKE2b-256 checksum
How to use checksums
5698d08d93755dc44bf74e64444f88881e0c404b5e91c913d124b0e49bcc1d3e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / harnesslab-0.2.0-py3-none-any.whl

Download URL harnesslab-0.2.0-py3-none-any.whl
Size 269.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d53ef426581943b015b29b84da2e23a938a9eb50f0e5fc636ae5fecf7e43b2da
BLAKE2b-256 checksum
How to use checksums
4aab856bba841886e346335d045732ba78a4251a6dbcc748742b9f5ac73420ea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page