Skip to main content

harnix

A complete, model-agnostic LLM agent harness — loop, tools, context, memory, permissions, hooks, sub-agents, skills, MCP — with deterministic record/replay and trajectory evaluation built in.

CI Python License: MIT Dependencies: 0

An agent harness is the software around a language model that turns its completions into actions: the control loop, the tool layer, context management, memory, safety/permissions, lifecycle hooks, sub-agents, and observability. A recent survey formalizes it as a six-component tuple H = (E, T, C, S, L, V) and argues that the harness — not the model — is the primary determinant of agent reliability at scale.

harnix implements all six components in a small, zero-dependency library — and adds the thing most harnesses lack: reproducibility and measurability by construction. Record a run's model calls once and replay them to get byte-identical trajectories, with no network and no cost. That makes agents testable in CI like ordinary code, and makes agent evaluation honest and comparable.

from harnix import Agent, Workspace, PermissionPolicy, Memory, standard_toolset

ws = Workspace("./sandbox")
agent = Agent(
    model=my_model,                                  # any LLM (see below)
    tools=standard_toolset(ws, memory=Memory()),     # file + shell + memory tools
    permissions=PermissionPolicy.allow_all(),        # blocks dangerous commands
)
traj = agent.run("Create a README and run the tests.")
print(traj.final_output, traj.num_steps, traj.usage.total_tokens)

The six harness components, and how we cover them

Component What harnix provides
E Execution loop Agent loop with budgets, retries, and a doom-loop guard (agent.py)
T Tool registry @tool (auto JSON-schema from type hints), built-in file/shell/memory tools, MCP tools
C Context manager WindowedContext: token-budget windowing + tool-output offloading + summary hook
S State store Memory (durable facts/notes, MEMORY.md-style digest) + Session save/resume
L Lifecycle hooks Hook pre/post-tool control points that block or modify calls; PermissionPolicy enforced as a hook
V Evaluation interface Trajectory + record/replay Cassette + scorers + process metrics

Plus the differentiators that live across all of them: deterministic record/replay, trajectory-level metrics, and sub-agents & skills.

Installation

pip install harnix          # core, zero dependencies
pip install "harnix[dev]"   # + pytest, ruff, build tooling

Safety is structural, not advisory

Permissions are a policy consulted by a hook before every tool call, so a blocked call cannot be reasoned around by the model. Dangerous shell commands are screened regardless of policy:

from harnix import PermissionPolicy, PermissionRule, Decision

policy = PermissionPolicy(
    rules=[
        PermissionRule("read_file", Decision.ALLOW),
        PermissionRule("bash", Decision.ASK),     # ASK -> resolved by an approver
        PermissionRule("*", Decision.DENY),
    ],
    default=Decision.DENY,
)
# `rm -rf /`, fork bombs, `curl ... | sh`, disk wipes, etc. are DENIED outright.

Lifecycle hooks add deterministic control:

from harnix import Hook
from harnix.hooks import HookDecision

class NoSecrets(Hook):
    def pre_tool_use(self, call):
        if "AWS_SECRET" in str(call.arguments):
            return HookDecision.block("refusing to handle secrets")
        return HookDecision.proceed()

Plugging in a model

A model maps messages (+ tool schemas) to a response — that's the only contract. Ready-made adapters ship for the common providers and speak each HTTP API directly, so no vendor SDK is required — the core stays zero-dependency:

from harnix.providers import OpenAIModel, AnthropicModel, OllamaModel, VLLMModel

model = OpenAIModel("gpt-4o-mini")                       # $OPENAI_API_KEY
model = AnthropicModel("claude-sonnet-4-20250514")       # $ANTHROPIC_API_KEY
model = OllamaModel("llama3.1")                           # local, no key
model = VLLMModel("meta-llama/Llama-3.1-8B-Instruct",    # your own server
                  base_url="http://gpu:8000/v1")

agent = Agent(model=model, tools=[...])

Any OpenAI-compatible gateway (Together, Groq, Fireworks, OpenRouter, LM Studio, DeepSeek, …) works via OpenAICompatibleModel, and a "provider:model" string resolves to a configured model:

from harnix import load_model, OpenAICompatibleModel

model = load_model("anthropic:claude-sonnet-4-20250514", max_tokens=2048)
model = OpenAICompatibleModel("llama-3.1-70b",
                              base_url="https://api.groq.com/openai/v1",
                              api_key_env="GROQ_API_KEY")

For a custom or in-house client, wrap any callable with FunctionModel:

from harnix import FunctionModel
from harnix.types import Message

def call_llm(messages, tools):
    return Message.assistant(content="...")   # translate your client's output

agent = Agent(model=FunctionModel(call_llm), tools=[...])

The headline feature: deterministic record / replay

Wrap any model — a live provider above included — to capture its calls once, then replay them forever with no network and no cost:

from harnix import Cassette, RecordingModel, ReplayModel
from harnix.providers import OpenAIModel

model = OpenAIModel("gpt-4o-mini")

cassette = Cassette()
Agent(model=RecordingModel(model, cassette), tools=tools).run(task)
cassette.save("fixtures/task.json")          # record once

cassette = Cassette.load("fixtures/task.json")
agent = Agent(model=ReplayModel(cassette, name=model.name), tools=tools)
agent.run(task)                               # replay forever: deterministic, offline, free
from harnix.assertions import assert_deterministic, assert_used_tool

def test_agent():
    traj = run_with_replay()
    assert_used_tool(traj, "run_tests")
    assert traj.succeeded

def test_reproducible():
    assert_deterministic(run_with_replay, times=3)

The differential test-bench: hold the model fixed, measure the harness

Deterministic replay pins the model's decisions. Once the model is fixed, you can change the harness — context policy, tool set, permissions, system prompt — and measure the effect offline and for free. This is the thing full-stack harnesses can't do: Harness-Bench establishes it with 5,194 live trajectories and real GPU spend; here it's a pure function of recorded runs. Three primitives build on the cassette:

1. Trajectory diff — a semantic, step-aligned diff (not a text diff): where two runs first diverge, which tool calls changed, and how the process metrics moved.

from harnix import diff_trajectories

d = diff_trajectories(baseline_traj, candidate_traj)
print(d.render())          # first divergence at step 1; tool errors 0 -> 1; ...
d.identical                # False

2. Golden-trajectory regression — snapshot testing for agent behavior. Record a run once, commit it as a fixture, and fail CI with a readable diff when behavior drifts:

from harnix import assert_matches_golden

def test_agent_behavior_is_stable():
    traj = run_with_replay()
    assert_matches_golden(traj, "fixtures/agent.golden.json")
    # first run bootstraps the golden; HARNIX_UPDATE_GOLDEN=1 re-records

3. Harness A/B ablation — run two harness configs over a suite (model pinned by the cassette) and get a verdict on which wins, and by how much, on accuracy and process cost:

from harnix import compare_harnesses

report = compare_harnesses(baseline_factory, candidate_factory, suite)
print(report.summary())
#   verdict     : candidate wins: +12.5% accuracy (not significant, p=0.5)
#   accuracy    : 75.0% -> 87.5% (delta +12.5%)  |  avg tokens 240 -> 190 (-50)
#   significance: accuracy not significant (McNemar p=0.5, 95% CI [-0.1, +0.4])
#                 total_tokens: -50 significant (sign test p=0.008, 95% CI [-62, -38])

The comparison is paired (the same tasks run under both harnesses), so the verdict is backed by the right paired tests, not a point estimate: McNemar's exact test for the accuracy delta, the sign test for each process metric, and a seeded bootstrap CI (reproducible run to run). That's the difference between "candidate wins +12.5%" and knowing whether that lead is real or noise on a small suite — often the accuracy delta isn't significant while the efficiency gain clearly is. When several metrics are screened at once the p-values are corrected for multiple comparisons (Holm by default; also Bonferroni or Benjamini–Hochberg), across only the metrics that actually varied:

report = compare_harnesses(a, b, suite, correction="holm")  # or "bonferroni" / "benjamini-hochberg" / "none"

Fixtures survive schema changes

Cassettes and goldens are committed and outlive the code that wrote them, so both carry a schema version and are auto-migrated on load — an old fixture keeps working after the format evolves, and a fixture newer than your installed version fails loudly instead of being misread. Upgrade committed fixtures in place with the CLI:

harnix migrate old.cassette.json         # v1 -> v2, in place
harnix migrate run.json --dry-run        # report the change without writing

When a candidate harness changes the request text (so the exact cassette key misses), enable resilient replay — it falls back to the nearest recorded request and reports the count of approximate hits, so the ablation stays honest:

model = ReplayModel(cassette, resilient=True)   # counterfactual replay

Run the whole workflow: examples/harness_ablation.py. See docs/STRATEGY.md for why this is the project's moat.

Golden regression in CI: the pytest plugin

Installing harnix registers a pytest plugin (via the pytest11 entry point), so golden-trajectory regression testing needs no boilerplate — just the golden fixture:

def test_agent_behaviour_is_stable(golden):
    traj = run_agent_with_replay()      # deterministic, offline
    golden.check(traj)                  # -> goldens/test_agent_behaviour_is_stable.json

The first run bootstraps the fixture; later runs diff against it and fail with a rendered TrajectoryDiff when behavior drifts. Flags:

pytest --update-golden          # re-record goldens after an intended change
pytest --golden-require-exists  # CI: a missing golden is a failure, not a bootstrap

Goldens live in a goldens/ folder beside each test file by default; set harnix_golden_dir in your pytest config for a central location.

harnix studio: a local web UI

The test-bench is visual by nature — a diff, a divergence point, an A/B verdict with confidence intervals. studio is a local, offline web UI that renders all of it. Point it at a directory of trajectory / cassette / report JSON:

harnix serve ./runs        # opens http://127.0.0.1:8765 in your browser
python -m harnix.web ./runs --no-browser --port 8080   # equivalent

It has five views:

  • Artifacts — every run in the directory, badged by kind.
  • Trajectory — a step timeline: assistant text, tool calls (pretty-printed args), tool results (ok/error, latency), and the metrics header.
  • Diff — pick two runs: the first divergence is highlighted, per-step changes are aligned, and metric deltas render as an inline-SVG chart.
  • Report — an A/B ablation report with accuracy bars, a delta chart with 95% CI error bars, and a significance table (adjusted p-values, CIs, verdict).
  • Live run — drive an agent against a provider and watch its steps stream in (Server-Sent Events); the recorded cassette + trajectory are saved into the served directory and appear in the sidebar. Defaults to an offline scripted provider so it works with zero setup — no API key, no network.

Same ethos as the rest of the library: zero runtime dependencies (stdlib http.server backend, vanilla-JS frontend — no npm, no framework, no CDN, works fully offline), and it ships inside the wheel.

Security: studio is a single-user local tool. It binds to 127.0.0.1 and refuses non-loopback Host headers (a DNS-rebinding guard), so a web page you visit can't reach it. Don't bind it to a public interface: the live-run endpoint can spend API-key money and, with the files toolset, run a sandboxed shell.

Sub-agents, skills, memory, MCP

from harnix import SubAgent, SkillRegistry, MCPClient

# Delegate an isolated subtask to a child agent (fresh context, narrow tools).
delegate = SubAgent(model=model, tools=[search], name="researcher").as_tool()

# Progressive-disclosure skills: catalog in context, full body loaded on demand.
skills = SkillRegistry(); skills.load_dir("./skills")
agent = Agent(model=model, tools=[delegate, skills.as_tool()])

# Connect external MCP tool servers (dependency-free stdio client).
with MCPClient(["python", "my_mcp_server.py"]) as mcp:
    agent = Agent(model=model, tools=mcp.list_tools())

Parallel tool execution

When a single step issues several independent (I/O-bound) tool calls — shell, HTTP, file reads — run them concurrently. Results are always returned in the original call order, so the recorded trajectory stays stable and replayable:

agent = Agent(model=model, tools=tools, parallel_tools=True, max_workers=8)

(Default is sequential; enable it per agent. Your tools and hooks should be thread-safe when you turn it on.)

Interfaces: the harness is headless

harnix is an engine, not a UI. There is no required frontend — you drive it however you like, because everything an interface needs is in the Trajectory and the Callback stream:

  • Library / API — agent.run(task) (the primary surface)
  • CLI — harnix inspect|diff|report|demo
  • Web UI — harnix serve ./runs (studio: a local, offline trace/diff/ablation explorer)
  • TUI — harnix tui (a built-in curses terminal UI, stdlib-only)
  • REPL — see examples/repl_frontend.py; a ~30-line stdin loop is a frontend
  • Web backend — call agent.run() from a request handler and stream Callback events
  • Chat bots — wire agent.run() to Slack/Telegram/Discord, like OpenHarness's Ohmo

The same headless engine backs all of them; swap the adapter, keep the harness.

The built-in TUI

┌──────────────────────────────────────────────────────────┐
│ you> What is 2 + 3?                                        │
│   · add({'a': 2, 'b': 3})                                  │
│       ↳ [ok] 5                                             │
│ assistant> The answer is 5.                                │
│ [completed] 2 steps · 8 tokens                             │
│                                                            │
│ ───────────────────────────────────────────────── ready  │
│ you> ▋                                                     │
└──────────────────────────────────────────────────────────┘

A scrolling transcript streams steps live as the agent runs (PgUp/PgDn to scroll); type a task and press Enter; /quit to exit. Run it with the bundled offline demo via harnix tui, or drive a real model from code:

from harnix import TUIApp, Agent
TUIApp(lambda: Agent(model=my_model, tools=my_tools)).run()

The rendering and transcript state are pure functions (tested without a terminal); only the curses driver touches a TTY.

Evaluating across a task suite

from harnix import Suite, Task, includes, run_suite, Pricing

suite = Suite(name="arith", tasks=[
    Task(id="add", prompt="What is 2 + 3?", scorer=includes("5")),
])
report = run_suite(lambda: build_agent(), suite, pricing=Pricing(0.003, 0.015))
print(report.summary())     # accuracy + process metrics (steps, tokens, cost, redundant calls)

Don't want to hand-enter rates? pricing_for("gpt-4o-mini") returns an approximate Pricing from a built-in table, so cost_usd populates automatically in harnix inspect and the studio UI (inferred from a run's recorded model). Prices are approximate and dated — override with your own Pricing, or extend the table via register_pricing(prefix, in_per_1m, out_per_1m).

What's in the box

Module Purpose
agent the loop, Budget, retries, doom-loop guard, parallel tool execution, Callback tracing
tools @tool decorator + ToolRegistry
builtins file / shell / memory tools, standard_toolset
workspace sandboxed filesystem root
permissions PermissionPolicy, rules, dangerous-command screening
hooks pre/post-tool hooks (block/modify), PermissionHook, TruncateOutputHook
context WindowedContext token budgeting + offloading
memory Memory + Session persistence
skills SkillRegistry progressive disclosure
subagent SubAgent delegation tool
mcp dependency-free MCP stdio client
model Model + Scripted/Function/Recording/Replay adapters
providers live adapters: OpenAI, Anthropic, Ollama, vLLM, any OpenAI-compatible endpoint (stdlib HTTP, no SDK)
cassette content-addressed record/replay store (+ resilient nearest-request replay)
diff semantic, step-aligned TrajectoryDiff between two runs
regression golden-trajectory snapshot testing (assert_matches_golden)
compare harness A/B ablation (compare_harnesses, ComparisonReport)
pricing approximate per-model price presets so cost_usd populates automatically (pricing_for, register_pricing)
stats paired significance: McNemar + sign test + bootstrap CIs + multiple-comparison correction
migrations schema versioning + auto-migration for cassettes/goldens
pytest_plugin golden fixture + --update-golden for CI regression testing
eval / metrics / assertions tasks, scorers, reports, trajectory metrics, pytest helpers
config layered HarnessConfig
tui built-in curses terminal UI (a frontend over agent.run())
web studio: local, offline web UI — trajectory/diff/ablation explorer + live run (stdlib server, vanilla JS)
cli `harnix version

Run the full offline demo:

python examples/full_harness.py     # workspace + tools + permissions + hooks + memory + replay

How this relates to other harnesses

Full-stack harnesses like OpenHands and HKUDS/OpenHarness are excellent at capability (rich tools, memory, multi-agent), but ship no record/replay, deterministic reproducibility, or trajectory-level evaluation. Harness-Bench shows the harness alone drives substantial, model-dependent swings in task completion — which is exactly why being able to reproduce and measure a run matters. harnix aims to be a complete harness where that reproducibility and measurability are first-class, while staying small enough (zero dependencies) to read in an afternoon.

See docs/RESEARCH.md for the full survey, the formal taxonomy, design rationale, and an honest self-critique; and docs/research-paper.md for the accompanying paper, "Deterministic Replay and Trajectory-Level Evaluation for LLM Agent Harnesses."

Development

pip install -e ".[dev]"
pytest                                   # 118 tests
ruff check src tests
python experiments/run_experiments.py    # reproduce the paper's numbers

License

MIT — see LICENSE.

Metadata

Release files for harnix 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harnix 0.9.0
File Size Uploaded
harnix-0.9.0.tar.gz 154.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harnix 0.9.0
File Interpreter ABI Platform
harnix-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 269.4 kB

Release files / harnix-0.9.0.tar.gz

Download URL harnix-0.9.0.tar.gz
Size 154.6 kB
Tags Source
SHA-256 checksum
How to use checksums
b40d36eaf1a083fd9ece0e7716717448d7b0e46c4731a6bcf0f25e8e82f07e11
BLAKE2b-256 checksum
How to use checksums
3c2a1216b79b0e9346e0e3eeb53b758211dc9b29431c65d32df5d82c33f06adb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.13

Release files / harnix-0.9.0-py3-none-any.whl

Download URL harnix-0.9.0-py3-none-any.whl
Size 114.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bd649058a4a8f558b3ccaa894f8c81c284517e2fbf5971522fe57127743cc890
BLAKE2b-256 checksum
How to use checksums
7f4874686090484b8edf3b599980bd8e74d9e49177e6bc4987edaf782812e8fc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.13

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page