harnix
A complete, model-agnostic LLM agent harness — loop, tools, context, memory, permissions, hooks, sub-agents, skills, MCP — with deterministic record/replay and trajectory evaluation built in.
An agent harness is the software around a language model that turns its completions into actions: the control loop, the tool layer, context management, memory, safety/permissions, lifecycle hooks, sub-agents, and observability. A recent survey formalizes it as a six-component tuple H = (E, T, C, S, L, V) and argues that the harness — not the model — is the primary determinant of agent reliability at scale.
harnix implements all six components in a small, zero-dependency
library — and adds the thing most harnesses lack: reproducibility and
measurability by construction. Record a run's model calls once and replay them
to get byte-identical trajectories, with no network and no cost. That makes agents
testable in CI like ordinary code, and makes agent evaluation honest and
comparable.
from harnix import Agent, Workspace, PermissionPolicy, Memory, standard_toolset
ws = Workspace("./sandbox")
agent = Agent(
model=my_model, # any LLM (see below)
tools=standard_toolset(ws, memory=Memory()), # file + shell + memory tools
permissions=PermissionPolicy.allow_all(), # blocks dangerous commands
)
traj = agent.run("Create a README and run the tests.")
print(traj.final_output, traj.num_steps, traj.usage.total_tokens)
The six harness components, and how we cover them
| Component | What harnix provides |
|
|---|---|---|
| E | Execution loop | Agent loop with budgets, retries, and a doom-loop guard (agent.py) |
| T | Tool registry | @tool (auto JSON-schema from type hints), built-in file/shell/memory tools, MCP tools |
| C | Context manager | WindowedContext: token-budget windowing + tool-output offloading + summary hook |
| S | State store | Memory (durable facts/notes, MEMORY.md-style digest) + Session save/resume |
| L | Lifecycle hooks | Hook pre/post-tool control points that block or modify calls; PermissionPolicy enforced as a hook |
| V | Evaluation interface | Trajectory + record/replay Cassette + scorers + process metrics |
Plus the differentiators that live across all of them: deterministic record/replay, trajectory-level metrics, and sub-agents & skills.
Installation
pip install harnix # core, zero dependencies
pip install "harnix[dev]" # + pytest, ruff, build tooling
Safety is structural, not advisory
Permissions are a policy consulted by a hook before every tool call, so a blocked call cannot be reasoned around by the model. Dangerous shell commands are screened regardless of policy:
from harnix import PermissionPolicy, PermissionRule, Decision
policy = PermissionPolicy(
rules=[
PermissionRule("read_file", Decision.ALLOW),
PermissionRule("bash", Decision.ASK), # ASK -> resolved by an approver
PermissionRule("*", Decision.DENY),
],
default=Decision.DENY,
)
# `rm -rf /`, fork bombs, `curl ... | sh`, disk wipes, etc. are DENIED outright.
Lifecycle hooks add deterministic control:
from harnix import Hook
from harnix.hooks import HookDecision
class NoSecrets(Hook):
def pre_tool_use(self, call):
if "AWS_SECRET" in str(call.arguments):
return HookDecision.block("refusing to handle secrets")
return HookDecision.proceed()
Plugging in a model
A model maps messages (+ tool schemas) to a response — that's the only contract. Ready-made adapters ship for the common providers and speak each HTTP API directly, so no vendor SDK is required — the core stays zero-dependency:
from harnix.providers import OpenAIModel, AnthropicModel, OllamaModel, VLLMModel
model = OpenAIModel("gpt-4o-mini") # $OPENAI_API_KEY
model = AnthropicModel("claude-sonnet-4-20250514") # $ANTHROPIC_API_KEY
model = OllamaModel("llama3.1") # local, no key
model = VLLMModel("meta-llama/Llama-3.1-8B-Instruct", # your own server
base_url="http://gpu:8000/v1")
agent = Agent(model=model, tools=[...])
Any OpenAI-compatible gateway (Together, Groq, Fireworks, OpenRouter, LM Studio,
DeepSeek, …) works via OpenAICompatibleModel, and a "provider:model" string
resolves to a configured model:
from harnix import load_model, OpenAICompatibleModel
model = load_model("anthropic:claude-sonnet-4-20250514", max_tokens=2048)
model = OpenAICompatibleModel("llama-3.1-70b",
base_url="https://api.groq.com/openai/v1",
api_key_env="GROQ_API_KEY")
For a custom or in-house client, wrap any callable with FunctionModel:
from harnix import FunctionModel
from harnix.types import Message
def call_llm(messages, tools):
return Message.assistant(content="...") # translate your client's output
agent = Agent(model=FunctionModel(call_llm), tools=[...])
The headline feature: deterministic record / replay
Wrap any model — a live provider above included — to capture its calls once, then replay them forever with no network and no cost:
from harnix import Cassette, RecordingModel, ReplayModel
from harnix.providers import OpenAIModel
model = OpenAIModel("gpt-4o-mini")
cassette = Cassette()
Agent(model=RecordingModel(model, cassette), tools=tools).run(task)
cassette.save("fixtures/task.json") # record once
cassette = Cassette.load("fixtures/task.json")
agent = Agent(model=ReplayModel(cassette, name=model.name), tools=tools)
agent.run(task) # replay forever: deterministic, offline, free
from harnix.assertions import assert_deterministic, assert_used_tool
def test_agent():
traj = run_with_replay()
assert_used_tool(traj, "run_tests")
assert traj.succeeded
def test_reproducible():
assert_deterministic(run_with_replay, times=3)
The differential test-bench: hold the model fixed, measure the harness
Deterministic replay pins the model's decisions. Once the model is fixed, you can change the harness — context policy, tool set, permissions, system prompt — and measure the effect offline and for free. This is the thing full-stack harnesses can't do: Harness-Bench establishes it with 5,194 live trajectories and real GPU spend; here it's a pure function of recorded runs. Three primitives build on the cassette:
1. Trajectory diff — a semantic, step-aligned diff (not a text diff): where two runs first diverge, which tool calls changed, and how the process metrics moved.
from harnix import diff_trajectories
d = diff_trajectories(baseline_traj, candidate_traj)
print(d.render()) # first divergence at step 1; tool errors 0 -> 1; ...
d.identical # False
2. Golden-trajectory regression — snapshot testing for agent behavior. Record a run once, commit it as a fixture, and fail CI with a readable diff when behavior drifts:
from harnix import assert_matches_golden
def test_agent_behavior_is_stable():
traj = run_with_replay()
assert_matches_golden(traj, "fixtures/agent.golden.json")
# first run bootstraps the golden; HARNIX_UPDATE_GOLDEN=1 re-records
3. Harness A/B ablation — run two harness configs over a suite (model pinned by the cassette) and get a verdict on which wins, and by how much, on accuracy and process cost:
from harnix import compare_harnesses
report = compare_harnesses(baseline_factory, candidate_factory, suite)
print(report.summary())
# verdict : candidate wins: +12.5% accuracy (not significant, p=0.5)
# accuracy : 75.0% -> 87.5% (delta +12.5%) | avg tokens 240 -> 190 (-50)
# significance: accuracy not significant (McNemar p=0.5, 95% CI [-0.1, +0.4])
# total_tokens: -50 significant (sign test p=0.008, 95% CI [-62, -38])
The comparison is paired (the same tasks run under both harnesses), so the verdict is backed by the right paired tests, not a point estimate: McNemar's exact test for the accuracy delta, the sign test for each process metric, and a seeded bootstrap CI (reproducible run to run). That's the difference between "candidate wins +12.5%" and knowing whether that lead is real or noise on a small suite — often the accuracy delta isn't significant while the efficiency gain clearly is. When several metrics are screened at once the p-values are corrected for multiple comparisons (Holm by default; also Bonferroni or Benjamini–Hochberg), across only the metrics that actually varied:
report = compare_harnesses(a, b, suite, correction="holm") # or "bonferroni" / "benjamini-hochberg" / "none"
Fixtures survive schema changes
Cassettes and goldens are committed and outlive the code that wrote them, so both carry a schema version and are auto-migrated on load — an old fixture keeps working after the format evolves, and a fixture newer than your installed version fails loudly instead of being misread. Upgrade committed fixtures in place with the CLI:
harnix migrate old.cassette.json # v1 -> v2, in place
harnix migrate run.json --dry-run # report the change without writing
When a candidate harness changes the request text (so the exact cassette key misses), enable resilient replay — it falls back to the nearest recorded request and reports the count of approximate hits, so the ablation stays honest:
model = ReplayModel(cassette, resilient=True) # counterfactual replay
Run the whole workflow: examples/harness_ablation.py.
See docs/STRATEGY.md for why this is the project's moat.
Golden regression in CI: the pytest plugin
Installing harnix registers a pytest plugin (via the pytest11 entry
point), so golden-trajectory regression testing needs no boilerplate — just the
golden fixture:
def test_agent_behaviour_is_stable(golden):
traj = run_agent_with_replay() # deterministic, offline
golden.check(traj) # -> goldens/test_agent_behaviour_is_stable.json
The first run bootstraps the fixture; later runs diff against it and fail with a
rendered TrajectoryDiff when behavior drifts. Flags:
pytest --update-golden # re-record goldens after an intended change
pytest --golden-require-exists # CI: a missing golden is a failure, not a bootstrap
Goldens live in a goldens/ folder beside each test file by default; set
harnix_golden_dir in your pytest config for a central location.
harnix studio: a local web UI
The test-bench is visual by nature — a diff, a divergence point, an A/B verdict with confidence intervals. studio is a local, offline web UI that renders all of it. Point it at a directory of trajectory / cassette / report JSON:
harnix serve ./runs # opens http://127.0.0.1:8765 in your browser
python -m harnix.web ./runs --no-browser --port 8080 # equivalent
It has five views:
- Artifacts — every run in the directory, badged by kind.
- Trajectory — a step timeline: assistant text, tool calls (pretty-printed args), tool results (ok/error, latency), and the metrics header.
- Diff — pick two runs: the first divergence is highlighted, per-step changes are aligned, and metric deltas render as an inline-SVG chart.
- Report — an A/B ablation report with accuracy bars, a delta chart with 95% CI error bars, and a significance table (adjusted p-values, CIs, verdict).
- Live run — drive an agent against a provider and watch its steps stream in (Server-Sent Events); the recorded cassette + trajectory are saved into the served directory and appear in the sidebar. Defaults to an offline scripted provider so it works with zero setup — no API key, no network.
Same ethos as the rest of the library: zero runtime dependencies (stdlib
http.server backend, vanilla-JS frontend — no npm, no framework, no CDN, works
fully offline), and it ships inside the wheel.
Security: studio is a single-user local tool. It binds to
127.0.0.1and refuses non-loopbackHostheaders (a DNS-rebinding guard), so a web page you visit can't reach it. Don't bind it to a public interface: the live-run endpoint can spend API-key money and, with thefilestoolset, run a sandboxed shell.
Sub-agents, skills, memory, MCP
from harnix import SubAgent, SkillRegistry, MCPClient
# Delegate an isolated subtask to a child agent (fresh context, narrow tools).
delegate = SubAgent(model=model, tools=[search], name="researcher").as_tool()
# Progressive-disclosure skills: catalog in context, full body loaded on demand.
skills = SkillRegistry(); skills.load_dir("./skills")
agent = Agent(model=model, tools=[delegate, skills.as_tool()])
# Connect external MCP tool servers (dependency-free stdio client).
with MCPClient(["python", "my_mcp_server.py"]) as mcp:
agent = Agent(model=model, tools=mcp.list_tools())
Parallel tool execution
When a single step issues several independent (I/O-bound) tool calls — shell, HTTP, file reads — run them concurrently. Results are always returned in the original call order, so the recorded trajectory stays stable and replayable:
agent = Agent(model=model, tools=tools, parallel_tools=True, max_workers=8)
(Default is sequential; enable it per agent. Your tools and hooks should be thread-safe when you turn it on.)
Interfaces: the harness is headless
harnix is an engine, not a UI. There is no required frontend — you
drive it however you like, because everything an interface needs is in the
Trajectory and the Callback stream:
- Library / API —
agent.run(task)(the primary surface) - CLI —
harnix inspect|diff|report|demo - Web UI —
harnix serve ./runs(studio: a local, offline trace/diff/ablation explorer) - TUI —
harnix tui(a built-incursesterminal UI, stdlib-only) - REPL — see
examples/repl_frontend.py; a ~30-line stdin loop is a frontend - Web backend — call
agent.run()from a request handler and streamCallbackevents - Chat bots — wire
agent.run()to Slack/Telegram/Discord, like OpenHarness's Ohmo
The same headless engine backs all of them; swap the adapter, keep the harness.
The built-in TUI
┌──────────────────────────────────────────────────────────┐
│ you> What is 2 + 3? │
│ · add({'a': 2, 'b': 3}) │
│ ↳ [ok] 5 │
│ assistant> The answer is 5. │
│ [completed] 2 steps · 8 tokens │
│ │
│ ───────────────────────────────────────────────── ready │
│ you> ▋ │
└──────────────────────────────────────────────────────────┘
A scrolling transcript streams steps live as the agent runs (PgUp/PgDn to
scroll); type a task and press Enter; /quit to exit. Run it with the bundled
offline demo via harnix tui, or drive a real model from code:
from harnix import TUIApp, Agent
TUIApp(lambda: Agent(model=my_model, tools=my_tools)).run()
The rendering and transcript state are pure functions (tested without a
terminal); only the curses driver touches a TTY.
Evaluating across a task suite
from harnix import Suite, Task, includes, run_suite, Pricing
suite = Suite(name="arith", tasks=[
Task(id="add", prompt="What is 2 + 3?", scorer=includes("5")),
])
report = run_suite(lambda: build_agent(), suite, pricing=Pricing(0.003, 0.015))
print(report.summary()) # accuracy + process metrics (steps, tokens, cost, redundant calls)
Don't want to hand-enter rates? pricing_for("gpt-4o-mini") returns an
approximate Pricing from a built-in table, so cost_usd populates automatically
in harnix inspect and the studio UI (inferred from a run's recorded
model). Prices are approximate and dated — override with your own Pricing, or
extend the table via register_pricing(prefix, in_per_1m, out_per_1m).
What's in the box
| Module | Purpose |
|---|---|
agent |
the loop, Budget, retries, doom-loop guard, parallel tool execution, Callback tracing |
tools |
@tool decorator + ToolRegistry |
builtins |
file / shell / memory tools, standard_toolset |
workspace |
sandboxed filesystem root |
permissions |
PermissionPolicy, rules, dangerous-command screening |
hooks |
pre/post-tool hooks (block/modify), PermissionHook, TruncateOutputHook |
context |
WindowedContext token budgeting + offloading |
memory |
Memory + Session persistence |
skills |
SkillRegistry progressive disclosure |
subagent |
SubAgent delegation tool |
mcp |
dependency-free MCP stdio client |
model |
Model + Scripted/Function/Recording/Replay adapters |
providers |
live adapters: OpenAI, Anthropic, Ollama, vLLM, any OpenAI-compatible endpoint (stdlib HTTP, no SDK) |
cassette |
content-addressed record/replay store (+ resilient nearest-request replay) |
diff |
semantic, step-aligned TrajectoryDiff between two runs |
regression |
golden-trajectory snapshot testing (assert_matches_golden) |
compare |
harness A/B ablation (compare_harnesses, ComparisonReport) |
pricing |
approximate per-model price presets so cost_usd populates automatically (pricing_for, register_pricing) |
stats |
paired significance: McNemar + sign test + bootstrap CIs + multiple-comparison correction |
migrations |
schema versioning + auto-migration for cassettes/goldens |
pytest_plugin |
golden fixture + --update-golden for CI regression testing |
eval / metrics / assertions |
tasks, scorers, reports, trajectory metrics, pytest helpers |
config |
layered HarnessConfig |
tui |
built-in curses terminal UI (a frontend over agent.run()) |
web |
studio: local, offline web UI — trajectory/diff/ablation explorer + live run (stdlib server, vanilla JS) |
cli |
`harnix version |
Run the full offline demo:
python examples/full_harness.py # workspace + tools + permissions + hooks + memory + replay
How this relates to other harnesses
Full-stack harnesses like OpenHands
and HKUDS/OpenHarness are excellent at
capability (rich tools, memory, multi-agent), but ship no record/replay,
deterministic reproducibility, or trajectory-level evaluation.
Harness-Bench shows the harness alone drives
substantial, model-dependent swings in task completion — which is exactly why
being able to reproduce and measure a run matters. harnix aims to be a
complete harness where that reproducibility and measurability are first-class,
while staying small enough (zero dependencies) to read in an afternoon.
See docs/RESEARCH.md for the full survey, the formal
taxonomy, design rationale, and an honest self-critique; and
docs/research-paper.md for the accompanying paper,
"Deterministic Replay and Trajectory-Level Evaluation for LLM Agent Harnesses."
Development
pip install -e ".[dev]"
pytest # 118 tests
ruff check src tests
python experiments/run_experiments.py # reproduce the paper's numbers
License
MIT — see LICENSE.
Metadata
Release files for harnix 0.9.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| harnix-0.9.0.tar.gz | 154.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| harnix-0.9.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 269.4 kB
Release files / harnix-0.9.0.tar.gz
| Download URL | harnix-0.9.0.tar.gz |
|---|---|
| Size | 154.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b40d36eaf1a083fd9ece0e7716717448d7b0e46c4731a6bcf0f25e8e82f07e11
|
|
BLAKE2b-256 checksum How to use checksums |
3c2a1216b79b0e9346e0e3eeb53b758211dc9b29431c65d32df5d82c33f06adb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.13
|
Release files / harnix-0.9.0-py3-none-any.whl
| Download URL | harnix-0.9.0-py3-none-any.whl |
|---|---|
| Size | 114.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bd649058a4a8f558b3ccaa894f8c81c284517e2fbf5971522fe57127743cc890
|
|
BLAKE2b-256 checksum How to use checksums |
7f4874686090484b8edf3b599980bd8e74d9e49177e6bc4987edaf782812e8fc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.13
|