Learn to improve executable AI-agent harnesses from failure trajectories.
Harnyx is an agent-agnostic Python library implementing the method from the paper Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories (Shao et al., 2026, arXiv:2608.02276). It mines a batch of target-agent failures, has an engineer model propose an executable runtime patch, sandboxes it, reruns the same tasks, and accepts it only when a real outcome improvement is measured.
target agent rollout
-> batch failure packet
-> harness engineer (LLM or scripted)
-> parse <think>...</think><patch>...</patch>
-> AST validate + sandbox
-> rerun the frozen target on the same task identities
-> reward = patched performance - baseline performance
-> accept / reject (with regression protection) -> versioned harness
Harnyx is not an agent framework. It is the optimization loop around an agent: the harness is the editable object, not the model weights.
Why you'd use it
Your agent's model is frozen (an API, or a self-hosted checkpoint you are not going to fine-tune), but it still fails in recurring, systematic ways: wrong tool arguments, dropped state, protocol violations, repeated actions, no recovery after an error. Today you patch that by hand (prompts, guards, retry logic) with no evidence it actually helps.
Harnyx automates that loop:
- It edits the runtime, not the model. The editable surface is four lifecycle hooks around your frozen policy: initialize context, hint before a decision, validate/rewrite/veto an action before it runs, recover after feedback.
- An engineer model writes the edits from your own failure traces, so the fix targets the failures you actually have instead of a generic prompt.
- Every edit is measured, not judged. Harnyx reruns the same tasks with and without the patch and keeps it only if task success/reward improves.
- It refuses to make things worse. A patch that breaks a previously solved task is rejected (regression protection); every accepted patch is versioned and reversible.
- You get a tiny artifact to load at runtime. The output is a validated
code-hook patch (
harness-vN) you ship alongside your agent.
Use it when you can measure task outcomes and want your agent's success rate to improve automatically, without training the model.
Architecture
harnyx/
├── core/ Task, Trajectory, Agent, Harness (hook contract), Result
├── engineering/ HarnessPatch, parser, PatchValidator, HarnessEngineer, prompts
├── sandbox/ AST policy, LocalSandbox, SubprocessSandbox, limits
├── optimization/ FailurePacket, PatchGenerator, OutcomeReward, selection, optimizer
├── evaluation/ LocalEvaluator, HarnessR1BenchmarkAdapter, run reports
├── adapters/ Nyvero adapter (no Nyvero dependency)
├── llm/ Provider protocol, OpenAI-compatible client, scripted provider
├── demo/ Deterministic toy end-to-end
└── cli/ harnyx <command>
Research-only code (SFT/GRPO training, ablations, the random baseline engineer)
lives in research/ at the repository root and is not part of the
installed package.
The hard boundary: frozen policy (never edited) vs editable harness (only four hooks, only structured effects). A hook never executes an environment action itself; the host runtime interprets its return value.
What it actually does (a concrete example)
Say your support agent has tools lookup_order, issue_refund, and reply,
and its failure traces show a pattern: it calls issue_refund before a
successful lookup_order, so the refund errors and the agent loops.
The engineer reads those traces and writes one hook:
def hook(ctx, nb): # on_before_action
action = str((ctx.get("action") or {}).get("name") or "")
verified = (ctx.get("state") or {}).get("order_verified")
if action == "issue_refund" and not verified:
return {"kind": "block_and_prompt",
"message": "Look up the order before refunding."}
return None
Harnyx then reruns the same tasks with and without that hook. If success
improves and nothing that used to pass breaks, it is saved as harness-v1 and
becomes a tiny artifact you load next to your agent:
harness = ExecutableHarness.from_patch(patch, sandbox=LocalSandbox())
agent.run(task, harness=harness)
No model change, no new prompt framework - one reviewed, reversible code hook.
Use cases
Any agent with an objective success signal can be optimized this way. Concrete places people apply it:
- Coding agents. Stop
rm -rfoutside the workspace, run the tests before editing, and recover when a patch fails (examples/nyvero/). - Customer support / CRM. Block a refund until the order is verified, escalate after repeated failures, avoid duplicate tickets.
- SQL and analytics. Inspect the schema before querying, refuse destructive DDL, retry a broken query with the database error text in context.
- RAG and research. Retrieve before answering, cite what was read, recover from an empty retrieval instead of answering anyway.
- Browser automation. Select required options before submitting, avoid repeated clicks, recover from a stalled page.
- DevOps and SRE. Require a health check before redeploy, back off after repeated failures, block destructive shell commands.
- Data pipelines / ETL. Validate the schema before writing, avoid duplicate loads, retry with backoff.
- IT and helpdesk automation. Do not close a ticket before the resolution is confirmed, and recover from a failed tool call.
- Game and embodied agents. Stop repeating no-op actions and track multi-step state such as find, take, transform, place.
- Multi-tool workflows. Enforce tool ordering, block protocol violations, and add recovery when a tool errors.
- Evaluation teams. Compare harness variants against a fixed model with a measured metric, without retraining anything.
If you can write "done means this command exits 0" (or any objective score), Harnyx can improve against it.
Relationship to Harness-R1
Harnyx is an independent implementation of the method in
"Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure
Trajectories" (Shao et al., 2026). The paper/repository and Harnyx map
component-by-component in docs/reproduction.md.
- Reproduced faithfully: the four executable lifecycle hooks
(
on_init,make_pre_hint,on_before_action,on_post_step), thehook(ctx, nb)contract, theadd_code_hook-only patch protocol, the same-batch outcome reward (Eq. 1), theK = 8candidate group, AST/sandbox restrictions, the engineer prompt/response protocol, and the SFT + GRPO hyperparameters. - Deliberate deviations: benchmark runtimes are not bundled (use an adapter);
training delegates to TRL instead of vendoring Relax; only the released
code-hook protocol is implemented (not the legacy six-action DSL). See
docs/reproduction.md. - Harnyx extensions (opt-in): patch caching, failure clustering, and an
explicit regression suite. See
docs/research.md.
Reference implementation: https://github.com/DeepExperience/Harness-R1 (used for behavioural verification only; no source is copied).
Install
pip install harnyx # or: uv add harnyx
Requires Python ≥ 3.11. The core library has zero runtime dependencies.
Optional extra: harnyx[yaml] for YAML config files.
Verify the install (self-test, no model needed)
This is a deterministic offline self-test that proves the whole pipeline works (parse → validate → sandbox → rerun → reward → version). It is not how you use the library in your app.
harnyx run
# baseline success 0/1 -> patched success 1/1, engineer reward +1.000, harness-v1
Use it in your agent
You provide three things: your environment, your policy, and a task batch. Harnyx supplies the harness, the engineer loop, the sandbox, the outcome reward, and selection.
from harnyx import Action, StepResult, Task, HarnessedAgent
from harnyx import HarnessOptimizer, LocalEvaluator, LLMHarnessEngineer
from harnyx.llm.openai import OpenAICompatibleProvider
from harnyx.optimization.optimizer import OptimizationConfig
class MyEnv:
"""Runs one task and reports an objective outcome."""
name = "support"
def reset(self, task: Task) -> str: ... # -> initial observation
def step(self, action: Action) -> StepResult: ... # -> next observation (+done)
def state(self) -> dict: ... # exposed to hooks as ctx["state"]
def predicates(self) -> dict: ... # exposed as ctx["predicates"]
def success(self) -> bool: ... # your success criterion
def episode_reward(self) -> float: ... # 0/1, or a shaped score
def admissible_actions(self) -> list[str]: ...
class MyPolicy:
"""Your frozen LLM. It only decides; it never edits the harness."""
name = "my-policy"
def act(self, messages, *, step, admissible) -> Action:
# call your model here and return Action(name=..., arguments=...)
...
agent = HarnessedAgent(MyPolicy(), MyEnv(), benchmark="support", max_steps=12)
# The engineer is an LLM that reads failure traces and writes harness patches.
engineer = LLMHarnessEngineer(
OpenAICompatibleProvider(
base_url="http://localhost:8000/v1", # OpenAI / OpenRouter / vLLM / SGLang
model="your-engineer-model",
env_key="HARNYX_ENGINEER_API_KEY",
),
benchmark="support",
)
# Mine failures -> propose patches -> rerun the same tasks -> keep what improves.
result = HarnessOptimizer(
agent,
engineer,
evaluator=LocalEvaluator(benchmark="support"),
benchmark="support",
config=OptimizationConfig(candidates=8, iterations=3),
run_dir="runs",
).optimize(tasks) # tasks: list[Task] with stable .id values
print(result.baseline.mean_reward, "->", result.final.mean_reward)
Already have an agent loop? Wrap it instead of using HarnessedAgent: call
harness.on_init, make_pre_hint, on_before_action, and on_post_step at
your lifecycle points. See
docs/custom-agent.md.
Ship an accepted patch
The optimizer writes runs/<timestamp>/accepted_patch.json and a versioned
harness. Load the patch into your production agent. No model change, and the
patch is inert until you choose to install it:
import json
from harnyx import ExecutableHarness, LocalSandbox, HarnessPatch
raw = json.load(open("runs/<timestamp>/accepted_patch.json"))["patch"]
patch = HarnessPatch.from_dict(raw)
harness = ExecutableHarness.from_patch(patch, sandbox=LocalSandbox())
outcome = agent.run(task, harness=harness) # your agent, now guarded
When not to use it
- No measurable outcome. If you cannot compute a success/score per task there is nothing to optimize - the reward is the rerun delta, not a judge.
- No failures. If the agent already succeeds, there is no signal to learn from.
- The model or prompt can change freely. That may be simpler; Harnyx is for a fixed model where you want the runtime improved from evidence.
- Same-batch only. The reward is transductive (the same tasks before/after), exactly as in the paper - not a held-out generalization guarantee.
- Integration cost. You must implement an
Environment(and usually a tool-callingPolicy) for your domain.
Reproduction instructions
# Deterministic local end-to-end (works in CI, no model)
harnyx run
# Full local reproduction: toy loop + A-F ablations + security smoke
python examples/reproduction/run_local_reproduction.py --run-root runs/reproduction
# Full paper reproduction (WebShop/ALFWorld/DBBench) is documented, not runnable
# here because benchmark assets, Qwen3.5 models, and 8xH800 are unavailable.
The honest reproduction status (expected paper numbers vs what was actually
observed on the available hardware) is in docs/research.md.
No benchmark result in this repository is fabricated.
To reproduce the paper protocol on real benchmark runtimes, point
HarnessR1BenchmarkAdapter at the reference AgentBench runtimes (or any
harness-aware runner) and use the released reference task splits.
Security model
Generated harness code is untrusted. Every candidate passes:
parse -> schema validation -> AST policy -> sandbox smoke test -> execution
harnyx.sandbox.policy rejects imports, filesystem/network/subprocess access,
dynamic evaluation (eval/exec/compile), dunder/attribute escapes,
getattr/setattr/globals/locals, async/class/with/lambda/while/yield,
generators, raise, and benchmark-answer leakage (e.g. numbered ALFWorld
instances). Execution uses a restricted builtin set, a wall-clock timeout, and a
line budget; runtime failures degrade to no intervention. SubprocessSandbox
adds process isolation, a sanitized environment (no credentials), and
RLIMIT_AS/RLIMIT_CPU. See docs/sandbox.md.
No API key is ever hard-coded or logged; providers read them from environment variables.
Benchmarks and adapters
Harnyx ships a deterministic local evaluator and a benchmark adapter. The reference
paper's WebShop (500 tasks), ALFWorld (500 tasks), and DBBench (300 tasks) are
supported through HarnessR1BenchmarkAdapter but their runtimes are not bundled.
It is deliberately not a core dependency of Harnyx.
Limitations
- No paper-level benchmark reproduction here. The benchmark runtimes, model
weights, and 8×H800 hardware are not available in this environment. See
docs/research.mdfor the exact gap and likely causes. - Training stages expose the reference hyperparameters and a TRL launcher; the authors' Relax/LLaMA-Factory stack is not vendored.
- The legacy six-action DSL from the repository is not implemented; only the
released
add_code_hookprotocol is. - Reward is transductive (same tasks before/after), exactly as in the paper.
- The in-process sandbox is hardened but not a full OS sandbox; use
SubprocessSandboxwhen stronger isolation is required.
Citation
@misc{shao2026harnessr1learningeditexecutable,
title={Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories},
author={Shuai Shao and Kangning Zhang and Qingyao Li and Shijian Wang and Hao Wang and Wenxiang Jiao and Yuan Lu and Yi Guo and Weiwen Liu and Weinan Zhang},
year={2026},
eprint={2608.02276},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.02276},
}
License
Apache-2.0. Harnyx is an independent implementation; see NOTICE for
attribution of the Harness-R1 reference implementation, Life-Harness/AgentBench,
Relax, and the benchmark environments.
Metadata
Release files for harnyx 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| harnyx-0.1.1.tar.gz | 113.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| harnyx-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 198.9 kB
Release files / harnyx-0.1.1.tar.gz
| Download URL | harnyx-0.1.1.tar.gz |
|---|---|
| Size | 113.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e1a684f2ae26e2c147290606769982af2370f16848334cdd9d7f86484de01650
|
|
BLAKE2b-256 checksum How to use checksums |
24ce46703c8f3709c0600e3b06db0aace3bbf579c96352d35036eb6dbce8ddbc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / harnyx-0.1.1-py3-none-any.whl
| Download URL | harnyx-0.1.1-py3-none-any.whl |
|---|---|
| Size | 85.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fdba1185a036b58f9fdc31a6f4ca0ae3de87b08592bc3b9634387c3725252486
|
|
BLAKE2b-256 checksum How to use checksums |
e17486c4cd7c3c43fb71061674d4b86b5d2930a65d96a9e7baa8824a3ad8600f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|