Harnyx
Learn to improve executable AI-agent harnesses from failure trajectories.
Harnyx is a clean, agent-agnostic Python library implementing the Harness-R1 methodology: mine a batch of target-agent failures, have a harness engineer propose an executable runtime patch, sandbox it, rerun the same tasks, and accept it only when a real outcome improvement is measured.
target agent rollout
-> batch failure packet
-> harness engineer (LLM or scripted)
-> parse <think>...</think><patch>...</patch>
-> AST validate + sandbox
-> rerun the frozen target on the same task identities
-> reward = patched performance - baseline performance
-> accept / reject (with regression protection) -> versioned harness
Harnyx is not an agent framework. It is the optimization loop around an agent: the harness is the editable object, not the model weights.
Demo
Recorded on this repository's deterministic toy benchmark: baseline reward 0.0
-> patched reward 1.0, accepting harness-v1 (no model, GPU, or benchmark
assets required). Video: assets/demo.mp4.
Relationship to Harness-R1
Harnyx is an independent implementation of the method in
"Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure
Trajectories" (Shao et al., 2026). The paper/repository and Harnyx map
component-by-component in docs/reproduction.md.
- Reproduced faithfully: the four executable lifecycle hooks
(
on_init,make_pre_hint,on_before_action,on_post_step), thehook(ctx, nb)contract, theadd_code_hook-only patch protocol, the same-batch outcome reward (Eq. 1), theK = 8candidate group, AST/sandbox restrictions, the engineer prompt/response protocol, and the SFT + GRPO hyperparameters. - Deliberate deviations: benchmark runtimes are not bundled (use an adapter);
training delegates to TRL instead of vendoring Relax; only the released
code-hook protocol is implemented (not the legacy six-action DSL). See
docs/reproduction.md. - Harnyx extensions (opt-in): patch caching, failure clustering, and an
explicit regression suite. See
docs/research.md.
Reference implementation: https://github.com/DeepExperience/Harness-R1 (used for behavioural verification only; no source is copied).
Architecture
harnyx/
├── core/ Task, Trajectory, Agent, Harness (hook contract), Result
├── engineering/ HarnessPatch, parser, PatchValidator, HarnessEngineer, prompts
├── sandbox/ AST policy, LocalSandbox, SubprocessSandbox, limits
├── optimization/ FailurePacket, PatchGenerator, OutcomeReward, selection, optimizer
├── evaluation/ LocalEvaluator, HarnessR1BenchmarkAdapter, run reports
├── adapters/ Nyvero adapter (no Nyvero dependency)
├── llm/ Provider protocol, OpenAI-compatible client, scripted provider
├── demo/ Deterministic toy end-to-end
└── cli/ harnyx <command>
Research-only code (SFT/GRPO training, ablations, the random baseline engineer)
lives in research/ at the repository root and is not part of the
installed package.
The hard boundary: frozen policy (never edited) vs editable harness (only four hooks, only structured effects). A hook never executes an environment action itself; the host runtime interprets its return value.
Installation
pip install -e . # core, no dependencies
pip install -e ".[yaml]" # + YAML configs
pip install -e ".[dev]" # + pytest/ruff/mypy
Requires Python ≥ 3.11. Core has zero runtime dependencies.
Minimal example
The deterministic toy demo needs no model, GPU, or benchmark assets:
harnyx run
# baseline success 0/1, patched success 1/1, engineer reward +1.000
Or in Python:
from harnyx.demo.toy import run_demo
print(run_demo("runs", candidates=3))
from harnyx import (
ExecutableHarness, FailurePacket, LocalEvaluator, LocalSandbox,
ScriptedHarnessEngineer, HarnessOptimizer,
)
A full optimization over your own agent:
from harnyx import HarnessOptimizer, LocalEvaluator, LLMHarnessEngineer
from harnyx.llm.openai import OpenAICompatibleProvider
from harnyx.optimization.optimizer import OptimizationConfig
provider = OpenAICompatibleProvider(
base_url="http://localhost:8000/v1", # OpenAI / OpenRouter / vLLM / SGLang
model="Qwen3.5-9B-engineer",
env_key="HARNYX_ENGINEER_API_KEY",
)
engineer = LLMHarnessEngineer(provider, benchmark="mybench")
optimizer = HarnessOptimizer(
agent, # any object with .run(task, harness, recorder)
engineer,
evaluator=LocalEvaluator(benchmark="mybench"),
benchmark="mybench",
config=OptimizationConfig(candidates=8, iterations=3),
run_dir="runs",
)
result = optimizer.optimize(tasks) # tasks: list[harnyx.Task]
print(result.final.mean_reward - result.baseline.mean_reward)
See docs/quickstart.md and
docs/custom-agent.md.
Nyvero example
Nyvero is never a dependency of Harnyx core. The adapter targets a documented
duck-typed contract (see docs/nyvero.md):
from harnyx.adapters.nyvero import NyveroAgentAdapter, NyveroHarnessAdapter
harnyx_agent = NyveroAgentAdapter(nyvero_agent, benchmark="nyvero")
harnyx_harness = NyveroHarnessAdapter(nyvero_harness) # expose Nyvero's harness
result = harnyx_agent.run(task, harness=harnyx_harness)
Reproduction instructions
# Deterministic local end-to-end (works in CI, no model)
harnyx run
# Full local reproduction: toy loop + A-F ablations + security smoke
python examples/reproduction/run_local_reproduction.py --run-root runs/reproduction
# Full paper reproduction (WebShop/ALFWorld/DBBench) is documented, not runnable
# here because benchmark assets, Qwen3.5 models, and 8xH800 are unavailable.
The honest reproduction status — expected paper numbers vs what was actually
observed on the available hardware — is in docs/research.md.
No benchmark result in this repository is fabricated.
To reproduce the paper protocol on real benchmark runtimes, point
HarnessR1BenchmarkAdapter at the reference AgentBench runtimes (or any
harness-aware runner) and use the released reference task splits.
Security model
Generated harness code is untrusted. Every candidate passes:
parse -> schema validation -> AST policy -> sandbox smoke test -> execution
harnyx.sandbox.policy rejects imports, filesystem/network/subprocess access,
dynamic evaluation (eval/exec/compile), dunder/attribute escapes,
getattr/setattr/globals/locals, async/class/with/lambda/while/yield,
generators, raise, and benchmark-answer leakage (e.g. numbered ALFWorld
instances). Execution uses a restricted builtin set, a wall-clock timeout, and a
line budget; runtime failures degrade to no intervention. SubprocessSandbox
adds process isolation, a sanitized environment (no credentials), and
RLIMIT_AS/RLIMIT_CPU. See docs/sandbox.md.
No API key is ever hard-coded or logged; providers read them from environment variables.
Benchmarks and adapters
Harnyx ships a deterministic local evaluator and a benchmark adapter. The reference
paper's WebShop (500 tasks), ALFWorld (500 tasks), and DBBench (300 tasks) are
supported through HarnessR1BenchmarkAdapter but their runtimes are not bundled.
It is deliberately not a core dependency of Harnyx.
Limitations
- No paper-level benchmark reproduction here. The benchmark runtimes, model
weights, and 8×H800 hardware are not available in this environment. See
docs/research.mdfor the exact gap and likely causes. - Training stages expose the reference hyperparameters and a TRL launcher; the authors' Relax/LLaMA-Factory stack is not vendored.
- The legacy six-action DSL from the repository is not implemented; only the
released
add_code_hookprotocol is. - Reward is transductive (same tasks before/after), exactly as in the paper.
- The in-process sandbox is hardened but not a full OS sandbox; use
SubprocessSandboxwhen stronger isolation is required.
Citation
@misc{shao2026harnessr1learningeditexecutable,
title={Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories},
author={Shuai Shao and Kangning Zhang and Qingyao Li and Shijian Wang and Hao Wang and Wenxiang Jiao and Yuan Lu and Yi Guo and Weiwen Liu and Weinan Zhang},
year={2026},
eprint={2608.02276},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.02276},
}
License
Apache-2.0. Harnyx is an independent implementation; see NOTICE for
attribution of the Harness-R1 reference implementation, Life-Harness/AgentBench,
Relax, and the benchmark environments.
Metadata
Release files for harnyx 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| harnyx-0.1.0.tar.gz | 99.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| harnyx-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 179.6 kB
Release files / harnyx-0.1.0.tar.gz
| Download URL | harnyx-0.1.0.tar.gz |
|---|---|
| Size | 99.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
dfbc58fbeb53d52ad4811ecda0328575d5bbc3cf5424973e5ff8a5e0af8bff36
|
|
BLAKE2b-256 checksum How to use checksums |
736044bfa7be20fc4ebe29dfd1acadd10090388802b6413bc44c1f33240da633
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / harnyx-0.1.0-py3-none-any.whl
| Download URL | harnyx-0.1.0-py3-none-any.whl |
|---|---|
| Size | 80.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
db50bc07ba0b13fbc199593395cb74c083f6a8c01f3a59650818be668c2e1a45
|
|
BLAKE2b-256 checksum How to use checksums |
4d8089c09fa21e2978f849d5fc0874b605f80864c644e8c1fc01e699d7ea600f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|