Skip to main content

Harnyx

Learn to improve executable AI-agent harnesses from failure trajectories.

Harnyx is a clean, agent-agnostic Python library implementing the Harness-R1 methodology: mine a batch of target-agent failures, have a harness engineer propose an executable runtime patch, sandbox it, rerun the same tasks, and accept it only when a real outcome improvement is measured.

target agent rollout
  -> batch failure packet
  -> harness engineer  (LLM or scripted)
  -> parse <think>...</think><patch>...</patch>
  -> AST validate + sandbox
  -> rerun the frozen target on the same task identities
  -> reward = patched performance - baseline performance
  -> accept / reject (with regression protection) -> versioned harness

Harnyx is not an agent framework. It is the optimization loop around an agent: the harness is the editable object, not the model weights.

Harnyx architecture


Demo

Harnyx CLI demo

Recorded on this repository's deterministic toy benchmark: baseline reward 0.0 -> patched reward 1.0, accepting harness-v1 (no model, GPU, or benchmark assets required). Video: assets/demo.mp4.


Relationship to Harness-R1

Harnyx is an independent implementation of the method in "Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories" (Shao et al., 2026). The paper/repository and Harnyx map component-by-component in docs/reproduction.md.

  • Reproduced faithfully: the four executable lifecycle hooks (on_init, make_pre_hint, on_before_action, on_post_step), the hook(ctx, nb) contract, the add_code_hook-only patch protocol, the same-batch outcome reward (Eq. 1), the K = 8 candidate group, AST/sandbox restrictions, the engineer prompt/response protocol, and the SFT + GRPO hyperparameters.
  • Deliberate deviations: benchmark runtimes are not bundled (use an adapter); training delegates to TRL instead of vendoring Relax; only the released code-hook protocol is implemented (not the legacy six-action DSL). See docs/reproduction.md.
  • Harnyx extensions (opt-in): patch caching, failure clustering, and an explicit regression suite. See docs/research.md.

Reference implementation: https://github.com/DeepExperience/Harness-R1 (used for behavioural verification only; no source is copied).

Architecture

harnyx/
├── core/          Task, Trajectory, Agent, Harness (hook contract), Result
├── engineering/   HarnessPatch, parser, PatchValidator, HarnessEngineer, prompts
├── sandbox/       AST policy, LocalSandbox, SubprocessSandbox, limits
├── optimization/  FailurePacket, PatchGenerator, OutcomeReward, selection, optimizer
├── evaluation/    LocalEvaluator, HarnessR1BenchmarkAdapter, run reports
├── adapters/      Nyvero adapter (no Nyvero dependency)
├── llm/           Provider protocol, OpenAI-compatible client, scripted provider
├── demo/          Deterministic toy end-to-end
└── cli/           harnyx <command>

Research-only code (SFT/GRPO training, ablations, the random baseline engineer) lives in research/ at the repository root and is not part of the installed package.

The hard boundary: frozen policy (never edited) vs editable harness (only four hooks, only structured effects). A hook never executes an environment action itself; the host runtime interprets its return value.

Installation

pip install -e .            # core, no dependencies
pip install -e ".[yaml]"    # + YAML configs
pip install -e ".[dev]"     # + pytest/ruff/mypy

Requires Python ≥ 3.11. Core has zero runtime dependencies.

Minimal example

The deterministic toy demo needs no model, GPU, or benchmark assets:

harnyx run
# baseline success 0/1, patched success 1/1, engineer reward +1.000

Or in Python:

from harnyx.demo.toy import run_demo

print(run_demo("runs", candidates=3))

from harnyx import (
    ExecutableHarness, FailurePacket, LocalEvaluator, LocalSandbox,
    ScriptedHarnessEngineer, HarnessOptimizer,
)

A full optimization over your own agent:

from harnyx import HarnessOptimizer, LocalEvaluator, LLMHarnessEngineer
from harnyx.llm.openai import OpenAICompatibleProvider
from harnyx.optimization.optimizer import OptimizationConfig

provider = OpenAICompatibleProvider(
    base_url="http://localhost:8000/v1",  # OpenAI / OpenRouter / vLLM / SGLang
    model="Qwen3.5-9B-engineer",
    env_key="HARNYX_ENGINEER_API_KEY",
)
engineer = LLMHarnessEngineer(provider, benchmark="mybench")

optimizer = HarnessOptimizer(
    agent,                       # any object with .run(task, harness, recorder)
    engineer,
    evaluator=LocalEvaluator(benchmark="mybench"),
    benchmark="mybench",
    config=OptimizationConfig(candidates=8, iterations=3),
    run_dir="runs",
)
result = optimizer.optimize(tasks)          # tasks: list[harnyx.Task]
print(result.final.mean_reward - result.baseline.mean_reward)

See docs/quickstart.md and docs/custom-agent.md.

Nyvero example

Nyvero is never a dependency of Harnyx core. The adapter targets a documented duck-typed contract (see docs/nyvero.md):

from harnyx.adapters.nyvero import NyveroAgentAdapter, NyveroHarnessAdapter

harnyx_agent = NyveroAgentAdapter(nyvero_agent, benchmark="nyvero")
harnyx_harness = NyveroHarnessAdapter(nyvero_harness)  # expose Nyvero's harness
result = harnyx_agent.run(task, harness=harnyx_harness)

Reproduction instructions

# Deterministic local end-to-end (works in CI, no model)
harnyx run

# Full local reproduction: toy loop + A-F ablations + security smoke
python examples/reproduction/run_local_reproduction.py --run-root runs/reproduction
# Full paper reproduction (WebShop/ALFWorld/DBBench) is documented, not runnable
# here because benchmark assets, Qwen3.5 models, and 8xH800 are unavailable.

The honest reproduction status — expected paper numbers vs what was actually observed on the available hardware — is in docs/research.md. No benchmark result in this repository is fabricated.

To reproduce the paper protocol on real benchmark runtimes, point HarnessR1BenchmarkAdapter at the reference AgentBench runtimes (or any harness-aware runner) and use the released reference task splits.

Security model

Generated harness code is untrusted. Every candidate passes:

parse -> schema validation -> AST policy -> sandbox smoke test -> execution

harnyx.sandbox.policy rejects imports, filesystem/network/subprocess access, dynamic evaluation (eval/exec/compile), dunder/attribute escapes, getattr/setattr/globals/locals, async/class/with/lambda/while/yield, generators, raise, and benchmark-answer leakage (e.g. numbered ALFWorld instances). Execution uses a restricted builtin set, a wall-clock timeout, and a line budget; runtime failures degrade to no intervention. SubprocessSandbox adds process isolation, a sanitized environment (no credentials), and RLIMIT_AS/RLIMIT_CPU. See docs/sandbox.md.

No API key is ever hard-coded or logged; providers read them from environment variables.

Benchmarks and adapters

Harnyx ships a deterministic local evaluator and a benchmark adapter. The reference paper's WebShop (500 tasks), ALFWorld (500 tasks), and DBBench (300 tasks) are supported through HarnessR1BenchmarkAdapter but their runtimes are not bundled. It is deliberately not a core dependency of Harnyx.

Limitations

  • No paper-level benchmark reproduction here. The benchmark runtimes, model weights, and 8×H800 hardware are not available in this environment. See docs/research.md for the exact gap and likely causes.
  • Training stages expose the reference hyperparameters and a TRL launcher; the authors' Relax/LLaMA-Factory stack is not vendored.
  • The legacy six-action DSL from the repository is not implemented; only the released add_code_hook protocol is.
  • Reward is transductive (same tasks before/after), exactly as in the paper.
  • The in-process sandbox is hardened but not a full OS sandbox; use SubprocessSandbox when stronger isolation is required.

Citation

@misc{shao2026harnessr1learningeditexecutable,
      title={Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories},
      author={Shuai Shao and Kangning Zhang and Qingyao Li and Shijian Wang and Hao Wang and Wenxiang Jiao and Yuan Lu and Yi Guo and Weiwen Liu and Weinan Zhang},
      year={2026},
      eprint={2608.02276},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.02276},
}

License

Apache-2.0. Harnyx is an independent implementation; see NOTICE for attribution of the Harness-R1 reference implementation, Life-Harness/AgentBench, Relax, and the benchmark environments.

Metadata

Release files for harnyx 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harnyx 0.1.0
File Size Uploaded
harnyx-0.1.0.tar.gz 99.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harnyx 0.1.0
File Interpreter ABI Platform
harnyx-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 179.6 kB

Release files / harnyx-0.1.0.tar.gz

Download URL harnyx-0.1.0.tar.gz
Size 99.0 kB
Tags Source
SHA-256 checksum
How to use checksums
dfbc58fbeb53d52ad4811ecda0328575d5bbc3cf5424973e5ff8a5e0af8bff36
BLAKE2b-256 checksum
How to use checksums
736044bfa7be20fc4ebe29dfd1acadd10090388802b6413bc44c1f33240da633
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / harnyx-0.1.0-py3-none-any.whl

Download URL harnyx-0.1.0-py3-none-any.whl
Size 80.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
db50bc07ba0b13fbc199593395cb74c083f6a8c01f3a59650818be668c2e1a45
BLAKE2b-256 checksum
How to use checksums
4d8089c09fa21e2978f849d5fc0874b605f80864c644e8c1fc01e699d7ea600f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page