Skip to main content
Harnyx

Learn to improve executable AI-agent harnesses from failure trajectories.

Harnyx is an agent-agnostic Python library implementing the method from the paper Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories (Shao et al., 2026, arXiv:2608.02276). It mines a batch of target-agent failures, has an engineer model propose an executable runtime patch, sandboxes it, reruns the same tasks, and accepts it only when a real outcome improvement is measured.

target agent rollout
  -> batch failure packet
  -> harness engineer  (LLM or scripted)
  -> parse <think>...</think><patch>...</patch>
  -> AST validate + sandbox
  -> rerun the frozen target on the same task identities
  -> reward = patched performance - baseline performance
  -> accept / reject (with regression protection) -> versioned harness

Harnyx is not an agent framework. It is the optimization loop around an agent: the harness is the editable object, not the model weights.

Why you'd use it

Your agent's model is frozen (an API, or a self-hosted checkpoint you are not going to fine-tune), but it still fails in recurring, systematic ways: wrong tool arguments, dropped state, protocol violations, repeated actions, no recovery after an error. Today you patch that by hand (prompts, guards, retry logic) with no evidence it actually helps.

Harnyx automates that loop:

  • It edits the runtime, not the model. The editable surface is four lifecycle hooks around your frozen policy: initialize context, hint before a decision, validate/rewrite/veto an action before it runs, recover after feedback.
  • An engineer model writes the edits from your own failure traces, so the fix targets the failures you actually have instead of a generic prompt.
  • Every edit is measured, not judged. Harnyx reruns the same tasks with and without the patch and keeps it only if task success/reward improves.
  • It refuses to make things worse. A patch that breaks a previously solved task is rejected (regression protection); every accepted patch is versioned and reversible.
  • You get a tiny artifact to load at runtime. The output is a validated code-hook patch (harness-vN) you ship alongside your agent.

Use it when you can measure task outcomes and want your agent's success rate to improve automatically, without training the model.

Architecture

Harnyx architecture

harnyx/
├── core/          Task, Trajectory, Agent, Harness (hook contract), Result
├── engineering/   HarnessPatch, parser, PatchValidator, HarnessEngineer, prompts
├── sandbox/       AST policy, LocalSandbox, SubprocessSandbox, limits
├── optimization/  FailurePacket, PatchGenerator, OutcomeReward, selection, optimizer
├── evaluation/    LocalEvaluator, HarnessR1BenchmarkAdapter, run reports
├── adapters/      Nyvero adapter (no Nyvero dependency)
├── llm/           Provider protocol, OpenAI-compatible client, scripted provider
├── demo/          Deterministic toy end-to-end
└── cli/           harnyx <command>

Research-only code (SFT/GRPO training, ablations, the random baseline engineer) lives in research/ at the repository root and is not part of the installed package.

The hard boundary: frozen policy (never edited) vs editable harness (only four hooks, only structured effects). A hook never executes an environment action itself; the host runtime interprets its return value.

What it actually does (a concrete example)

Say your support agent has tools lookup_order, issue_refund, and reply, and its failure traces show a pattern: it calls issue_refund before a successful lookup_order, so the refund errors and the agent loops.

The engineer reads those traces and writes one hook:

def hook(ctx, nb):                          # on_before_action
    action = str((ctx.get("action") or {}).get("name") or "")
    verified = (ctx.get("state") or {}).get("order_verified")
    if action == "issue_refund" and not verified:
        return {"kind": "block_and_prompt",
                "message": "Look up the order before refunding."}
    return None

Harnyx then reruns the same tasks with and without that hook. If success improves and nothing that used to pass breaks, it is saved as harness-v1 and becomes a tiny artifact you load next to your agent:

harness = ExecutableHarness.from_patch(patch, sandbox=LocalSandbox())
agent.run(task, harness=harness)

No model change, no new prompt framework - one reviewed, reversible code hook.

Use cases

Any agent with an objective success signal can be optimized this way. Concrete places people apply it:

  • Coding agents. Stop rm -rf outside the workspace, run the tests before editing, and recover when a patch fails (examples/nyvero/).
  • Customer support / CRM. Block a refund until the order is verified, escalate after repeated failures, avoid duplicate tickets.
  • SQL and analytics. Inspect the schema before querying, refuse destructive DDL, retry a broken query with the database error text in context.
  • RAG and research. Retrieve before answering, cite what was read, recover from an empty retrieval instead of answering anyway.
  • Browser automation. Select required options before submitting, avoid repeated clicks, recover from a stalled page.
  • DevOps and SRE. Require a health check before redeploy, back off after repeated failures, block destructive shell commands.
  • Data pipelines / ETL. Validate the schema before writing, avoid duplicate loads, retry with backoff.
  • IT and helpdesk automation. Do not close a ticket before the resolution is confirmed, and recover from a failed tool call.
  • Game and embodied agents. Stop repeating no-op actions and track multi-step state such as find, take, transform, place.
  • Multi-tool workflows. Enforce tool ordering, block protocol violations, and add recovery when a tool errors.
  • Evaluation teams. Compare harness variants against a fixed model with a measured metric, without retraining anything.

If you can write "done means this command exits 0" (or any objective score), Harnyx can improve against it.


Relationship to Harness-R1

Harnyx is an independent implementation of the method in "Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories" (Shao et al., 2026). The paper/repository and Harnyx map component-by-component in docs/reproduction.md.

  • Reproduced faithfully: the four executable lifecycle hooks (on_init, make_pre_hint, on_before_action, on_post_step), the hook(ctx, nb) contract, the add_code_hook-only patch protocol, the same-batch outcome reward (Eq. 1), the K = 8 candidate group, AST/sandbox restrictions, the engineer prompt/response protocol, and the SFT + GRPO hyperparameters.
  • Deliberate deviations: benchmark runtimes are not bundled (use an adapter); training delegates to TRL instead of vendoring Relax; only the released code-hook protocol is implemented (not the legacy six-action DSL). See docs/reproduction.md.
  • Harnyx extensions (opt-in): patch caching, failure clustering, and an explicit regression suite. See docs/research.md.

Reference implementation: https://github.com/DeepExperience/Harness-R1 (used for behavioural verification only; no source is copied).

Install

pip install harnyx        # or: uv add harnyx

Requires Python ≥ 3.11. The core library has zero runtime dependencies. Optional extra: harnyx[yaml] for YAML config files.

Verify the install (self-test, no model needed)

This is a deterministic offline self-test that proves the whole pipeline works (parse → validate → sandbox → rerun → reward → version). It is not how you use the library in your app.

harnyx run
# baseline success 0/1  ->  patched success 1/1, engineer reward +1.000, harness-v1

Use it in your agent

You provide three things: your environment, your policy, and a task batch. Harnyx supplies the harness, the engineer loop, the sandbox, the outcome reward, and selection.

from harnyx import Action, StepResult, Task, HarnessedAgent
from harnyx import HarnessOptimizer, LocalEvaluator, LLMHarnessEngineer
from harnyx.llm.openai import OpenAICompatibleProvider
from harnyx.optimization.optimizer import OptimizationConfig

class MyEnv:
    """Runs one task and reports an objective outcome."""
    name = "support"
    def reset(self, task: Task) -> str: ...          # -> initial observation
    def step(self, action: Action) -> StepResult: ... # -> next observation (+done)
    def state(self) -> dict: ...                      # exposed to hooks as ctx["state"]
    def predicates(self) -> dict: ...                 # exposed as ctx["predicates"]
    def success(self) -> bool: ...                    # your success criterion
    def episode_reward(self) -> float: ...            # 0/1, or a shaped score
    def admissible_actions(self) -> list[str]: ...

class MyPolicy:
    """Your frozen LLM. It only decides; it never edits the harness."""
    name = "my-policy"
    def act(self, messages, *, step, admissible) -> Action:
        # call your model here and return Action(name=..., arguments=...)
        ...

agent = HarnessedAgent(MyPolicy(), MyEnv(), benchmark="support", max_steps=12)

# The engineer is an LLM that reads failure traces and writes harness patches.
engineer = LLMHarnessEngineer(
    OpenAICompatibleProvider(
        base_url="http://localhost:8000/v1",   # OpenAI / OpenRouter / vLLM / SGLang
        model="your-engineer-model",
        env_key="HARNYX_ENGINEER_API_KEY",
    ),
    benchmark="support",
)

# Mine failures -> propose patches -> rerun the same tasks -> keep what improves.
result = HarnessOptimizer(
    agent,
    engineer,
    evaluator=LocalEvaluator(benchmark="support"),
    benchmark="support",
    config=OptimizationConfig(candidates=8, iterations=3),
    run_dir="runs",
).optimize(tasks)          # tasks: list[Task] with stable .id values

print(result.baseline.mean_reward, "->", result.final.mean_reward)

Already have an agent loop? Wrap it instead of using HarnessedAgent: call harness.on_init, make_pre_hint, on_before_action, and on_post_step at your lifecycle points. See docs/custom-agent.md.

Ship an accepted patch

The optimizer writes runs/<timestamp>/accepted_patch.json and a versioned harness. Load the patch into your production agent. No model change, and the patch is inert until you choose to install it:

import json
from harnyx import ExecutableHarness, LocalSandbox, HarnessPatch

raw = json.load(open("runs/<timestamp>/accepted_patch.json"))["patch"]
patch = HarnessPatch.from_dict(raw)
harness = ExecutableHarness.from_patch(patch, sandbox=LocalSandbox())
outcome = agent.run(task, harness=harness)   # your agent, now guarded

When not to use it

  • No measurable outcome. If you cannot compute a success/score per task there is nothing to optimize - the reward is the rerun delta, not a judge.
  • No failures. If the agent already succeeds, there is no signal to learn from.
  • The model or prompt can change freely. That may be simpler; Harnyx is for a fixed model where you want the runtime improved from evidence.
  • Same-batch only. The reward is transductive (the same tasks before/after), exactly as in the paper - not a held-out generalization guarantee.
  • Integration cost. You must implement an Environment (and usually a tool-calling Policy) for your domain.

Reproduction instructions

# Deterministic local end-to-end (works in CI, no model)
harnyx run

# Full local reproduction: toy loop + A-F ablations + security smoke
python examples/reproduction/run_local_reproduction.py --run-root runs/reproduction
# Full paper reproduction (WebShop/ALFWorld/DBBench) is documented, not runnable
# here because benchmark assets, Qwen3.5 models, and 8xH800 are unavailable.

The honest reproduction status (expected paper numbers vs what was actually observed on the available hardware) is in docs/research.md. No benchmark result in this repository is fabricated.

To reproduce the paper protocol on real benchmark runtimes, point HarnessR1BenchmarkAdapter at the reference AgentBench runtimes (or any harness-aware runner) and use the released reference task splits.

Security model

Generated harness code is untrusted. Every candidate passes:

parse -> schema validation -> AST policy -> sandbox smoke test -> execution

harnyx.sandbox.policy rejects imports, filesystem/network/subprocess access, dynamic evaluation (eval/exec/compile), dunder/attribute escapes, getattr/setattr/globals/locals, async/class/with/lambda/while/yield, generators, raise, and benchmark-answer leakage (e.g. numbered ALFWorld instances). Execution uses a restricted builtin set, a wall-clock timeout, and a line budget; runtime failures degrade to no intervention. SubprocessSandbox adds process isolation, a sanitized environment (no credentials), and RLIMIT_AS/RLIMIT_CPU. See docs/sandbox.md.

No API key is ever hard-coded or logged; providers read them from environment variables.

Benchmarks and adapters

Harnyx ships a deterministic local evaluator and a benchmark adapter. The reference paper's WebShop (500 tasks), ALFWorld (500 tasks), and DBBench (300 tasks) are supported through HarnessR1BenchmarkAdapter but their runtimes are not bundled. It is deliberately not a core dependency of Harnyx.

Limitations

  • No paper-level benchmark reproduction here. The benchmark runtimes, model weights, and 8×H800 hardware are not available in this environment. See docs/research.md for the exact gap and likely causes.
  • Training stages expose the reference hyperparameters and a TRL launcher; the authors' Relax/LLaMA-Factory stack is not vendored.
  • The legacy six-action DSL from the repository is not implemented; only the released add_code_hook protocol is.
  • Reward is transductive (same tasks before/after), exactly as in the paper.
  • The in-process sandbox is hardened but not a full OS sandbox; use SubprocessSandbox when stronger isolation is required.

Citation

@misc{shao2026harnessr1learningeditexecutable,
      title={Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories},
      author={Shuai Shao and Kangning Zhang and Qingyao Li and Shijian Wang and Hao Wang and Wenxiang Jiao and Yuan Lu and Yi Guo and Weiwen Liu and Weinan Zhang},
      year={2026},
      eprint={2608.02276},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.02276},
}

License

Apache-2.0. Harnyx is an independent implementation; see NOTICE for attribution of the Harness-R1 reference implementation, Life-Harness/AgentBench, Relax, and the benchmark environments.

Metadata

Release files for harnyx 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harnyx 0.1.1
File Size Uploaded
harnyx-0.1.1.tar.gz 113.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harnyx 0.1.1
File Interpreter ABI Platform
harnyx-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 198.9 kB

Release files / harnyx-0.1.1.tar.gz

Download URL harnyx-0.1.1.tar.gz
Size 113.1 kB
Tags Source
SHA-256 checksum
How to use checksums
e1a684f2ae26e2c147290606769982af2370f16848334cdd9d7f86484de01650
BLAKE2b-256 checksum
How to use checksums
24ce46703c8f3709c0600e3b06db0aace3bbf579c96352d35036eb6dbce8ddbc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / harnyx-0.1.1-py3-none-any.whl

Download URL harnyx-0.1.1-py3-none-any.whl
Size 85.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fdba1185a036b58f9fdc31a6f4ca0ae3de87b08592bc3b9634387c3725252486
BLAKE2b-256 checksum
How to use checksums
e17486c4cd7c3c43fb71061674d4b86b5d2930a65d96a9e7baa8824a3ad8600f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page