Skip to main content

rlm-harness

A clean, reusable harness for building any task on top of DSPy's Recursive Language Model module (dspy.RLM).

RLMs (Zhang & Khattab, MIT, arXiv:2512.24601) let a model explore unbounded context by treating it as a variable in a sandboxed Python REPL and recursively calling sub-LLMs over it. DSPy's dspy.RLM is the first-party implementation (Khattab co-authored both DSPy and the RLM paper) — it works with existing Signatures and is optimizer-compatible (GEPA/MIPRO). This kit distills the boilerplate around it into one small, opinionated layer.

rlm-harness is domain-agnostic — anything dspy.RLM can do fits: multi-hop "deep research", an RSS-digest agent that posts to a webhook, structured extraction, detection authoring, you name it. Security happens to be the author's own first use of it, but it isn't the kit's scope.

Why this exists

Using dspy.RLM directly leaves you re-writing the same plumbing for every task: model/sub-model config, a retry+validation loop, a sandbox choice, observability. rlm-harness makes a task a declaration:

from rlm_harness import RLMConfig, RLMTask, configure
from rlm_harness.tools import make_schema_validator
from pydantic import BaseModel

class Article(BaseModel):
    title: str
    summary: str

class Summarize(RLMTask):
    signature = "document: str -> article: Article"
    output_field = "article"
    output_model = Article
    instructions = "Read the document and produce a title and a one-paragraph summary."
    tools = [make_schema_validator(Article)]

configure(RLMConfig.from_env())
article = Summarize().run(document=long_text)   # validated Article

The retry loop, pydantic validation, sandbox selection, and budget caps are all inherited.

Installation

pip install rlm-harness
# or with uv:
uv add rlm-harness

rlm-harness needs Python ≥ 3.11 and pulls in dspy + pydantic; extras are opt-in — observability (pip install "rlm-harness[observe]") and running on a Claude Pro/Max subscription instead of an API key (pip install "rlm-harness[subscription]"rlm_harness.ClaudeAgentLM, injected via configure(main_lm=…)). A live dspy.RLM run additionally needs model credentials (see the guide's Configuration) and a Deno sandbox — the logic and tests run without either. dspy requires Deno >=2.0.0,<3.0.0: brew install deno, or let dspy manage it with pip install "dspy[deno]".

What's in the box

  • Tasks as declarations. Subclass RLMTask — the retry+validation loop, sandbox selection, budget caps, and observability are inherited.
  • The whole trajectory, recorded. TraceRecorder writes main steps, every sub-LM call, and every tool call into one append-only JSONL stream — replayable and exportable as SFT/RL datasets (reward-free: scoring belongs to your trainer).
  • The recursion seat, interceptable. intercept_sub_lm traces every sub-LM escalation (plus optional deterministic validate/post-process); model_as_tool lets the main LM choose to consult another named model, in the trajectory.
  • Tools, the base/wrap way. Pydantic/JSON-Schema validators, an SSRF-guarded fetch_url, provider-agnostic web search, the generic model-as-tool core, a run_command seam over your isolated runner, an MCP client bridge, and skills-as-tools progressive disclosure.
  • Sandboxed by default. The pyodide/deno interpreter; the local interpreter is refused unless explicitly opted into; an opt-in Docker container interpreter for when the REPL itself needs real subprocesses.
  • Offline-testable. rlm_harness.testing drives the real dspy.RLM forward loop with no model, no Deno, no network.

Documentation — the guide

The deep documentation lives in rlm_harness/README.md:

Built with rlm-harness

Real projects using rlm-harness as their RLM scaffold:

  • ctx-distillery: distils an AI coding agent's session transcripts and memory store into a judgement-only distillation plan — what to prune, cross-reference, or promote into durable memory or a reusable Skill. It proposes; it writes nothing.
  • cve-reverser: reverses publicly disclosed CVEs from their patches into local-lab PoCs and Nuclei detection templates. A traced, trainable RLM harness.
  • diff-sentry: classifies GitHub changes (PRs, issues, pushes) for malicious intent — the diff is read as untrusted data in the sandboxed REPL, emitting evidence-backed benign / suspicious / malicious verdicts into a SIEM.
  • toolscout: an ATLAS-style rollout harness — a small planner progressively discovers a large MCP toolspace and computes over tool results as code, emitting reward-free trajectories + per-criterion facts for a downstream trainer.

Built something on rlm-harness? Open a PR to add it here.

Security note — the sandbox is the boundary

RLM executes model-written code. When that code processes untrusted scraped content, the interpreter choice is your attack surface. The default (pyodide/deno) is the sandboxed DSPy interpreter. The local interpreter runs code on the host and is refused unless you set allow_insecure_sandbox=True / RLM_ALLOW_INSECURE_SANDBOX=1. Don't.

The default sandbox is built by the kit (not handed straight to dspy) so it can pre-bind the JSON literals true/false/null to True/False/None in the REPL namespace — a JSON-trained instruct model otherwise writes SUBMIT({"ok": true}) and the REPL raises NameError: name 'true' is not defined, which the model tends to retry verbatim. Isolation is unchanged; RLMTask owns the teardown.

Develop

uv sync --group dev
uv run pytest          # logic tests (no live LLM needed)

Tests cover config parsing, the retry/validation engine, the sandbox guard, the tools, the sub-LM-hook/trace/replay/dataset layer, and a real-dspy.RLM construction check (dspy-bearing tests use DummyLM or skip if dspy is absent). A live run additionally needs real credentials and a Deno sandbox (brew install deno, or pip install "dspy[deno]" for dspy's managed binary; it requires Deno >=2.0.0,<3.0.0); examples/mini_run.py shows it. To drive the real forward loop offline (no model, no Deno), see the guide's Testing the forward path offline. See CLAUDE.md for invariants when modifying the kit.

Status

v1.3.0 — the current release. Scaffold + harness-engineering layer (sub-LM hook, skills-as-tools, trajectory recording, replay, dataset export, harness delegation on both the client and server side), plus fast-failing non-retryable LM errors and a filesystem/process layer: reading, searching, writing and editing a bounded directory, safe git clone and archive extraction, an isolated-subprocess primitive, and deterministic quote grounding. Hardened by dogfooding against real downstream consumers; the changes that surfaced are in CHANGELOG.md.

1.0.0 means the public surface is a contract: __init__.__all__, the rlm-harness/trace/v1 wire format, and RLMTask's declaration fields are frozen under SemVer and pinned by tests/test_contract.py. Additions ship in a minor release; a rename or removal ships with an alias and a DeprecationWarning first, and the removal itself waits for the next major. The trace format carries its own version and evolves additive-only within v1. A _-prefixed name or module internal is not part of that promise.

Next: enable optimize.compile_task against a labelled trainset to actually GEPA-compile tasks (currently a documented stub).

License

MIT © Boik Su (@boik_su). See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rlm_harness-1.5.0.tar.gz (562.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

rlm_harness-1.5.0-py3-none-any.whl (225.4 kB view details)

Uploaded Python 3

File details

Details for the file rlm_harness-1.5.0.tar.gz.

File metadata

  • Download URL: rlm_harness-1.5.0.tar.gz
  • Upload date:
  • Size: 562.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for rlm_harness-1.5.0.tar.gz
Algorithm Hash digest
SHA256 3511bcf043c0bd6e72a191df507bf3829a20f53983c8ecfa2a7f41479f3a7203
MD5 76790e60b3642299a210710f11c1a800
BLAKE2b-256 9de8cca04fb53a308870e83dccbe7bc6ea5bbc473c062a82b0d4f42888b2c76c

See more details on using hashes here.

Provenance

The following attestation bundles were made for rlm_harness-1.5.0.tar.gz:

Publisher: release.yml on qazbnm456/rlm-harness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file rlm_harness-1.5.0-py3-none-any.whl.

File metadata

  • Download URL: rlm_harness-1.5.0-py3-none-any.whl
  • Upload date:
  • Size: 225.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for rlm_harness-1.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 eb0595b9ee8b3d14311fe87366310c1230b69af62e71cfb02226b77740a46034
MD5 72f8a0afa62ca784e00e934ae7bf67c8
BLAKE2b-256 8d7a2bdde64af940b67726091b6b941de99777753edb58256f7b7903113f7003

See more details on using hashes here.

Provenance

The following attestation bundles were made for rlm_harness-1.5.0-py3-none-any.whl:

Publisher: release.yml on qazbnm456/rlm-harness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.10.0

2 files

1.9.1

2 files

1.9.0

2 files

1.8.4

2 files

1.8.3

2 files

1.8.2

2 files

1.8.1

2 files

1.8.0

2 files

1.7.0

2 files

1.6.1

2 files

1.6.0

2 files

This release

1.5.0 This release

2 files

1.4.0

2 files

1.3.0

2 files

1.2.1

2 files

1.2.0

2 files

1.1.0

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page