Skip to main content


😇
HALO

✨ RLM-based agent optimizer using production traces✨

X (formerly Twitter) License GitHub

QuickstartWhat is this?BenchmarksDevelopmentContributing

Quickstart

Install the HALO desktop app with:

curl -fsSL https://inference.net/halo/install.sh | sh
Read HALO reports

The installer downloads the latest release for your platform and sets up the desktop app. macOS uses a signed, notarized DMG. You can also install directly from the GitHub releases page.

For a full walkthrough of the desktop app, from loading demo traces to running HALO and shipping fixes with a coding agent, see the HALO Desktop guide.

If you're looking for a hosted, plug-and-play version of HALO, please sign up for inference.net and follow the instructions here.

What is this?

HALO is a methodology for building recursively self-improving agent harnesses using RLMs. This repository contains:

  • The HALO Desktop App for running HALO locally on your machine.
  • Information on HALO methodology.
  • A Python package that implements the core HALO-RLM engine. View on PyPI
  • A demo project that shows how to build HALO loops for your agents using the Python package. View demo
  • Benchmarking examples applying HALO to popular agent benchmarks. (View AppWorld).

HALO Loop

The core HALO loop is surprisingly simple:

  1. Collect execution traces from your agent harness. HALO uses OpenTelemetry-compatible tracing.
  2. Feed traces into HALO-RLM engine.
  3. The engine decomposes the traces to understand common failure modes across harness executions and produces a report with its findings.
  4. This report is fed into a coding agent like Cursor or Claude Code to generate and apply a set of changes to your harness.
  5. The harness is then re-deployed, more traces are gathered, and the cycle repeats.

HALO is great at finding issues in production agent deployments. We find high-traffic environments tend to generate more data with higher variance across executions, creating the type of issues that HALO is great at identifying.

Why an RLM?

A general-purpose harness like Claude Code is the wrong tool for trace analysis. This isn’t because the model isn’t smart, but because traces can get extremely long, and you need a specialized toolkit in order to make observations about systemic agentic behavior. We noticed in our testing that harnesses like CC would often overfit to an error present in a single/few traces rather than generalize to harness-level problems. This led us to creating a specialized form of a RLM.

rlm

Get Started

Install

Install the HALO engine + CLI from PyPI:

pip install halo-engine

# Verify installation
halo --help

Usage

  1. Integrate Tracing
  2. Collect traces by running your agent
  3. Run the HALO engine
export OPENAI_API_KEY=...
# Optional: point HALO at another OpenAI-compatible provider.
export OPENAI_BASE_URL=https://openrouter.ai/api/v1

halo path_to_your_traces.jsonl -p "Diagnose errors you find and suggest fixes"

HALO uses the canonical OpenAI env vars: OPENAI_API_KEY for credentials and OPENAI_BASE_URL for OpenAI-compatible providers. If OPENAI_BASE_URL is unset, HALO uses https://api.openai.com/v1. Run halo --help to see all CLI options. The CLI mirrors the model/provider settings exposed by the Python SDK's ModelConfig and ModelProviderConfig.

CLI options

Flag Default Description
TRACE_PATH required JSONL trace file
--prompt, -p required User prompt sent to the root agent
--model, -m gpt-5.4-mini Model name for root and subagent calls; also the fallback for synthesis and compaction
--synthesis-model --model Model for synthesis calls (trace summarization). A small, cheap model (e.g. gpt-4.1-nano) is recommended
--compaction-model --model Model for compaction calls (context summarization) — the biggest token consumer in large runs. A small, cheap model (e.g. gpt-4.1-nano) is recommended
--max-depth 2 Max subagent recursion depth
--max-turns 20 Max turns per agent
--max-parallel 10 Max concurrent subagents
--base-url OPENAI_BASE_URL / https://api.openai.com/v1 OpenAI-compatible API base URL
--api-key OPENAI_API_KEY Provider API key
--header, -H unset Provider header as NAME: VALUE. Repeat for multiple headers, matching curl's -H convention
--temperature provider default Sampling temperature forwarded to the model
--max-output-tokens provider default Maximum output tokens forwarded to the model
--parallel-tool-calls / --no-parallel-tool-calls enabled Allow models to issue parallel tool calls
--refusal-retries 0 Retry an agent model request this many times when the model refuses
--reasoning-effort model/provider default Reasoning effort for root and subagent calls.
--telemetry off Emit OpenInference traces of HALO's own LLM, tool, and agent activity

For example:

halo path_to_your_traces.jsonl \
  -p "Diagnose errors you find and suggest fixes" \
  --base-url https://openrouter.ai/api/v1 \
  -H "HTTP-Referer: https://example.com"

Telemetry

HALO can emit OpenInference-shaped traces of its own LLM, tool, and agent activity. It is off by default; nothing is emitted unless you pass --telemetry.

halo TRACE_PATH --prompt "..." --telemetry

When telemetry is enabled, setting INFERENCE_API_KEY uploads spans to inference.net over OTLP. If it is not set, spans are written to a local JSONL file at ./halo-telemetry-{run_id}.jsonl in the current working directory.

Var Default Purpose
INFERENCE_API_KEY unset inference.net API key. If set, uploads spans over OTLP
INFERENCE_OTLP_ENDPOINT SDK default OTLP endpoint base URL, for example https://telemetry.inference.net
INFERENCE_DEBUG unset Set to 1 to surface OTLP export errors
HALO_TRACING_RUN_ID unset Uses this HALO run id instead of a generated uuid
HALO_TRACING_* unset Generic resource-attribute passthrough (HALO_TRACING_TEAM_IDhalo.team.id)
HALO_TELEMETRY_PATH ./halo-telemetry-{run_id}.jsonl Local fallback file path. Only used when no ingest token is set

We have provided a simple demo and an AppWorld demo.

Python API

The engine exposes four entry points from engine.main. Use whichever matches the trade-off you want between observability and code simplicity. The yielded types (AgentOutputItem and AgentTextDelta) are defined in engine/models/engine_output.py:

Function Sync / async Returns When to use
stream_engine_async async AsyncIterator[AgentOutputItem | AgentTextDelta] You want every event including streaming-token deltas (live UI, custom rendering).
stream_engine_output_async async AsyncIterator[AgentOutputItem] You want to log / persist each completed step (assistant message, tool call, tool result) as it lands.
run_engine_async async list[AgentOutputItem] You want the final list at the end and don't care about per-step observability.
stream_engine sync Iterator[AgentOutputItem | AgentTextDelta] Sync generator; yields every event including deltas. Drives the async iterator on a private event loop.
stream_engine_output sync Iterator[AgentOutputItem] Sync generator; yields completed items only. Same shape as the async variant for sync callers.
run_engine sync list[AgentOutputItem] Sync, collects to a list. Pure convenience over asyncio.run(run_engine_async(...)).
from engine.main import stream_engine_output_async

async for item in stream_engine_output_async(messages, cfg, trace_path):
    logger.info("step", extra={"sequence": item.sequence, "agent": item.agent_name})
    # item.item is an AgentMessage (assistant / tool / etc.)

Benchmarks

HALO is consistently capable of driving improvements on benchmarks, solely by optimizing the harness.

AppWorld

We applied HALO to the AppWorld benchmark, a set of agentic tasks that assess the LLM’s ability to use multi-app services like Spotify, Venmo, file systems, and phone contacts. We tested HALO’s ability to improve harnesses for both Gemini 3 Flash and Sonnet 4.6. We iterated on the harness using the dev split, and then used the test_normal split as a proxy to verify that improvements did not come from overfitting.

The feedback from HALO Engine surfaced failures in the harnesses such as hallucinated tool calls, redundant arguments in tools, refusal loops, and semantic correctness issues. Each issue mapped cleanly to a direct prompt edit. HALO’s claims were independently verified from the source trace files with the findings holding up under scrutiny.

app-world-sgc The peak improvements over baseline were substantial for both models. For Gemini 3 Flash, dev SGC went from 36.8% to 52.6% (+15.8 points) and test_normal SGC went from 37.5% to 48.2% (+10.7 points). For Sonnet 4.6, dev SGC went from 73.7% to 89.5% (+15.8 points) and test_normal SGC went from 62.5% to 73.2% (+10.7 points).

Development

Local development against this repo uses uv for dependency management and go-task as the task runner.

Setup

git clone https://github.com/context-labs/HALO
cd HALO
task env:setup

task env:setup installs uv (if missing), syncs the venv from uv.lock, and configures the repo's git hooks. After that, the halo CLI is available via uv run halo ... (or activate .venv/).

Common tasks

Run task --list for the full list. The ones you'll use most:

Task What it does
task check Run all pre-commit checks: pinned-versions, lint, format, typecheck, unit tests
task check:fix Same, but auto-fix lint/format issues
task test:unit Unit tests under tests/unit/
task test:integration Integration tests under tests/integration/

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a pull request.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

halo_engine-0.3.5.tar.gz (10.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

halo_engine-0.3.5-py3-none-any.whl (157.3 kB view details)

Uploaded Python 3

File details

Details for the file halo_engine-0.3.5.tar.gz.

File metadata

  • Download URL: halo_engine-0.3.5.tar.gz
  • Upload date:
  • Size: 10.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.4 {"installer":{"name":"uv","version":"0.12.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for halo_engine-0.3.5.tar.gz
Algorithm Hash digest
SHA256 c37e59b3e6e277e030542c2324577088363bd1ca29032d6a2797b150395268a1
MD5 f727c2b1b8c9034e08fc75c629e428c7
BLAKE2b-256 4b8c3ea49359737ca24ff435c7d4020e13e404bc0467ad65a9e28b08c5c007df

See more details on using hashes here.

File details

Details for the file halo_engine-0.3.5-py3-none-any.whl.

File metadata

  • Download URL: halo_engine-0.3.5-py3-none-any.whl
  • Upload date:
  • Size: 157.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.4 {"installer":{"name":"uv","version":"0.12.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for halo_engine-0.3.5-py3-none-any.whl
Algorithm Hash digest
SHA256 45daef661473163e7a2db5c11fac2d37eb1d812c99923dd643b569c86fc5a0fb
MD5 7fc7a2171a113085109fe8487115f94c
BLAKE2b-256 078d11b008e023ef09ff08403f18e70226fff55c99d7b75d7c10bf52d17f9d52

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page