Black Box: a flight recorder for AI agents
When an AI agent fails, the mistake usually happened several steps before the wrong answer. Black Box records every step of an agent run, finds the step that caused the failure, and proves it: it re-runs the agent from its recording with only that step repaired. If the run now passes, that step was the cause. Unchanged steps come from the recording, so a replay costs zero model calls and a fix re-runs only what changed.
pip install halfabyte-blackbox # recorder, replay, proof, CLI, CI, MCP server, trace import
pip install "halfabyte-blackbox[ui]" # + the API / dashboard backend
pip install "halfabyte-blackbox[all]" # + step checker, ranking model, test agents, dataset loaders
Add it to an existing agent (3 lines)
import blackbox
from openai import OpenAI
client = blackbox.wrap(OpenAI()) # 1. every model call is recorded
@blackbox.tool # 2. every tool call is recorded and replayable
def get_weather(city: str) -> dict:
...
def my_agent(question: str) -> str: # your agent, unchanged
...
trace = blackbox.run(my_agent, "Will it rain in Pune?") # 3. run it under the recorder
blackbox.replay(trace) # re-run from the recording: 0 model calls
child = blackbox.fork(trace, 2, blackbox.Change(kind="output", value={"result": {"rain": True}}))
blackbox.savings(child) # steps re-run vs re-used, tokens saved
Works with agents you did not write (shown on Hugging Face smolagents), and with traces you already export:
blackbox import otel traces.json / blackbox import langfuse trace.json.
What it does
| Capability | How |
|---|---|
| Record every model call, tool call, search and memory change, with a save-point after each step | blackbox.wrap, @blackbox.tool, blackbox.run |
| Diagnose: rank the steps most likely to have caused a failure, with plain-language evidence | blackbox show <run>, API /api/runs/{id}/diagnosis |
| Prove the cause by re-running with one suspect repaired at a time | Verify (API, MCP verify) |
| Causal report: necessary vs sufficient, joint causes, every recovery path, blast radius | API /causal, MCP causal_report |
| One fix for many: test one rule on similar past failures and passing runs; APPROVE only if nothing breaks | blackbox fleet <run> |
| Crash test: plant every known kind of mistake into a working agent, grade it 1–5 stars | blackbox crash-test <run> |
| Seen this before?: failure fingerprints, look-alike failures, novel failures, known fixes | MCP similar_failures |
| Guardian: incidents in plain words on email, WhatsApp, Slack, Telegram; money-moving agents paused until approved | blackbox.Guardian(...), blackbox guardian test |
| Black Box CI: replay pinned recorded runs on every pull request; fail on regressions | blackbox ci pin / blackbox ci run, GitHub Action |
| MCP server: let Claude Code, Cursor or VS Code investigate failures with 28 tools | blackbox mcp |
| Live Lab: ask any question, plant a mistake live, watch the 14-stage investigation | blackbox lab "question" |
Safety: tools that move money never execute during any re-run, and Guardian never retries them without a human.
Command line
blackbox runs --fail # failed runs
blackbox show <run_id> # step by step, with who-used-whose-output links
blackbox replay <run_id> # replay from the recording
blackbox crash-test <run_id> # red-team a passing run
blackbox fleet <run_id> --rule "..." # one fix for many, with a regression firewall
blackbox ci pin && blackbox ci run # regression firewall for agent code
blackbox guardian channels # which alert channels are configured
blackbox mcp # MCP server on stdio
blackbox ui # API on http://127.0.0.1:8000
Models are named in one file, models.yaml (Ollama locally, Groq or OpenRouter hosted); keys live in .env.
Development (this repository)
pip install -e ".[all,dev]"
py -3 -m pytest -q -m "not live"
cd web && npm install && npm run dev # dashboard on http://localhost:5173
Design: SYSTEM.md · rules: CLAUDE.md · UI contract: docs/API_FOR_UI.md · MCP: docs/MCP.md · CI: docs/CI.md ·
Guardian setup: docs/GUARDIAN_SETUP.md. Branches <name>/<feature> → PR into dev → main at milestones.
Built by team Half a Byte: Aryan Lomte, Radhesh, Aditya, Advay Chavan.
Metadata
Release files for halfabyte-blackbox 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| halfabyte_blackbox-0.1.0.tar.gz | 264.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| halfabyte_blackbox-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 523.7 kB
Release files / halfabyte_blackbox-0.1.0.tar.gz
| Download URL | halfabyte_blackbox-0.1.0.tar.gz |
|---|---|
| Size | 264.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1d3115d8f56bc4b09ffd3c339a62f7b37be503dec7f86f5391cbc50b72a9db9a
|
|
BLAKE2b-256 checksum How to use checksums |
22aa6f468814f84398b74ca4302ba63fa6f91245ca8cf19d9ef2f85ca88e3002
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.4
|
Release files / halfabyte_blackbox-0.1.0-py3-none-any.whl
| Download URL | halfabyte_blackbox-0.1.0-py3-none-any.whl |
|---|---|
| Size | 259.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5a570143512b30c6cb57c132f2fcbaef235436aa5baa3f7aeb37899d09913fa8
|
|
BLAKE2b-256 checksum How to use checksums |
1a8bd7c335a4f40b5c0723a35713daf2f82a56ade2d6cd6523408c6f3ef65903
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.4
|