🧵 Loom
English · 中文
The black box, firewall & debugger for AI agents.
Your agent ran — touched files, called tools, spent tokens — and you have no idea what it did or why. Loom records every action, replays it byte-for-byte for $0, firewalls dangerous calls before they run, and lets you step through the whole run like a debugger. Works with any Claude/OpenAI-API agent — Claude Code, LangGraph, CrewAI, your own.
pip install loom-harness # zero dependencies
loom record claude "fix the failing test" --safe
recorded 17 steps · 42k tokens → session.loom.json
🛡 firewall blocked 1 risky call: Read(".env")
🔬 loom debug session.loom.json # step through it, fork any turn live
Step through any agent run — the reasoning, the tools, the exact context the model saw, and fork any turn live.
Why Loom
- 🎥 Record any agent — proxy Claude Code / Codex / Cursor / your own, one command, zero code changes.
- ⏪ Replay for $0 — every call recorded at one boundary → byte-identical, offline. Deterministic CI for a stochastic agent.
- 🔬 Step-debug it — walk each step, see the exact context the model saw, then edit a turn and re-run it live.
- 🕸 Any multi-agent framework — LangGraph · CrewAI · AutoGen · OpenAI-Agents · Claude-SDK, recovered into one agent tree from the wire, zero code changes.
- 🛡 Firewall it — deny / confirm dangerous calls before they run, by capability (
cap:money_movement) or sequence (after Read(.env): deny network). - 🕵 Catch exfiltration — a secret flowing to an egress, even base64-encoded or paraphrased, confirmed by an LLM judge.
- ↩ Undo the world — revert the files an agent changed, or snapshot & restore a whole workspace + database.
The debugger
loom debug run.loom.json (or loom live to watch it run) opens a step-debugger in your browser:
- Step through every action — the model's reasoning, the tool call + args, the world-diff (file / SQL row / DOM), risk, tokens.
- Context frame — the exact conversation the model saw at each step: the debugger's stack & variables.
- Fork & re-run live — inject a message or switch the model at any turn; only the divergent tail costs a call, and the branch appears beside the original.
- Multi-agent tree — a supervisor/sub-agent system (yours or a third-party framework) recovered from the wire and shown as a collapsible tree, laned by agent.
- Ask & assert — send the live agent a new message, or check plain-English expectations (
never issue_refund,output contains …) as a CI gate.
loom studio <trace> freezes the whole UI into one shareable HTML file (no server, no agent).
Debug a live agent
loom live --agent app:agent # watch it run, send follow-ups, fork any turn
Behind a gRPC / HTTP endpoint? Point your server at the recording proxy and drive it from the same debugger — no code, just your grpcurl:
loom live --proxy-port 9000 \
--trigger 'grpcurl -d "{\"prompt\": $LOOM_PROMPT_JSON}" -plaintext :50051 agent.Agent/Run'
# then start your server with ANTHROPIC_BASE_URL=http://127.0.0.1:9000
Loom reconstructs the agent's full internal hierarchy even though it's behind an endpoint.
Use it as a Python harness
from loom import Agent, tool, Policy
@tool
def search(q: str) -> str:
"Search the docs."
return db.search(q)
agent = Agent(model="claude-opus-4-8", tools=[search],
policy=Policy(deny=["issue_refund*"], budget_tokens=50_000)) # in-loop firewall
run = agent.run("What changed in the API last week?")
run.replay() # byte-identical, no API calls
run.fork(at=3) # rewind to turn 3, continue live on a new branch
One effect boundary records every model + tool call — so replay, fork, free CI, human-in-the-loop, the firewall, and every analyzer fall out of the same primitive. The kernel is zero-dependency.
A few more commands
loom replay <trace> |
re-run byte-identical, $0, offline |
loom taint · loom dlp --judge |
exfiltration lineage · semantic DLP |
loom redteam run --generate <m> |
AI red-teamer — invents attacks for your tool surface |
loom mcp gateway -- <server> |
firewall + record any MCP server |
loom undo <trace> |
revert the files the agent changed |
Run loom --help for the full set — record, replay, debug, live, studio, firewall,
taint, dlp, redteam, mcp, undo, cost, rootcause, experiment, and more.
Install
pip install loom-harness # kernel + CLI, zero deps
pip install "loom-harness[anthropic]" # + live Claude
pip install "loom-harness[mcp]" # + MCP gateway
Python 3.10–3.13 · MIT · import loom
Loom shrinks an agent's blast radius and makes its behavior inspectable — it is not a guarantee a model can't misbehave. See the threat model.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file loom_harness-0.33.11.tar.gz.
File metadata
- Download URL: loom_harness-0.33.11.tar.gz
- Upload date:
- Size: 2.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d2045d6edccd6a21042bd3f988e32fb83e1b4e9ba70ab896c461d7da06ee1274
|
|
| MD5 |
e6abb0168c00b1c09598ef51c55bc4f4
|
|
| BLAKE2b-256 |
1377d20a20a99223e165a0c501b7e2595d9e88ef2e1e9a263eaa79bcdcb1cadb
|
File details
Details for the file loom_harness-0.33.11-py3-none-any.whl.
File metadata
- Download URL: loom_harness-0.33.11-py3-none-any.whl
- Upload date:
- Size: 396.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
79b8e0a0b3b06eafa0addf33b39d87cd55d15fffdd931457b0f5bd129afd2efe
|
|
| MD5 |
05c7da7f595624d9eab5ce5cc147453b
|
|
| BLAKE2b-256 |
fdb0fba18a01d1e5356cc62c6eae61694910385d8566058f9ab8bcdffe2799f1
|