runtape
runtape records AI agent runs to local JSONL files and finds which part of the model's context caused a given decision.
runtape why removes parts of the context, reruns the recorded model call
several times per variant, and reports the parts whose removal changes the
decision, with a significance test.
In the example above, an email assistant forwards an invoice to an external
address. runtape why attributes the call to one sentence in an HTML comment
inside a vendor email. With that sentence removed, the agent does not forward
in 10 of 10 reruns (p = 5e-6).
Install
pip install runtape
Requires Python 3.10+. Supports the OpenAI and Anthropic SDKs, LangChain and LangGraph, OpenAI-compatible local servers (Ollama, LM Studio, vLLM), and custom agent loops.
Examples
The repository includes two example agents. Both run offline with a
rule-based stand-in model, passed to why with --model-fn.
git clone https://github.com/RehanMohammed985/runtape
cd runtape
pip install . openai anthropic
python examples/inbox_agent.py
runtape why last tool:forward_email --model-fn examples/inbox_agent.py:simulated_model
python examples/refund_bot.py
runtape why last tool:issue_refund --model-fn examples/refund_bot.py:simulated_model
To run the refund example against a local model:
ollama pull llama3.1:8b
python examples/refund_bot.py --local llama3.1:8b
runtape why last tool:issue_refund
With llama3.2 (3B), the refund agent paid order B-2290 $64, the amount from a different customer's order earlier in the conversation. On the recorded context this happened in 9 of 40 reruns. With the earlier order lookup removed, 0 of 40 (Fisher exact test, p = 0.001).
Recording
import runtape
from anthropic import Anthropic
rec = runtape.record(name="support-bot")
client = rec.wrap(Anthropic())
@rec.tool
def lookup_order(order_id: str):
...
- OpenAI:
rec.wrap(OpenAI())records Chat Completions and the Responses API. - LangChain / LangGraph:
graph.invoke(inputs, config={"callbacks": [rec.langchain()]}). - Other frameworks:
rec.log_llm_request(...)andrec.log_llm_response(...).
Traces are written to ./traces/, one file per run, as events happen. Sync,
async and streaming calls are recorded, including Anthropic's
messages.stream(), OpenAI's .stream() and .parse(), and beta endpoints.
For OpenAI-compatible servers, the base URL is recorded so reruns go to the
same server.
runtape why
runtape why <trace> <event>
<trace> is a path or last. <event> is an event number, tool:NAME (the
last call to that tool), or last (the last model response).
Procedure:
- Rerun the recorded call on the unchanged context to measure how often the model makes the same decision.
- Split the context into pieces: system prompt, messages, tool results.
- Remove each piece and rerun (2 runs to screen, more where the decision changes).
- Confirm candidates with a one-sided Fisher exact test, Bonferroni-corrected over all variants tried.
- Narrow each confirmed piece to JSON items, paragraphs and sentences.
- Search inside pieces whose removal has no effect, for causes masked by other content in the same piece.
- Detect causes that are duplicated or individually sufficient.
- Separate pieces that change the action from pieces the agent only needs as input (without them it stops or repeats a lookup).
Only the selected model call is rerun. Tools are not executed.
Options: --runs (reruns per variant, default 5), --budget (max model
calls, default 400), --dry (rank suspects without model calls), --match REGEX (explain a text answer), --max-pieces, --expand, --json.
Replies are cached in ./.runtape/.
runtape rerun
runtape rerun <trace> <event> --drop "9[1]"
runtape rerun <trace> <event> --replace "old=>new"
runtape rerun <trace> <event> --system-file fixed.txt
runtape rerun <trace> <event> --model <name>
runtape odds <trace> <event>
rerun applies an edit to the recorded context and reports the distribution
of decisions. odds reports the distribution on the unchanged context.
The same check can be used as a test:
def test_no_forwarding_from_injected_email():
runtape.rerun("traces/injected.jsonl", 22, drop=["12.body para 4"]).never_calls("forward_email")
always_calls, rate(tool) and counts() are also available.
Replay
with runtape.replay("traces/bad-refund.jsonl") as rp:
client = rp.wrap(Anthropic())
@rp.tool
def lookup_order(order_id): ... # returns the recorded result
run_agent(client)
Model responses and tool results are served from the trace, so the run is
offline and deterministic. Tool results are restored to their original types
(dataclasses, Pydantic models, tuples). If a request differs from the
recording, replay stops and reports the first difference, or continues
against the live model with on_diverge="live".
Browsing a trace
runtape # newest trace
runtape traces/run.jsonl
| command | |
|---|---|
Enter, step [N or TYPE] |
next event |
back, goto N |
move |
show [N] |
full event |
context [N] |
the model's input at that point |
grep TERM, next, prev |
search |
diff [A] [B] |
context changes between two events |
why, rerun, odds |
run on the current event |
Other subcommands: ls, summary, timeline, show, context, grep,
diff.
MCP server
pip install "runtape[mcp]"
claude mcp add runtape -- runtape mcp
Exposes trace inspection, why and rerun as MCP tools.
Trace format
Append-only JSONL, one event per line. Message history is delta-encoded. See SPEC.md.
Pass redact=fn to runtape.record to filter events before they are
written. If fn raises, the event content is dropped.
Limitations
whycalls the model, typically 100 to 250 times per decision. Use--dry,--budgetand--max-piecesto limit cost.- On a simulated model that ignores its context, false positives occurred in 0 to 5 of 100 runs (alpha = 0.05). A cause that moves the decision rate from 90% to 10% was found in all runs; 90% to 30%, in about 4 of 5.
- Decisions made in fewer than about 1 in 5 reruns are too rare to attribute
automatically. Use
oddsandrerun --dropinstead. - At temperature 0, each variant is run once and each candidate twice.
- Removed content is replaced with
[content removed], which can itself affect the model. - By default, 80 pieces are tested (ranked by word overlap with the decision, always including the system prompt, the first user message and the latest message), and 6 are searched for masked causes. Untested pieces are listed in the report.
- Replay does not serve streamed calls. LangChain runs can be recorded and analyzed but not replayed.
- The
--model-fncache key includes the function's source file. Use--no-cacheif the function depends on other code that changed. - Providers other than OpenAI, Anthropic and OpenAI-compatible servers need
--model-fn.
Related work
ContextCite and TracLLM attribute single model responses to context by ablation. Causal Agent Replay, AgentDebugX and AgentDoG apply attribution and counterfactual reruns to agents. AttriGuard uses reruns to detect prompt injection at runtime. LangSmith, Laminar and Langfuse record and replay agent runs as hosted platforms.
License
MIT
Release files for runtape 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| runtape-0.3.0.tar.gz | 346.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| runtape-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 418.8 kB
Release files / runtape-0.3.0.tar.gz
| Download URL | runtape-0.3.0.tar.gz |
|---|---|
| Size | 346.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bbe71093a72d5f57be84de8301882c99799e90875d3b861326fa2f25b2ed9271
|
|
BLAKE2b-256 checksum How to use checksums |
ea93aa84df158a4b736a7e209050214fe6f1660ac7213060deecc4c1b0de6fc0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.2
|
Release files / runtape-0.3.0-py3-none-any.whl
| Download URL | runtape-0.3.0-py3-none-any.whl |
|---|---|
| Size | 72.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4b80ec71e984f4c33a96abe955fbd315a2234d8b7a5f0b27474b18d3de82486e
|
|
BLAKE2b-256 checksum How to use checksums |
f7c66230d9c9d02f935bda7e593378633285b0908576471d1014eb394648846c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.2
|