Skip to main content

backspin

CI PyPI Python License: MIT Code style: ruff

The flight recorder for AI agents. Record every LLM call and tool call of an agent run into one portable file, replay the run deterministically with zero API access, and diff two runs to find the exact step where behavior diverged. 100% local, zero core dependencies.

Think rr, but for agents instead of processes.

from backspin import Recorder

with Recorder(agent="my-agent") as rec:
    client = rec.capture_openai(OpenAI())   # any code that talks to chat.completions
    ...                                     # your agent, unchanged

That's the whole integration. Everything the agent did — prompts, completions, tool calls, timings, token counts — is now in a single runs/*.backspin.jsonl file you can open, replay, diff, or attach to a bug report.

中文文档

Why

Your agent made forty LLM calls, called three tools, and then did something weird at 2am. Good luck reproducing that from a chat log.

Cloud observability tools (Langfuse, LangSmith, AgentOps…) answer "what happened?" on a dashboard. backspin answers "make it happen again, exactly." It's a debugger, not a dashboard:

  • Record — one context manager captures every OpenAI-shaped call (sync, async, streaming) plus your own tool calls and logs.
  • Replay — the run becomes a cassette: your agent re-runs offline with recorded responses injected, deterministically. Perfect for regression tests and reproducing bugs without API keys or cost.
  • Diff — replay the same agent against a fix and diff the two runs; backspin pinpoints the first step where they stopped matching.
  • Local-first — runs are plain JSONL files. No server, no account, no telemetry. git attach-friendly: the failing run is the bug report.

Install

pip install "backspin[ui]"     # SDK + CLI + local viewer

Core has zero dependencies; [ui] adds FastAPI + uvicorn for the local viewer.

Record

from openai import OpenAI
from backspin import Recorder

rec = Recorder(agent="support-bot")

with rec:
    client = rec.capture_openai(OpenAI())

    @rec.tool
    def lookup_order(order_id: str) -> str:
        return "shipped"

    rec.log("user asks about order #1234")
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Where is order #1234?"}],
    )

print(rec.path)   # runs/20260829-142300-support-bot-9f31c2.backspin.jsonl

Streaming and async clients are captured too; a streamed response is transparently reconstructed into one recorded completion.

Replay

from backspin import Cassette, load_run, stub_client

cassette = Cassette.from_run(load_run(rec.path))
stub = stub_client(cassette)

# Same agent code, zero network: responses come from the recording.
answer = run_agent(stub)

Requests are matched by fingerprint (model + messages) and fall back to call order with a warning. Use backspin.replay.patch_openai(cassette) to patch openai.OpenAI itself when you can't inject a client.

In tests this becomes deterministic agent regression testing: record once, assert forever, at zero token cost. The pytest plugin does the asserting for you:

def test_my_agent(backspin):                       # pip-installed = auto-loaded
    with backspin.record(agent="t") as rec:
        run_agent(rec.capture_openai(client))
    backspin.assert_replays_identically()          # strict: fingerprint-exact replay

What-if branching

The debugger superpower: change one answer, keep everything else constant, and see what the timeline looks like downstream.

from backspin import branch, diff_runs, load_run

branch_path = branch("runs/live.backspin.jsonl", {0: {"content": "Rome it is."}})
report = diff_runs(load_run("runs/live.backspin.jsonl"), load_run(branch_path), llm_only=True)
print(report.first_divergence)   # the step where the two timelines split

Or from the CLI: backspin branch runs/live.jsonl --step 0 --content "Rome it is." — writes a branch run (marked branch_of) and prints the divergence report.

Spans: structure, not just a flat list

with rec.span("research", meta={"topic": "weather"}):
    with rec.span("tool:search"):
        ...
    resp = client.chat.completions.create(...)   # recorded inside the span

Every event inside a span carries its span_id and nesting depth; spans are safe under concurrency (each asyncio task gets its own stack) and the viewer renders the tree. Spans never inflate duration totals.

Zero-code integration: backspin proxy

Can't (or don't want to) instrument code? Run the OpenAI-compatible local proxy and point any agent at it — any framework, any language:

backspin proxy --upstream https://api.openai.com --port 8840
# client: base_url = http://127.0.0.1:8840/v1   ← that's the whole integration

Every call is forwarded and captured, streaming included. Flip the same proxy into replay mode and it serves a recorded run back as an API — deterministic replay for agents written in any language, no SDK required:

backspin proxy --replay runs/live.backspin.jsonl --port 8840

Multi-provider: OpenAI, Anthropic, and everything OpenAI-compatible

Claude natively? Same story, two lines:

rec.capture_anthropic(Anthropic())   # sync/async/streaming, tool_use included

Anthropic events record with provider: "anthropic" and usage normalized to the same token fields, so costs and diffs work across providers. And because the proxy speaks /v1/messages too, Claude-native agents get the same record-or-replay treatment with zero code changes.

Anything that speaks the OpenAI protocol (DeepSeek, Qwen, Kimi, GLM, vLLM/Ollama, OpenRouter, …) is covered by capture_openai / the proxy out of the box.

Export, share, TUI

backspin export runs/live.jsonl --format sft -o train.jsonl   # eval/SFT datasets
backspin share runs/live.jsonl        # one self-contained HTML: run + viewer
backspin tui                          # keyboard-driven viewer for the terminal

share bundles the entire viewer and the run into a single .html — send it to a teammate, they open it in a browser and step through the run. Nothing is uploaded anywhere.

Costs

A built-in price table (gpt-4o, claude, gemini, deepseek, …) turns token counts into money: run.totals()["cost_usd"], a cost card in the viewer, ~$0.0142 in backspin show. Extending the table is a one-dict PR.

Diff

backspin diff runs/live.backspin.jsonl runs/replay.backspin.jsonl
# runs diverge at step #14
#   #13  llm   gpt-4o-mini        gpt-4o-mini        yes
#   #14  llm   gpt-4o-mini        gpt-4o-mini        NO

Steps are aligned and signed by what the agent chose to do (LLM request fingerprint / tool name), so the first mismatch marks exactly where two runs stopped matching — before costs and latencies are even considered.

CLI & local viewer

backspin ls                  # list runs: agent, steps, tokens
backspin show runs/...jsonl  # print a run's timeline
backspin show runs/... --step 7   # dump one step as JSON
backspin diff a b            # diff two runs (exit 1 if they differ)
backspin ui                  # http://127.0.0.1:8787 — timeline, inspector, diff

The viewer is a zero-build vanilla JS app served by the CLI: a waterfall timeline, a step inspector (request / response / raw), and side-by-side run diffing.

backspin timeline viewer

backspin diff view

Keeping secrets out of recordings

Recordings contain full prompts and completions. When that's not okay, pass a redact function — every payload value goes through it before touching disk:

from backspin import Recorder
from backspin.redaction import mask, redact_strings

rec = Recorder(
    agent="support-bot",
    redact=redact_strings(mask(r"sk-[A-Za-z0-9]{8,}")),
)

Structural fields (model, tool name, fingerprint, durations) stay readable so the viewer and diff keep working; everything else — including unknown custom payload keys — passes through the redactor. Fingerprints are computed pre-redaction, so replay matching is unaffected. Note the tradeoff: a redacted run still replays, but replayed values are the redacted ones.

The run file

One run = one self-contained JSONL file. First line is a header, every following line is a step:

{"kind": "llm", "seq": 3, "ts": 1756448402.1, "model": "gpt-4o-mini",
 "duration_ms": 812.4, "fingerprint": "9f31c2ab77e01d44",
 "request": {"messages": [...]}, "response": {"choices": [...]},
 "usage": {"prompt_tokens": 120, "completion_tokens": 45}}

Kinds: llm, tool, log, error, plus whatever custom events you record via rec.event(kind, **payload). Because a run is one file, "please attach the failing run" finally works.

How backspin compares

Langfuse / LangSmith / AgentOps backspin
Question it answers what happened? make it happen again, exactly
Where cloud SaaS your machine, plain files
Replay with recorded responses no yes, deterministic
First-divergence diffing no yes
Setup SDK + account + ingestion one context manager
Token cost of debugging full price zero after recording

They compose well: keep the dashboard if you like it, attach a backspin run when someone says "I can't reproduce it."

Status & roadmap

backspin is a young project (v0.5) — record → replay → what-if → diff → view works end to end across OpenAI and Anthropic protocols, from Python and TypeScript, verified against the real SDKs. Next:

  • Async + streaming capture, spans, redaction, costs, pytest plugin (0.2/0.3)
  • Sidecar proxy: record + replay, OpenAI protocol (0.3)
  • Anthropic native: SDK capture + proxy /v1/messages; TypeScript SDK; export/share/TUI (0.4)
  • Agent-level what-if (re-run the whole agent against a mutated cassette) (0.5)
  • Docs site with runnable examples
  • Deterministic clock/random stubs for full boundary capture

Development

pip install -e ".[dev]"
pytest                      # full suite incl. real-SDK integration tests
ruff check backspin/ tests/ # lint
mypy backspin/              # types
python examples/mock_agent.py   # zero-setup demo: record → replay → diff

See CONTRIBUTING.md for ground rules (core stays dependency-free; the run format is a contract) and SECURITY.md for how to report vulnerabilities.

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

backspin-0.5.1.tar.gz (256.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

backspin-0.5.1-py3-none-any.whl (55.7 kB view details)

Uploaded Python 3

File details

Details for the file backspin-0.5.1.tar.gz.

File metadata

  • Download URL: backspin-0.5.1.tar.gz
  • Upload date:
  • Size: 256.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for backspin-0.5.1.tar.gz
Algorithm Hash digest
SHA256 025d1ae32d81dabfcb0fbfe3e0f409c5560778e10bc09676f1b5ff7c1d29c617
MD5 f655c956d7b48022cc3b82915bafe4c9
BLAKE2b-256 aab7d2324c9f6a724bab2c9fd5dfb9cab62741d6611bbd970d9aa940f8f3af33

See more details on using hashes here.

Provenance

The following attestation bundles were made for backspin-0.5.1.tar.gz:

Publisher: release.yml on zaibuchihuoji/backspin

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file backspin-0.5.1-py3-none-any.whl.

File metadata

  • Download URL: backspin-0.5.1-py3-none-any.whl
  • Upload date:
  • Size: 55.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for backspin-0.5.1-py3-none-any.whl
Algorithm Hash digest
SHA256 2bcc2b25930e308b2a9cf2469ddcb4f5b1c06384000762d70f29506ee5edfb72
MD5 1ee52afe0b82a4f042256ad517271ea2
BLAKE2b-256 1a948320d93b4f571e0ec267e6fbefd1e445cf24de05c609b1043f5467052e40

See more details on using hashes here.

Provenance

The following attestation bundles were made for backspin-0.5.1-py3-none-any.whl:

Publisher: release.yml on zaibuchihuoji/backspin

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page