ACID transactions, time-travel debugging, and zero-cost Ghost Replay for AI agents. Rollback filesystem + state. Works with LangGraph, CrewAI, or raw Python.
Project description
๐ผ Agent VCR
Your AI agent corrupted the codebase. Now what?
The only tool that physically deletes hallucinated files โ git reset --hard, not just state rollback.
pip install ai-agent-vcr
No API keys. No cloud. No vendor lock-in. Works with TERX โ memory layer for browser agents.
The Problem Every AI Agent Developer Has Hit
You ran Claude Code, OpenHands, or a LangGraph agent autonomously.
It wrote 40 files. It failed at step 8. Now you have:
- 23 files that shouldn't exist
- A broken import chain
- State that says "success" on steps that half-ran
- No way to know what the repo looked like before step 6
Your repo after a bad autonomous run:
/src/handlers.py โ hallucinated, breaks import
/src/auth_v2.py โ duplicate of auth.py, never needed
/src/models_refactor.py โ partial rewrite, syntax error
/tests/test_fake.py โ tests for code that doesn't exist
/config/settings_new.py โ overwrote working config
Every other tool shows you logs. Agent VCR runs git reset --hard and deletes every one of those files.
ACID Rollback โ The Feature Nobody Else Has
from agent_vcr import VCRRecorder
from agent_vcr.integrations.openhands import ACIDWorkspace
recorder = VCRRecorder()
acid = ACIDWorkspace("/my/workspace", recorder=recorder)
acid.begin(session_id="task-001") # isolated git branch
acid.savepoint(state, node_name="coder") # checkpoint state + filesystem
acid.savepoint(state, node_name="tester")
# Agent writes 47 files. 23 are hallucinated garbage. Step 6 failed.
acid.rollback(to_frame_index=1)
# git reset --hard
# All 23 files: physically deleted from disk. Not hidden. Gone.
# Pre-existing ignored files like .env are preserved.
acid.commit() # merge only the clean branch
Before rollback: After rollback:
/src/handlers.py โ DELETED
/src/auth_v2.py โ DELETED
/tests/fake_test.py โ DELETED
/src/utils.py โ kept
/src/models.py โ kept
LangSmith shows you what happened. LangFuse shows you what happened. Arize shows you what happened.
Agent VCR changes what happened.
Ghost Replay โ Never Pay for the Same Task Twice
Agent succeeds? Save it. Run it again for free forever.
from agent_vcr.golden_cache import GoldenRunCache
cache = GoldenRunCache()
identity = GoldenRunCache.build_identity(
model="gpt-4o",
prompt_hash="prompt:v3",
code_commit="abc123",
tool_schema_hash="tools:v1",
)
cache.save_golden_run("Build a REST API with JWT auth", recorder, identity=identity)
# Every future run of the same task:
outputs, ledger = cache.replay("Build a REST API with JWT auth", identity=identity)
print(ledger)
RUN 1 (original) RUN 2 (Ghost Replay)
โโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโ
Tokens: 4,100 Tokens: 0
Cost: $0.0123 Cost: $0.00
Time: 2,350ms Time: 1ms
๐ฐ 100% savings ยท $0.0123 saved ยท 4,100 tokens ยท 2,349ms faster
Time-Travel Debugging
Agent fails at step 8 of 10? Don't re-run from zero.
from agent_vcr import VCRPlayer
from agent_vcr.models import ResumeConfig
player = VCRPlayer.load(".vcr/my_run.vcr")
# See exact state at every step
print(player.goto_frame(6)) # {'files_written': [...], 'plan': '...'}
print(player.get_errors()) # what broke and where
# Fix the prompt. Resume from step 6. Skip steps 0-5.
player.resume(
agent_callable=coder,
config=ResumeConfig(
from_frame=6,
state_overrides={"plan": "use SQLAlchemy instead of raw SQL"}
)
)
Without Agent VCR With Agent VCR
โโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโ
Agent fails step 8 Agent fails step 8
Patch the code player.goto_frame(7)
Re-run ALL 10 steps Fix the state
$0.04 + 2 min wasted Resume from step 7
Repeat for every bug Done. $0.00 extra.
Who This Is For
You need this if you're running:
- Claude Code / Cursor autonomous mode
- OpenHands on real codebases
- LangGraph agents that write files
- CrewAI pipelines with filesystem access
- Any autonomous coding agent on a repo you care about
You don't need this if you're only:
- Doing RAG / chatbots (no filesystem risk)
- Already happy with LangSmith for tracing
Quick Start
Record
from agent_vcr import VCRRecorder
recorder = VCRRecorder()
recorder.start_session("my_run")
state = {"query": "build a REST API"}
state = planner(state)
recorder.record_step("planner", input_state, state)
state = coder(state)
recorder.record_step("coder", input_state, state)
recorder.save() # โ .vcr/my_run.vcr
Or use the context manager โ frames are saved even if the agent crashes:
with VCRRecorder() as recorder:
recorder.start_session("my_run")
# ... your agent code ...
Rewind & Fix
player = VCRPlayer.load(".vcr/my_run.vcr")
diff = player.compare_frames(5, 6)
# {'added': {'bad_file': '...'}, 'modified': {'plan': '...'}}
player.resume(
agent_callable=coder,
config=ResumeConfig(from_frame=5, state_overrides={"plan": "fixed"})
)
Integrations
LangGraph โ one line
from langgraph.graph import StateGraph
from agent_vcr import VCRRecorder
from agent_vcr.integrations.langgraph import VCRLangGraph
recorder = VCRRecorder()
graph = VCRLangGraph(recorder).wrap_graph(graph) # โ one line, that's it
result = graph.invoke({"query": "Build a todo app"})
recorder.save()
CrewAI
from agent_vcr.integrations.crewai import VCRCrewAI
recorder = VCRRecorder()
recorder.start_session("crew_run")
result = VCRCrewAI(recorder).kickoff(crew)
recorder.save()
pip install "ai-agent-vcr[crewai]"
pip install "ai-agent-vcr[langgraph]"
Raw Python (decorator)
from agent_vcr.integrations.langgraph import vcr_record
@vcr_record(recorder, node_name="research_step")
def research(state: dict) -> dict:
return {"findings": search(state["query"])}
Sentinel โ Real-Time Code Guardian
Catches what the agent wrote before it moves to the next step.
from openhands_sentinel import Sentinel
sentinel = Sentinel(recorder=recorder)
sentinel.attach(runtime.event_stream) # 3 lines. auto-intercepts every write.
STEP 2: Agent writes handlers.py
๐ก๏ธ SENTINEL: VIOLATIONS DETECTED
CRITICAL hash_password() already exists in auth/utils.py:8 โ reuse it
CRITICAL handle_auth_request() is 109 lines (limit: 40) โ break it up
CRITICAL Cyclomatic complexity: 32 (limit: 8)
STEP 3: Agent self-corrects
๐ก๏ธ SENTINEL: handlers.py โ CLEAN โ
| Without Sentinel | With Sentinel | |
|---|---|---|
| Agent writes bad code | โ | โ |
| Sentinel catches it | โ | < 10ms |
| Agent self-corrects | โ | done |
| Human reviews PR | manual | zero |
| Cost | 2ร LLM + human time | 1 extra LLM call |
Standalone scan:
sentinel scan ./my-ai-project
sentinel watch ./my-ai-project
TUI Debugger
vcr .vcr/my_run.vcr
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ผ Agent VCR Session: my_run ยท 8 frames โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โถ Frame 0 โ planner โ 100ms โ โ โ
โ Frame 1 โ researcher โ 250ms โ โ โ
โ Frame 2 โ coder โ 480ms โ โ ERROR โ
โ Frame 3 โ tester โ 80ms โ โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ { "query": "build a todo app", "plan": null } โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ/โ navigate โ e edit โ d diff โ r resume โ q quit โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Keybindings: โ/โ or j/k navigate ยท e edit state ยท 1/2/3 input/output/diff ยท r resume ยท s search ยท q quit
Claude Code hooks:
vcr init --claude-code
DAG Visualization
pip install "ai-agent-vcr[dashboard]"
vcr-server --vcr-dir .vcr --auth-token local-secret
# http://127.0.0.1:8000
original_run โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโบ [done]
โ frame 3
โฐโโโบ fork_v1 โโโบ [coder] โโโบ [tester] โโโบ [done]
โฐโโโบ fork_v2 โโโบ [coder] โโโบ [done]
Live WebSocket streaming. Every fork is a branch. Errors in red.
vs Everything Else
Honest take: LangSmith, Langfuse, Arize Phoenix, and AgentOps are serious platforms with large teams. They are observability tools โ they show you what happened. Agent VCR is an intervention tool โ it lets you change what happened. Different category. The overlap is tracing. Everything else diverges.
| Capability | ๐ผ Agent VCR | LangSmith | LangFuse | AgentOps | Arize Phoenix |
|---|---|---|---|---|---|
| Record execution traces | โ | โ | โ | โ | โ |
| Production dashboards | Local | โ best-in-class | โ | โ | โ |
| Eval / scoring pipelines | โ | โ | โ | โ | โ |
| Time-travel / session replay | โ | โ | โ | โ (view only) | โ |
| Edit state & resume mid-chain | โ | โ | โ | โ | โ |
| ACID filesystem rollback | โ | โ | โ | โ | โ |
| Ghost Replay (zero tokens) | โ | โ | โ | โ | โ |
| Sentinel (real-time code guard) | โ | โ | โ | โ | โ |
| Fork from any frame | โ | โ | โ | โ | โ |
| TUI debugger | โ | โ | โ | โ | โ |
| Fully local / self-hosted | โ | โ Cloud | โ | โ Cloud | โ |
| Framework-agnostic | โ | โ ๏ธ LangChain | โ | โ | โ |
AgentOps โ closest competitor on time-travel. It lets you view past sessions. It does not let you edit state and resume, fork a session, rollback the filesystem, or replay for zero tokens. If you need view-only replay, AgentOps is mature. If you need to actually intervene, you need Agent VCR.
Use LangSmith/Langfuse/Phoenix/AgentOps for production tracing and evals. Use Agent VCR when you need to actually fix a broken run without re-running it, rollback filesystem damage, or replay a successful run for free.
vs LangGraph's Built-In Checkpointer
LangGraph's checkpointer is solid if you're 100% LangGraph and only need state inspection.
The gap: when your agent writes files to disk and fails, the checkpointer rolls back the state object. The files stay. Agent VCR runs git reset --hard. The files are gone.
| LangGraph Checkpointer | Agent VCR | |
|---|---|---|
| Checkpoint in-memory state | โ | โ |
| Rollback files from disk | โ | โ |
| Ghost Replay (zero tokens) | โ | โ |
| Sentinel (code guardian) | โ | โ |
| Works with CrewAI, raw Python | โ | โ |
| JSONL format (git-diffable) | โ | โ |
| Session forking | โ | โ |
Performance
Every benchmark is enforced in CI. If it regresses, CI fails.
pip install -e ".[dev]"
pytest tests/benchmarks/ -v --benchmark-only
| Benchmark | Limit | What it measures |
|---|---|---|
test_benchmark_recorder_overhead |
< 5ms mean | Serialize + buffer one state snapshot |
test_benchmark_file_write_speed |
> 1,000 frames/sec | Sustained write throughput (10K frames) |
test_benchmark_load_speed |
< 500ms | Load a 10,000-frame session from disk |
test_benchmark_goto_frame |
< 1ms | Random-access time-travel to any frame |
Historical results: ixchio.github.io/agent-vcr/dev/bench/
Storage Format
Plain JSONL. One object per line.
{"type": "session", "data": {"session_id": "my_run", "created_at": "..."}}
{"type": "frame", "data": {"node_name": "planner", "input_state": {...}, "output_state": {...}}}
{"type": "frame", "data": {"node_name": "coder", ...}}
- Human-readable โ open in any text editor
- Git-diffable โ review agent state in PRs
- Append-only โ safe for concurrent agents, no full-file rewrites
- Streamable โ parse line-by-line without loading the full file
API Reference
VCRRecorder
recorder = VCRRecorder(
output_dir=".vcr",
auto_save=True,
diff_mode=False,
)
recorder.start_session(session_id="my_run", tags=["prod"])
recorder.record_step(node_name, input_state, output_state, metadata)
recorder.record_llm_call(model, messages, response, tokens_input, tokens_output, latency_ms)
recorder.record_tool_call(tool_name, tool_input, tool_output, latency_ms)
recorder.record_error(node_name, input_state, error)
recorder.save() -> Path
recorder.fork(from_frame=3) -> VCRRecorder
VCRPlayer
player = VCRPlayer.load(".vcr/my_run.vcr")
player.goto_frame(index) # โ output state at frame N
player.get_input_state(index) # โ input state at frame N
player.get_errors() # โ [Frame, ...]
player.compare_frames(a, b) # โ {'added': {}, 'removed': {}, 'modified': {}}
player.get_total_cost() # โ float (USD)
player.resume(
agent_callable,
config=ResumeConfig(
from_frame=7,
state_overrides={"k": "v"},
mode=ResumeMode.FORK, # FORK | REPLAY | MOCK
)
)
ACIDWorkspace
acid = ACIDWorkspace("/workspace", recorder=recorder)
acid.begin(session_id="task-001")
acid.savepoint(state, node_name="coder")
acid.rollback(to_frame_index=2) # git reset --hard
acid.commit()
GoldenRunCache
cache = GoldenRunCache(cache_dir=".vcr/golden")
identity = GoldenRunCache.build_identity(model="gpt-4o", code_commit="abc123")
cache.save_golden_run(task_description, recorder, identity=identity)
outputs, ledger = cache.replay(task_description, identity=identity)
cache.invalidate(task_description)
cache.list_golden_runs()
Examples
# ACID rollback + Ghost Replay โ start here
python examples/acid_golden_run.py
# Time-travel: rewind, edit state, resume
python examples/time_travel_demo.py
# Sentinel: watch agent self-correct in real time
python examples/sentinel_demo.py
# LangGraph auto-instrumentation
python examples/langgraph_integration.py
# Basic recording and playback
python examples/basic_usage.py
Roadmap
- Core recording and playback
- Time-travel resume with state injection
- LangGraph + CrewAI integrations
- Async recorder and player
- Terminal TUI debugger (
vcr) - Claude Code hook scaffolding (
vcr init --claude-code) - Live dashboard with DAG visualization
- ACID Transactions (git-backed filesystem rollback)
- Ghost Replay (zero-cost replay of successful runs)
- Sentinel โ real-time code quality guardian
- Context manager (
with VCRRecorder() as r:) - Claude Code / Cursor integration
- AutoGen integration
- Replay regression tests (golden paths as CI assertions)
- Collaborative debugging (share sessions)
- Cloud storage backend (S3, GCS)
Community
If Agent VCR saved your repo from a bad autonomous run, share it:
- OpenHands โ Discord
#toolschannel - LangGraph โ Discord
#communitychannel - r/LocalLLaMA โ post your ACID rollback story
- Hacker News โ Show HN posts with real before/after diffs get traction
The best growth comes from developers sharing the moment it saved them. If that's you, a post with your actual corrupted-repo story (even anonymized) is worth more than any ad.
Contributing
git clone https://github.com/ixchio/agent-vcr.git
cd agent-vcr
pip install -e ".[dev,tui]"
make verify
See CONTRIBUTING.md for guidelines.
License
MIT โ see LICENSE.
๐ผ
Every AI agent developer has had a bad run trash their codebase. Agent VCR is the undo button.
pip install ai-agent-vcr
โญ Star on GitHub ยท ๐ฆ PyPI ยท ๐ Docs
Built by ixchio ยท MIT License
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai_agent_vcr-0.8.0.tar.gz.
File metadata
- Download URL: ai_agent_vcr-0.8.0.tar.gz
- Upload date:
- Size: 924.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a8a6d67bf75d80d87ae007554bf2264e2852acabbc205c216257bfc05d004320
|
|
| MD5 |
b0a131bcf7422b2a697bff3541982307
|
|
| BLAKE2b-256 |
b3c0d15179b25e18a264070b9d2d14a71f19b86e2e39942d518969228a188751
|
File details
Details for the file ai_agent_vcr-0.8.0-py3-none-any.whl.
File metadata
- Download URL: ai_agent_vcr-0.8.0-py3-none-any.whl
- Upload date:
- Size: 170.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0f371e4cea35e19744f000773c670124f10ee2e4ad606a4c6b80e90c6bb5f339
|
|
| MD5 |
81efee7be392c73e72223e0d4a1c8ce3
|
|
| BLAKE2b-256 |
a8adc67a8451af256be3859099bc369083c327d61706b5c53b17dce5c1e35054
|