Skip to main content

BunkerVM

BunkerVM

Time-travel debugging for AI agent sandboxes.
Hardware-isolated Firecracker microVMs with snapshot, replay, and diff — not containers.

PyPI CI Stars Isolation Python License

Sandbox variable x goes from 1100 to 11 after sb.restore(2) — a real VM rewind, not a re-run

That's a real run: three commands mutate x to 1100, one line rewinds the sandbox to step 2, and x is 11 again — actual state restored, not the script re-executed. That's what this repo does. Works on macOS too — see Two ways to run it.


Is this for you?

If you're building with LangChain, LangGraph, the OpenAI Agents SDK, or an MCP client like Claude Desktop or VS Code Copilot, and you've ever stared at an agent that did fifteen things and failed, with zero visibility into what happened or a way back to before it broke — yes. BunkerVM gives every sandboxed run a rewind button and a diff tool, on your own machine, for free.

If you need managed infrastructure for thousands of concurrent sandboxes, this isn't that — see Why not E2B / Daytona / Modal? below.

The problem

AI agents execute code on your machine. When something goes wrong — and it will — you have no way to see what the agent actually did, rewind to the moment before it broke, or tell why one agent succeeded and another failed on the same task.

Most debugging here means re-running and hoping, or reading a transcript the agent wrote about itself. BunkerVM records every step as it actually happens — commands, exit codes, filesystem changes — so you can rewind to any point and inspect real state, not a self-report.

It also happens to run each sandbox in a hardware-isolated Firecracker microVM when your machine supports it (same tech as AWS Lambda) — because containers share your kernel and cloud sandboxes send your data to someone else's server. On machines that can't do that (macOS, plain Windows), a local no-isolation fallback keeps the record/rewind/diff workflow available — see below.


Two ways to run it

Firecracker (Sandbox()) Local (Sandbox(backend="local"))
Isolation Hardware (KVM microVM) None — a plain subprocess
Platforms Linux, Windows+WSL2 Anywhere Python runs, incl. macOS
Record / rewind / diff ✅ full VM state ✅ namespace + working directory
Setup /dev/kvm or WSL2, ~100MB bundle Nothing — pip install and go
Use for Running agent-generated code you don't fully trust Trying the workflow, debugging on a machine without KVM

The local backend is never selected automatically — you have to ask for it (backend="local", --local), and BunkerVM tells you which one is active every time. It exists because the record/rewind/diff value doesn't require a hypervisor, only the isolation does — and a lot of development happens on machines that can't run one.

bunkervm demo --local     # works everywhere, no KVM/WSL2 needed
bunkervm demo             # real hardware isolation (Linux, or Windows+WSL2)

What it does

In Firecracker mode, each sandbox is a Firecracker microVM — the same technology behind AWS Lambda. Own kernel, own filesystem, hardware-level (KVM) isolation. Not a container. The examples below use this mode; swap in backend="local" and everything except true VM-level restore works the same way.

On top of that, BunkerVM adds capabilities that no other sandbox provides:

Record every execution

from bunkervm import Sandbox

with Sandbox(record=True) as sb:
    sb.run("import pandas as pd")
    sb.run("df = pd.read_csv('/data/input.csv')")
    sb.run("df['total'] = df.price * df.qty")
    sb.run("df.to_csv('/output/result.csv')")

# Every step recorded: command, output, filesystem changes, VM snapshot

Rewind to any point

sb.restore(step=2)  # VM state rewinds to after read_csv
sb.run("df.describe()")  # explore from that exact point

The VM's memory, CPU registers, filesystem — everything reverts to exactly what it was after step 2. Not a re-run. An actual restore from a Firecracker snapshot.

See what changed

for cp in sb.history():
    print(f"step {cp['step']}: {cp['command']}")
    if cp['trace']:
        for f in cp['trace']['files_created']:
            print(f"  + {f['path']} ({f['size']} bytes)")
step 1: import pandas as pd
step 2: df = pd.read_csv('/data/input.csv')
  ~ /data/input.csv (read)
step 3: df['total'] = df.price * df.qty
step 4: df.to_csv('/output/result.csv')
  + /output/result.csv (1247 bytes)

Compare two agents

Every recorded session gets an auto-generated ID (printed when the sandbox exits, or via sb.session_id). Run the same task through two agents, then:

bunkervm diff d0c13cb74d85 f29a61bb02e7
Agent Diff
  Session A: d0c13cb74d85  (12 steps, 3400ms)
  Session B: f29a61bb02e7  (8 steps, 1200ms)

  Files only in A:  /tmp/debug.log, /tmp/retry_3.py
  Files only in B:  /output/result.csv

  step  1  [same]  import pandas as pd
  step  2  [same]  df = pd.read_csv('/data/input.csv')
  step  3  [diff]
    A: df = df.dropna()
    B: df = df.fillna(0)
  step  4  [diff]
    A: # crashed — KeyError: 'total'
    B: df['total'] = df.price * df.qty  ← OK

Agent A dropped rows and lost a required column. Agent B filled missing values and succeeded. Without diff, you'd never know why.

Rank multiple agents

diff is pairwise. To score and rank several runs at once — which model, which prompt, which agent actually did the job:

bunkervm compare gpt4-run claude-run llama-run --html report.html
Agent Comparison

  #1  claude-run  [direct]  6 steps  completed   1900ms
      files: +1 created  ~0 modified  -0 deleted
  #2  gpt4-run    [direct]  8 steps  completed   3400ms  (1 destructive/blocked)
      files: +1 created  ~0 modified  -1 deleted
  #3  llama-run   [direct]  4 steps  failed (step 3)  1100ms
      files: +0 created  ~0 modified  -0 deleted

  Divergence from baseline (gpt4-run):
    claude-run: diverged at step 2
    llama-run: diverged at step 3

  Ranked by: completed without a failed step, then fewest destructive/blocked commands, then total time.

Every column is a fact already captured by record=True — exit codes, timing, the safety classifier's risk tier for each command that ran, and the filesystem trace. There's no LLM judge and no rubric to configure: this grades what each agent actually did, not a transcript it wrote about itself. --html renders the same data as a shareable report.


Quick start

pip install bunkervm
from bunkervm import run_code

result = run_code("print('Hello from a microVM!')")
print(result)  # Hello from a microVM!

VM boots, code runs, VM dies. Your host was never touched.


How it works

AI Agent
   │
   ▼
bunkervm (host)  ──vsock──▶  Firecracker MicroVM
   │                          ┌────────────────────┐
   │  record=True             │  Alpine Linux       │
   │  ─────────▶              │  Own kernel         │
   │  snapshot()              │  exec_agent.py      │
   │  trace()                 │  (filesystem trace) │
   │  restore()               └────────────────────┘
   │                          KVM hardware isolation
   ▼
~/.bunkervm/sessions/         ~/.bunkervm/snapshots/
  d0c13cb74d85.json             d0c13cb74d85-step1/ vmstate + memory
                                 d0c13cb74d85-step2/ vmstate + memory

Firecracker provides the isolation. BunkerVM adds the instrumentation layer:

Layer What it does
exec_agent (inside VM) Traces filesystem changes per command — files created, modified, deleted, bytes written
Firecracker API (host→VM) Pauses VM, snapshots CPU + memory state to disk, resumes — all via Firecracker's built-in snapshot API
Snapshot manager (host) Stores and indexes snapshots at ~/.bunkervm/snapshots/, manages lifecycle
Session recorder (host) Chains commands → traces → snapshots into a replayable session JSON

No custom kernel modules. No eBPF. No ptrace. The VM is the isolation boundary; the API socket is the control plane. Pure Python, stdlib-only transport.


Named checkpoints & replaying a session

restore(step=N) rewinds to an auto-recorded step. For a checkpoint you want to name and return to deliberately — e.g. right after a slow setup step — use checkpoint():

with Sandbox() as sb:
    sb.run("import torch; model = torch.load('bert.pt')")
    sb.checkpoint("model-loaded")        # snapshot: 45ms
    sb.run("output = model(bad_input)")  # crashes
    sb.restore(step=1)                   # restore: <100ms
    sb.run("output = model(good_input)") # works

Every record=True session is saved to ~/.bunkervm/sessions/<id>.json on exit and can be replayed from the CLI, independent of the process that created it:

bunkervm replay d0c13cb74d85 --trace
Session: d0c13cb74d85
  Steps: 5
  Recorded: 2026-03-29 23:15

     step   1  [ok]      34ms  x = 42
     step   2  [ok]      23ms  print(x * 2)
     step   3  [ok]      22ms  import os; os.makedirs('/tmp/output', exist_ok=True)
     step   4  [ok]      21ms  open('/tmp/output/result.txt', 'w').write(str(x))
     step   5  [ok]      21ms  print(open('/tmp/output/result.txt').read())

Why not E2B / Daytona / Modal?

Those are hosted sandbox platforms — good at giving your agent a place to run. BunkerVM is a local, self-hosted debugger for whatever sandbox your agent already runs in. As of writing, none of the major hosted sandboxes ship automatic action recording, mid-session VM snapshot/restore, and cross-run diffing together:

BunkerVM E2B / Daytona / Modal
Isolation Firecracker microVM (hardware/KVM) Firecracker or container, depending on provider
Hosting Local, self-hosted — nothing leaves your machine Cloud-hosted
Auto-records every command ❌ (manual snapshot primitives at best)
Mid-session restore ✅ full VM state (memory + fs) Fork-from-snapshot, not automatic rewind
Diff two agent runs bunkervm diff
Cost Free, open source Usage-billed

Trade-off: it won't scale to thousands of concurrent sandboxes the way a hosted platform will. If you need managed multi-tenant infra, use one of those. If you need to see exactly what your agent did and rewind to before it broke, that's what this is for — with real hardware isolation where your machine supports it (/dev/kvm or WSL2), or the same record/rewind/diff workflow with no isolation anywhere else, including macOS.


Integrations

MCP (Claude Desktop, VS Code Copilot, any MCP client)

bunkervm vscode-setup     # generates .vscode/mcp.json, works on Windows WSL2
bunkervm server            # stdio for Claude Desktop
bunkervm server --transport sse  # SSE for web

8 MCP tools: sandbox_exec, sandbox_write_file, sandbox_read_file, sandbox_list_dir, sandbox_upload_file, sandbox_download_file, sandbox_status, sandbox_reset.

Any agent framework

secure_agent() wraps a single-tool adapter around whatever you already have, no BunkerVM-specific toolkit required:

from bunkervm import secure_agent

runtime = secure_agent()
tool = runtime.as_tool()          # LangChain-compatible tool (requires langchain-core)
tool = runtime.as_openai_tool()   # OpenAI Agents SDK tool (requires openai-agents)

Install

pip install bunkervm
bunkervm demo --local     # macOS / no KVM — works immediately, no download
bunkervm demo             # real hardware isolation — Linux, or Windows+WSL2

For hardware isolation: Linux with /dev/kvm, or Windows WSL2 (enable nested virtualization). Python 3.10+. The Firecracker binary + kernel + rootfs (~100MB) auto-download on first run, or download from Releases.

For the local backend: nothing beyond Python 3.10+. No isolation — see Two ways to run it.

WSL2 setup (Windows)

Add to %USERPROFILE%\.wslconfig:

[wsl2]
nestedVirtualization=true

Then: wsl --shutdown

Troubleshooting
Problem Fix
/dev/kvm not found sudo modprobe kvm or enable nested virtualization
Permission denied sudo usermod -aG kvm $USER then re-login
Bundle download fails Manual download from Releases~/.bunkervm/bundle/
VM won't start bunkervm info — diagnoses all prerequisites
Build from source
git clone https://github.com/ashishgituser/bunkervm.git
cd bunkervm
sudo bash build/setup-firecracker.sh
sudo bash build/build-sandbox-rootfs.sh
pip install -e ".[dev]"
pytest tests/

CLI

bunkervm demo                              # see it in action (real isolation)
bunkervm demo --local                      # see it in action (no KVM needed — macOS, etc.)
bunkervm run script.py                     # run a script in a sandbox
bunkervm run -c "print(42)"               # inline code
bunkervm run script.py --local             # run without isolation, no KVM/WSL2 required
bunkervm replay <session-id> --trace       # replay recorded session
bunkervm diff <session-a> <session-b>      # compare two agent runs
bunkervm compare <a> <b> <c> --html out.html  # rank multiple agent runs
bunkervm snapshot list                     # list VM snapshots
bunkervm snapshot delete <name>            # delete a snapshot
bunkervm server --transport sse            # MCP server
bunkervm info                              # system readiness check

Contributing

See CONTRIBUTING.md.

Security

See SECURITY.md.

License

MIT


If BunkerVM helps you build safer agents, star the repo

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bunkervm-0.11.0.tar.gz (128.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bunkervm-0.11.0-py3-none-any.whl (122.9 kB view details)

Uploaded Python 3

File details

Details for the file bunkervm-0.11.0.tar.gz.

File metadata

  • Download URL: bunkervm-0.11.0.tar.gz
  • Upload date:
  • Size: 128.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for bunkervm-0.11.0.tar.gz
Algorithm Hash digest
SHA256 f8f9254caf0b5e5e54f30c58747060ad52480f375f3d877a533e5f9c57824359
MD5 ef9d8d4626cfc91fe8a2dab136778cd7
BLAKE2b-256 7e9192150bc9d211fc5865a35c6274ae9167619be642907ad84f2d43340fd429

See more details on using hashes here.

File details

Details for the file bunkervm-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: bunkervm-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 122.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for bunkervm-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 1766b288d2e374ed8885632702d0379869d403699e1b73b9f57f424a397a3cfa
MD5 27d0e25c01009af73c3d50e14b252e77
BLAKE2b-256 ff1a79b75f833839b9d8cdbaa1a342048e6e3e1a2211a9f671160d4ba7d1b688

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page