Skip to main content

Agent Seatbelt

A seatbelt for AI coding agents. One self-contained Python file that reads every tool call before it runs — shell commands, file writes, MCP calls — and denies or escalates the ones that end careers: recursive deletes, secret exfiltration, production deploys, agent-bypass launches, memory poisoning. Every verdict names its rule, its blast radius, and the safer path.

Seatbelt Bench v1   206/206 cases   ·   GuardFall fuzz 1431/1431
Naive rm-rf regex    42/206         ·   Injection 15/15, 0 FP

No network. No telemetry. No dependencies. The hook never phones home — check tests/test_no_network.py, it AST-scans the source.

Install (verify, then run)

Read install.py before you run it — it's short, and installing a security tool you haven't read is how the incidents in docs/INCIDENTS.md happen.

# 1. Verify the download against the release's SHA256SUMS
sha256sum -c SHA256SUMS

# 2. Install as a Claude Code plugin
claude plugin marketplace add littlestjames82-sys/agent-seatbelt
claude plugin install agent-seatbelt@agent-seatbelt

Or wire the hook directly into ~/.claude/settings.json:

python3 install.py --dry-run   # shows the diff first
python3 install.py

Optional extras: --harden adds a native permissions deny/ask layer alongside the hook; --canaries plants decoy credentials (see below). Team rollout: docs/TEAM_ADOPTION.md.

See it work

$ python3 hooks/seatbelt_hook.py --check "rm -rf ~/project"
DENY  rule=seatbelt-locked-rm-rf
  reason: recursive forced deletion (rm with -r and -f) is irreversible
  Safer path: List the targets first (ls), move them to a
  quarantine/trash directory, and delete only after checking.

Full captured demo — selftest, an ask with a rehearsal preview, an off-plan stop, a brain-drift report: docs/DEMO.md. Try the policy in your browser: docs/playground.html (open from disk).

What it does

  • Deobfuscation before judgment. Quote fragments (r"m" -rf), $IFS, backslash escapes, command substitution, base64 blobs, interpreter -c code, and aliases — npm run deploy is judged by what the script actually runs. The 2026 GuardFall bypass shapes are the test corpus, not a surprise.
  • Locked rules, strictest wins. Bash + PowerShell destruction, git destruction, cloud/DB deletion, deploys/publishes, pipe-to-shell, secret reads, secret+network exfiltration, sudo, persistence, reverse shells, tunnels, TLS bypass, crypto mining, agent-bypass launches. Project overlays and policy packs (solo / team-strict / ci / paranoid) can only add.
  • Blast radius in every serious reason — git status counts, rm target counts, symlink/junction escapes out of the target tree.
  • Snapshots before destruction. Ask-tier destructive actions first copy their enumerable targets (bounded: 2,000 files / 256 MB, caps recorded honestly in the manifest). One command restores — additively, hash-verified (docs/SNAPSHOTS.md).
  • Injection flagging. Tool output and pasted text are scored for agent-directed instructions; a flag taints the session for 30 minutes and escalates the next destructive/exfil verdict, naming the source (docs/INJECTION.md).
  • Memory drift watch. Your CLAUDE.md/AGENTS.md/settings are baselined; SessionStart reports drift only when it exists, quoting instruction-shaped added lines verbatim. The baseline moves only by a human running --baseline --accept (docs/BRAIN.md).
  • MCP rug-pull watch. Server configs (env key names only — never values) and tool surfaces are fingerprinted at baseline; drift escalates MCP verdicts until a human re-baselines.
  • Judge escalation seam (opt-in, off by default). Point Seatbelt at a local judge command you configure: on allow/ask verdicts it may raise the verdict exactly one tier — and it can never lower, clear, or suppress anything. Malformed, slow, or verdict-shaped replies are ignored and audit-logged; the deterministic verdict is always the floor (docs/JUDGE.md).
  • Skill & plugin drift watch. A human-filed baseline of sha256+size fingerprints (never file contents) for installed skills, commands, hooks, and plugin manifests. Drift is reported at SessionStart and via --skills, audit-logged, and — in strict/CI mode only — escalates the first call after the change one tier. Detection, not prevention (docs/SKILL_DRIFT.md).
  • Flight plans. Declare the session's scope (.seatbelt/plan.json): ask-tier actions inside it stop prompting; denials and locked rules are never covered; off-plan consequential actions ask (solo) or deny (ci/paranoid).
  • Rehearsal previews. Ask-tier infra commands run their honest dry-run twin (terraform plan, git push --dry-run, kubectl --dry-run=client…) and the ask carries a Preview: line. Previews never change verdicts.
  • Canary honeytokens. Opt-in decoy credentials (install.py --canaries); any read of one, or any command carrying one's value, denies in every mode. Registry stores hashes only.
  • Tamper-evident audit. Every decision is written first to .seatbelt/audit.jsonl — hash-chained, sanitized, secret-redacting (type+length, never values). --verify-log walks the chain; --report --evidence emits the compliance bundle (docs/COMPLIANCE.md).
  • Self-protection. The agent cannot edit the hook, the policies, the audit log, the baselines, or its own governance files without a human. Reads of them stay open — transparency cuts both ways.

Snapshots and injection flags are seatbelts, not roll cages: bounded, local, and honest about what they can't catch.

Agents

One core, per-agent envelopes (full matrix: docs/AGENTS.md). All non-Claude surfaces are simulated-tested, not live-tested.

Agent Ask semantics Notes
Claude Code True ask Reference surface (plugin)
Codex ask→deny under strict/CI Hooks can fail open on ask
Cursor True ask preToolUse contract
Cline Ask delivered as block for review Documented hooks; host fails open on errors
Gemini LEGACY Consumer CLI retired 2026-06-18 → Antigravity (agy); support pending Google's documented contract
OpenCode Ask delivered as block Experimental JS plugin shim

Built from real disasters

Each incident below is written up in docs/INCIDENTS.md and reproduced as bench cases (sources cited there):

Incident Seatbelt answer Bench case
Home-directory wipe (Claude Code, Dec 2025) rm -rf deny + blast radius incident-home-wipe
Wrong-folder photo deletion (Cowork, Feb 2026) Delete asks + snapshots incident-cowork-photos-rm
Alias pushed a DB change to prod (Replit, Jul 2025) Alias resolution + prod signals incident-replit-dbpush-alias
Over-scoped token deleted a prod volume (PocketOS, Apr 2026) MCP gating + flight plans incident-pocketos-mcp-volume
Junction cleanup ate a live tree (Sep 2026) Link-aware blast radius incident-junction-rm
Terraform applied to production (DataTalks, Feb 2026) Prod-target escalation incident-terraform-prod
Supply-chain payload hired an agent (Nx, Aug 2025) Agent-bypass deny incident-nx-agent-bypass
4.6-hour retry loop (Dec 2025) Loop guard incident-loop-sequence

Scope honesty: Seatbelt gates the tool-call path only. The PocketOS deletion, for instance, arrived through an MCP tool call — which is why MCP calls are gated — but no hook can govern what a separately-credentialed service does on its own.

How it compares (honestly)

  • Single-regex rm blockers are the common baseline; the bench scores one at 20.4%. They're better than nothing and blind to everything else.
  • Shellter covers PowerShell/cmd and persistence well; Seatbelt matches that ground and adds deobfuscation, snapshots, and the memory/injection layers.
  • Anthropic's security-guidance plugin reviews code patterns; Seatbelt gates actions at runtime — different layer, composable.
  • Classifier layers (e.g. Cursor's Auto-review) are, in Cursor's own documentation's words, "steering, not enforcement" and "not a security boundary". Classifiers advise; Seatbelt is the deterministic gate underneath.
  • AgentGuards and similar policy tools overlap on policy gating; Seatbelt's distinguishing bets are the public bench, the snapshot/restore pair, and the memory-drift watch.

No "first" or "only" claims: the field is active and good. The bench is the argument — run your current tool through bench/run_external.py and compare.

Known limits

  • Deterministic pattern policy, not a sandbox: a determined adversary with arbitrary code execution can route around any hook. Seatbelt raises the cost of accidents and common attack shapes; it is not a kernel boundary (see docs/THREAT_MODEL.md).
  • Snapshots cover enumerable local files only — not databases, remote state, or shell-expanded globs.
  • Injection flagging is a heuristic scorer with a measured false-positive budget, not a proof.
  • Claude Mods (and the experimental experimental/ editions) run unsandboxed in-process and can approve what a hook blocked — treat mods as trusted code; the hook is the floor.

The Python library

The original 0.1 library remains: agent_seatbelt (policy kernel with strictest-wins tiers, locked rules, named decisions, human handoff, redacted audit), the governor, and a CLI.

pip install ghost-seatbelt
python3 -m agent_seatbelt --help

Docs map

docs/BENCH.md (all the numbers) · docs/DEMO.md · docs/INCIDENTS.md · docs/THREAT_MODEL.md · docs/AGENTS.md · docs/SNAPSHOTS.md · docs/INJECTION.md · docs/BRAIN.md · docs/JUDGE.md · docs/SKILL_DRIFT.md · docs/COMPLIANCE.md · docs/TEAM_ADOPTION.md · docs/FEED_RUNBOOK.md · docs/VERIFICATION.md · docs/CLAIMS_AUDIT.md · bench/README.md (score anything)

Slash commands: /agent-seatbelt:status, :report, :selftest, :check, :doctor, :restore, :plan, :skills.

MIT licensed. Contributions follow one law — no detector without bench cases (destructive AND benign near-miss) and a remediation string (CONTRIBUTING.md).


From Ghost Developer Studio, the makers of GhostGuard — the governance layer for autonomous systems.

Metadata

Release files for ghost-seatbelt 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ghost-seatbelt 0.3.0
File Size Uploaded
ghost_seatbelt-0.3.0.tar.gz 515.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ghost-seatbelt 0.3.0
File Interpreter ABI Platform
ghost_seatbelt-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 533.9 kB

Release files / ghost_seatbelt-0.3.0.tar.gz

Download URL ghost_seatbelt-0.3.0.tar.gz
Size 515.1 kB
Tags Source
SHA-256 checksum
How to use checksums
c83e151074c128d413228a969d8ae546f7c0dc2754f0abc298cc9355871cb824
BLAKE2b-256 checksum
How to use checksums
37a5ae34c487b081944661a60dd499b7fa02bf18019049db060838eb5f01a0bf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / ghost_seatbelt-0.3.0-py3-none-any.whl

Download URL ghost_seatbelt-0.3.0-py3-none-any.whl
Size 18.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c141d72eaf786c6100a235579069ad30948ac40f1074a874376bd12b4d0f36e1
BLAKE2b-256 checksum
How to use checksums
a417aac631ad1307ad71e6451e6ea0f1be713605d5eb7547c8ca57ab460a4ea1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.1

2 release files

This release

0.3.0 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page