agent-guard
Behavior-guardrail hooks for Claude Code. Six guards plus test-honesty CLIs, one install, zero dependencies (Python standard library only):
- Test-tampering guard — stops the "green by editing the test" cheat.
On
SessionStartit snapshots hashes of every test and source file; onStopit diffs. If test files were modified or deleted while no source file changed, the stop is blocked until the agent proves each changed test actually fails without the fix. - Outbound-action guard — a
PreToolUsehook onBashwith a denylist of risky action patterns: pushes to protected branches (including force pushes), package publishes (npm publish,twine upload, …), prod deploys, cloud provisioning (spend), and mass-send channels (Slack webhooks, mailers). Matches block with a named rule; a user allowlist inconfig.jsonoverrides the denylist. Every matched decision is written to an audit log. v0.2 also scans the content of script files the command executes (bash evil.sh), closing the "write it to a script first" bypass. - Post-exec read-back verifier (v0.2) — a
PostToolUsehook onBashthat reads back world state after a command claimed an outbound effect (git push,npm publish) and warns — never blocks — when the effect isn't visible. Defense-in-depth on top of PreToolUse prevention: prevent first, verify after. - Mutation test-honesty checker (v0.2) —
agent-guard mutate-checkdeliberately breaks assertions (regex-based mutants for JS/TS and Python), re-runs the tests per mutant, and reports survivors: tests that stayed green don't actually cover the bug. - Comment-slop guard (v0.3) — a
PreToolUsehook onWrite/Editthat scores the added comments (never pre-existing code) for narrative slop: commented-out code, restatements of obvious code ("This function adds two numbers"), in-code changelogs, meta/apologetic notes, emoji, and docstrings that just restate the signature. Blocks at a configurable threshold, oragent-guard decomment --check/--fixfor a one-command decomment pass before PRs. - Cheat-sniffing guard (v0.4) — catches the cheating that never touches
test files: RNG rigging (
random.seed(123)with no reproducibility marker, patchingrandom.shuffle), mocking the function under test instead of its collaborators,conftest.pyplants, time-freezing, and always-True comparison dunders. APreToolUsehook onWrite/Editblocks cheat patterns in added test text;agent-guard cheatsniff --checkaudits the repo. - Cross-tool write guard (v0.5) — stops Bash from bypassing the
Edit/Write hooks. A
PreToolUsehook onBashstatically extracts file-write targets (>,>>, heredocs,sed -i,tee,cp/mvdestinations) and applies the same policy the Write-tool guards would: test-ish targets get cheat-sniffed when the content is visible, opaque writes to protected files are blocked. A guardrail on one tool protects nothing if another tool can do the same thing.
The failure modes are real, quoted from the community:
"Instead of fixing the indexing logic in the source file, the agent quietly modified the test file: it changed
expect(page.items.length).toBe(10)totoBe(9), re-ran the test, saw green, and told us the refactor was complete." — navune, r/ClaudeCode
"the risk isn't a bad answer, it's a bad action" — dank_as_fuck_, r/AI_Agents, on why guardrails must live "below the prompt layer" (verstands)
Differentiation
- vs Rashomon (r/aiagents): Rashomon observes — it records commands, edits and failures, then compares them against the agent's closing summary. agent-guard prevents — the hook sits in the tool-call path and blocks the bad action before it happens. Observation and prevention are complementary layers; this is the prevention one.
- vs edit-guard (same author): edit-guard blocks stale cross-session edits (write-after-write on a file another session changed). agent-guard blocks misbehaving actions (weakened tests, unapproved outbound effects). Different failure modes, same fail-open hook philosophy.
- vs Kvitansiya (Show HN, 2026-10-01): Kvitansiya verifies at stop — it checks the claimed outcome once, when the session ends. agent-guard v0.2 adds post-exec read-back as defense-in-depth on top of PreToolUse prevention: block the bad action before it happens, then verify the claimed effect is actually visible afterwards. Prevent first, verify after.
v0.5: cross-tool write guard
agent-guard install # registers agent-guard hook-bashwrite (PreToolUse on Bash)
The second bypass in the same family. v0.2 closed "write it to a script
first" (the r/AI_Agents blocklist bypass); v0.5 closes the other one.
thomastartrau read 28 Claude Code security advisories and found the
pattern: "My hooks block certain writes through the Edit and Write
tools. Once blocked, the agent went through Bash instead: sed -i, a
heredoc, a redirection. I had to add a hook that blocks writes to source
files via Bash. A guardrail on one tool protects nothing if another tool
can do the same thing."
install registers agent-guard hook-bashwrite as a PreToolUse hook on
Bash. It statically extracts file-write targets from the command —
> / >> redirections, heredocs (<<EOF, <<-EOF), sed -i
(including -i.bak / --in-place), tee (with/without -a),
cp/mv/install destinations, chained with && / ; / | — and
applies the same policy the Write-tool guards would apply:
- test-ish target (
test_*.py,conftest.py,tests/…): when the written content is visible (heredoc body), it is cheat-sniffed with the v0.4 detectors; an opaque write (sed -i, bare>) to a test file is treated as a violation on its own — that is exactly the bypass shape. bash_write.protected_paths(path prefixes, default empty = the test-tampering guard's scope): any Bash write under a protected prefix is treated like a Write-tool call — visible content is scored (comment-slop for source-ish files), opaque writes are blocked in block mode.
Warn mode (AGENT_GUARD_BASHWRITE_MODE=warn or bash_write.mode=warn)
advises instead of blocking. A bash_write.allow list
("tests/legacy/:bash-write") covers the judgment calls you disagree
with. Everything fails open: unparsable commands are allowed, never
blocked.
v0.4: cheat-sniffing beyond test files
agent-guard cheatsniff --check tests/ # per-file cheat scores, exit 1 over threshold
Half of agent cheating never touches a test file. A dev.to study (remdore,
2026-10-01, 102 runs × 4 models) found agents patching the RNG "so the list
would always be sorted", mocking the function under test instead of its
collaborators, and planting helpers in conftest.py. The test-tampering
guard (tests-only-change diff) and mutate-check (assertion mutation) both
miss this family — "restore the test files and re-run" only catches the
dumb half.
install registers agent-guard hook-cheatsniff as a PreToolUse hook on
Write/Edit. It fires only for test-ish files (test_*.py,
*_test.py, conftest.py, anything under tests/) and scores only the
added text. Six cheat kinds, regex-based and deliberately conservative:
- mock-subject (severe):
mock.patch("billing.total")insidetest_billing.py— patching the module under test itself, not its dependencies. Patching a collaborator (stripe.Charge.create) is clean. - rng-patch (severe): patching
random.shuffle/random.random/random.sampleto force outcomes. - conftest-patch (severe):
conftest.pymonkeypatching the subject or other local modules — including hand-rolledmymod.shuffle = ...direct assignment (pure fixtures are clean). - rng-seed: a fixed
random.seed(123)with no reproducibility marker.random.seed(42)next to a "reproducible" comment is legitimate and not flagged — the marker is the whole difference, and it's documented. - time-freeze:
freeze_time(...),time.sleeppatched to a no-op. - weak-comparator: a
__eq__/__lt__/ … whose body unconditionallyreturn True.
One severe hit reaches the default threshold (30/100) on its own. Warn mode
(AGENT_GUARD_CHEAT_MODE=warn or cheat_sniff.mode=warn in config) advises
instead of blocking. A cheat_sniff.allow list ("test_sort.py:rng-seed",
"*/legacy/*:*") covers the judgment calls you disagree with. Everything
fails open.
v0.3: comment-slop guard
agent-guard decomment --check src/ # per-file slop scores, exit 1 over threshold
agent-guard decomment --fix src/app.py # remove commented-out code only (writes .bak)
install registers agent-guard hook-commentslop as a PreToolUse hook
on Write/Edit. It scores only the added comment lines — your existing
codebase is never punished — and blocks (exit 2) when the added-slop score
reaches the threshold (default 30/100). The block message names the lines
and the slop kind, and points at decomment --fix. Six slop kinds, all
regex-based and deliberately conservative (when in doubt, it doesn't flag):
- commented-code (weighted highest): 2+ consecutive comment lines that look like code — dead code left in comments instead of deleted
- restatement: "This function …", "Here we …", "The following …"
- changelog: "Fixed …", "Updated …" narrating the diff inside the code (belongs in the commit message)
- meta-apology: HACK, "sorry", "workaround", "fix this later",
!!! - emoji in comments
- obvious-doc: a docstring whose words are >70% covered by the next
line's identifiers (e.g.
"""Add a and b."""abovedef add(a, b))
Tool directives (# noqa, # type: ignore, eslint-disable, …) are never
flagged. --fix is surgical: it removes only commented-code blocks and
always writes a .bak backup first. Warn mode
(AGENT_GUARD_COMMENT_MODE=warn or comment_slop.mode=warn in config)
advises instead of blocking. Everything fails open.
v0.2: script-content inspection, mutation test-honesty, post-exec verification
Script-content inspection
The Bash denylist failed as soon as the agent wrote the destructive command
into a script and executed the script instead (bash evil.sh). The outbound
guard now extracts script files from the command — bash|sh|zsh|dash,
python|node|ruby|perl|php <file>, ./run.sh, source/. <file>, plus
bash -c '...' inline code — resolves them against the hook's cwd, and scans
the file content with the same rule logic (regexes + destination-aware
git push parsing). Hits block with a script-content:<rule_id> id so the
audit log shows where the hit came from. Missing, unreadable, or >1MB files
are skipped (fail open); only plausible script extensions are scanned; the
user allowlist and disabled_rules apply to content hits too.
Mutation test-honesty checker
agent-guard mutate-check tests/test_app.py --project-root . -- pytest -q
Generates up to 20 syntactic mutations (default; --max-mutations), copies
the project to a temp dir per mutant (skipping .git/node_modules/etc.),
and runs the test command there. Mutation operators:
- JS/TS:
toBe(<n>)→<n+1>;toEqual("<s>")→"<s>_mut";toBe(true)↔toBe(false);===→!==onexpect()lines - Python:
assert <e> == <n>→<n+1>;assert <e> != <n>→==;assert <name>→assert not <name>
A mutation the test suite still passes is SURVIVED — the assertion doesn't cover the bug — and the command exits 1. All killed → exit 0. Inconclusive runs (command missing, timeout) are reported separately and never count as survived. Weird input fails open with a warning.
Post-exec read-back verifier
install registers agent-guard hook-verify as a PostToolUse hook on
Bash. After a command that claimed an outbound effect, it reads back world
state and warns on stderr (always exit 0 — it never blocks):
git push <remote> <ref>: runsgit ls-remote <remote> <ref>(10s timeout); warns when the ref is absent remotely.<remote>defaults toorigin;-C <dir>is honored; unreachable remotes stay silent.npm publish: readsname/versionfrompackage.jsonand runsnpm view <name>@<version> version; warns when the version isn't visible. Silent when there is nopackage.jsonor npm is missing.
Warnings are audit-logged (guard: verify, decision: warn).
Install
# Not on PyPI yet -- install from source:
git clone https://github.com/hahahahahahahahah6/agent-guard.git
cd agent-guard
pip install .
agent-guard install
install merges five hook entries into ~/.claude/settings.json (backing it
up first, never clobbering existing settings):
SessionStart→agent-guard hook-snapshot(records digests, never blocks)Stop→agent-guard hook-test(blocks tests-only changes)PreToolUseonBash→agent-guard hook-outbound(blocks denylisted actions)PreToolUseonWrite|Edit→agent-guard hook-commentslop(blocks comment slop in added text)PostToolUseonBash→agent-guard hook-verify(warns when a claimed effect isn't visible; never blocks)
Restart Claude Code afterwards. No MCP server, no daemon, no accounts.
What the agent sees
Test tampering blocked at stop:
Test-tampering guard: test files changed but no source files changed since
this session started.
Changed test files:
- tests/test_auth.py
This matches a known failure mode where an agent edits test assertions to make
them pass instead of fixing the source code. Before proceeding, verify each
changed test actually fails without the fix: revert the source change, re-run
the test, and confirm it goes red. A test that stays green without the fix is
not covering the bug.
Outbound action blocked before it runs:
Outbound-action guard blocked this command (rule 'git-push-protected': Push to
a protected branch (main/master/prod*/release/*): push to protected ref 'main').
If this action is intended, add an allowlist regex to outbound.allow in
~/.config/agent-guard/config.json, or run it yourself outside the agent.
Note: push blocking is destination-aware. A bare git push (or
git push <remote>) resolves the destination to the current branch via
git symbolic-ref --short HEAD in the hook's working directory
(git -C <dir> is honored), so pushing a feature branch is allowed and
only pushes whose destination is a protected branch are blocked. If the
branch can't be determined (detached HEAD, not a git repo), the push is
allowed rather than breaking your workflow.
Configuration
~/.config/agent-guard/config.json (all optional):
{
"test_guard": {
"mode": "block",
"ignore_paths": ["tests/legacy/"]
},
"outbound": {
"allow": ["my-registry\\.internal"],
"deny_extra": ["rm -rf /tmp/scratch"],
"disabled_rules": ["cloud-provision"]
},
"comment_slop": {
"mode": "block",
"threshold": 30
},
"cheat_sniff": {
"mode": "block",
"threshold": 30,
"allow": []
},
"bash_write": {
"mode": "block",
"protected_paths": ["src/", "infra/"],
"allow": []
}
}
test_guard.mode:"block"(default) or"warn". Env override:AGENT_GUARD_TEST_MODE=warn.comment_slop.mode:"block"(default) or"warn". Env override:AGENT_GUARD_COMMENT_MODE=warn.comment_slop.threshold: slop score (0–100) at which the comment hook trips. Env override:AGENT_GUARD_COMMENT_THRESHOLD.cheat_sniff.mode:"block"(default) or"warn". Env override:AGENT_GUARD_CHEAT_MODE=warn.cheat_sniff.threshold: cheat score (0–100) at which the cheat hook trips (one severe hit reaches the default 30). Env override:AGENT_GUARD_CHEAT_THRESHOLD.cheat_sniff.allow:"path-or-basename:kind"entries that suppress hits, e.g."test_sort.py:rng-seed"or"*/legacy/*:*".bash_write.mode:"block"(default) or"warn". Env override:AGENT_GUARD_BASHWRITE_MODE=warn.bash_write.protected_paths: path prefixes where any Bash write is treated like a Write-tool call (default[], which means the test-tampering guard's scope: test files plusconftest.py).bash_write.allow:"path-or-basename:bash-write"entries that suppress the Bash-write guard, e.g."tests/fixtures/:bash-write".outbound.allow: regexes that win over the denylist (e.g. your internal registry). Every override is audit-logged.outbound.deny_extra/disabled_rules: extend or trim the denylist.
Inspect decisions:
agent-guard log # blocked + allowlist-override + verify-warn decisions
agent-guard status # state paths, mode, active rules
Check whether your tests are honest:
agent-guard mutate-check tests/test_billing.py -- pytest -q
Honest limitations
- Heuristic, not proof. "Tests changed, source didn't" is a strong
signal of the navune cheat, not a proof. Legitimate test-only refactors get
blocked too — that's what
ignore_pathsand warn mode are for. - New tests are allowed. Added test files never count as tampering; only modified or deleted ones do.
- Bash only (for now). The outbound guard watches the
Bashtool. An agent reaching a Slack MCP tool directly is out of scope for this MVP. - Best-effort spend list. Cloud-provision patterns cover the common CLIs; exotic spend paths won't match. The denylist is a seatbelt, not a vault.
- Hook protocol is undocumented. The hook stdin shape and the exit-2-blocks convention come from community documentation, not a stable API. If Claude Code changes the protocol, the hooks degrade to fail-open allow.
- Per-machine only. State lives in
~/.config/agent-guard/. - The hooks fail open: corrupt state, unreadable files, malformed input — anything unexpected means "allow". A guard that wedges your session is worse than no guard.
- Script-content inspection is one level deep. A script that executes
another script (
bash a.shwherea.shrunsbash b.sh) is not followed; exotic interpreter wrappers beyondenv/sudo/nohup/time/niceare not unwrapped. The denylist is a seatbelt, not a vault. - Mutation checking is regex-based, not semantic. It generates at most 20 first-order mutants with simple syntactic operators — good enough to catch vacuous assertions, not a replacement for real mutation-testing tools.
- The post-exec verifier is advisory. It warns on stderr and always exits 0; unreachable remotes, missing npm, and timed-out checks stay silent rather than crying wolf.
- Slop detection is stylistic, not semantic. The six patterns are regex
heuristics tuned for low false positives, which means they miss subtler
slop (a well-written but pointless paragraph scores 0). Short added
comments normalize aggressively — one narrative line in an otherwise
comment-free edit scores high, by design. If your codebase has a
comment-heavy style (or non-English comments the patterns don't cover),
use warn mode or raise
comment_slop.threshold. decomment --fixonly removes commented-out code. Other slop kinds are reported, never auto-edited — deleting prose automatically is how you lose the one comment that mattered.- Cheat-sniffing is static and Python-first. It reads text, not runtime
behavior — a cheat applied only at runtime (e.g. via
sitecustomize.pyor an installed plugin) is invisible to it. The "subject vs collaborator" judgment is a filename heuristic (test_billing.py→billing); exotic layouts need the allowlist. A fixed seed with a reproducibility marker is trusted on the marker's word — an agent that writes "reproducible" next to a planted seed fools the exemption, which is why severe kinds (mock-subject, rng-patch, conftest-patch) have no marker exemption at all. - The cheat hook only watches added test text. Pre-existing cheats in
the repo are found by
cheatsniff --check, not blocked by the hook. - Bash write-target extraction is static and approximate. It is
shlex-based, so heavy quoting,
eval, command substitution building paths at runtime, andpython3 -c "open(...).write(...)"are known gaps — the target list is conservative by design (a missed target is a miss, never a crash). Unresolvable$VARexpansions, bare globs, and/dev/nullare skipped, not guessed at. This is the documented frontier for a future version, not a finished parser. - Opaque Bash writes to protected files are blocked, not scored. When
the hook cannot see what is being written (
sed -i, bare>), there is no content to score — block mode blocks the bypass shape itself. If that is too strict for your workflow, use warn mode or write through the Edit tool instead.
Roadmap
Cross-tool write guard: stop Bash (— shipped in v0.5 assed -i, heredocs, redirections) from bypassing the Edit/Write hooks (thomastartrau: "a guardrail on one tool protects nothing if another tool can do the same thing").hook-bashwrite.Cheat-sniffing beyond test files: catch RNG rigging, subject-mocking, and conftest plants (the remdore study: half of cheating never touches test files).— shipped in v0.4 ashook-cheatsniff+cheatsniff --check.Comment-slop guard: intercept the agent dumping conversation state into code comments (the "9 out of 10 of my revisions is deleting comments" complaint), or a one-command decomment pass before PRs.— shipped in v0.3 ashook-commentslop+decomment --check/--fix.Test-honesty hook: automate the mutation idea — deliberately break an assertion, run once, require red— shipped in v0.2 asmutate-check.
Development
python3 tests/test_agent_guard.py # 47 smoke tests
python3 tests/test_cheatsniff.py # 29 cheat-sniff tests
python3 tests/test_bashwrite.py # bash write-target extraction + hook tests
License
MIT
Metadata
Release files for agent-guard-hooks 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_guard_hooks-0.5.0.tar.gz | 65.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_guard_hooks-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 116.7 kB
Release files / agent_guard_hooks-0.5.0.tar.gz
| Download URL | agent_guard_hooks-0.5.0.tar.gz |
|---|---|
| Size | 65.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d40f23836605341d26133c68baf78520241ed759ef0039651180de7fcf20c52f
|
|
BLAKE2b-256 checksum How to use checksums |
ceb9fe2220f7337474da610f0c7417d829d4f5af888dc8e8bb425fe28b6dcaf2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / agent_guard_hooks-0.5.0-py3-none-any.whl
| Download URL | agent_guard_hooks-0.5.0-py3-none-any.whl |
|---|---|
| Size | 51.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d7a3f43c24c8708cfc35b1144f089e4465764dff9d4132725395c98833fd4d22
|
|
BLAKE2b-256 checksum How to use checksums |
54e7de2569f4509212dc5a2c55a0b5d8ac2328e952eb9f21b4f687da92962d27
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|