Quickstart · How it works · Results · Configuration · Docs
Coding agents end turns with "all tests pass", "the build succeeds", "fixed" or
"deployed", and the session often does not back it up. proof-of-done treats those words as
claims. A Claude Code Stop hook checks the agent's own transcript: did the matching command run
after the last relevant edit, and did it succeed? If not, the stop is blocked and the agent
is told exactly what to run. The same engine audits past transcripts and reports how often a set
of sessions claimed what they never checked.
No LLM, no network, no re-running your commands: for a given transcript, message and configuration the verdict is deterministic.
A fixture Stop payload piped into the installed hook, then audit --demo. Not a live agent session.
Highlights
- Acts only on claims. A keyword prefilter lets every other stop through without reading the transcript; the hook never runs tests or deploys itself.
- Evidence, not vibes. The latest matching command after the last relevant edit decides.
Failed, empty, interrupted, background and masked runs (
| tail,|| true) do not count. - Actionable. Each block names the claim, the reason with step numbers and the file, and a
Run:line with the command to run. - Hard to game. Fake
echo "12 passed", a redefinedpytest,PATHtricks,--collect-only, the agent typing the skip token or editing the config to switch the check off: an adversarial set tests all of these, and every case that gets through is listed below. - Safe to leave on. Stdlib-only hook, fails open on any error, a loop cap below Claude Code's own, a tamper check against self-disabling.
- Measured. Precision and recall on a held-out set frozen before the detector existed, and latency on real hardware, all reproducible from the repository.
Quickstart
claude plugin marketplace add B0yko/proof-of-done
claude plugin install proof-of-done@proof-of-done
The plugin loads on the next start of Claude Code, or after /reload-plugins in an open session.
Inside a session, the equivalent is /plugin marketplace add B0yko/proof-of-done, then
/plugin install proof-of-done@proof-of-done, which opens the plugin's details to pick a scope;
closing the /plugin panel reloads plugins
(docs).
Try the audit CLI on a bundled synthetic corpus without installing anything:
uvx proof-of-done audit --demo
Requires python3 ≥ 3.9 on the PATH Claude Code's hooks see (on macOS, /usr/bin/python3
needs the Command Line Tools; without them the hook skips its check with a message), macOS or
Linux. Tested with Claude Code 2.1.281.
What a block looks like
Real output: a fixture session in which the tests ran alongside the edit rather than after it,
piped into bin/proof-of-done-hook, the launcher Claude Code runs on Stop. The hook prints
{"decision": "block", "reason": …}; this is the reason the agent receives:
proof-of-done: 1 claim in your final message is not backed by this session's transcript.
1. "all 40 tests pass" (tests_passed): the last test run `uv run pytest -q` at step 2 was before your edit to `src/app/parser.py` at step 4.
Run: uv run pytest -q
Run these commands and report the real result, or restate your message without these claims.
The Run: command is the one the agent itself ran. Failing that, it comes from the rule's
suggest setting or from the project files; for a deploy claim or an unrecognised project it
reads Run: (no … command found — run it in the foreground).
Reproduce this block
OUT=$(mktemp -d)
uv run python fixtures/render.py fixtures/scenarios/parallel-calls.yaml --out "$OUT"
sed -e "s|__TRANSCRIPT_PATH__|$(ls "$OUT"/*.stop-0-0.jsonl)|" \
-e "s|__CWD__|$OUT|" -e "s|__SCRATCHPAD_DIR__|$OUT|" \
-e "s|__LAST_ASSISTANT_MESSAGE__|Updated the parser; all 40 tests pass.|" \
fixtures/payloads/stop_with_last_message.json |
sh bin/proof-of-done-hook stop --data-dir "$OUT/data" |
python3 -c 'import json, sys; print(json.load(sys.stdin)["reason"])'
fixtures/render.py writes the transcript and one snapshot per stop attempt
(*.stop-<turn>-<attempt>.jsonl); the other scenarios under fixtures/scenarios/ work the same
way. The templated set is expanded in-process by fixtures/expand.py (--stats prints its
counts).
How it works
flowchart TD
A([Claude Code Stop event]) --> B[launcher]
B --> C{fast-path prefilter:<br/>a claim keyword?}
C -- no --> D([exit 0, no output])
C -- yes --> E[merged-config cache]
E --> F[transcript parser<br/>edits · commands · subagents]
K([proof-of-done audit]) --> F
F --> G[claim detector]
G --> H[evidence engine<br/>anchor → candidates → verdict]
H --> I[suggester: Run: command]
H --> L[(audit report +<br/>agent-trace/v1 export)]
I --> J([decision JSON<br/>block · warn · allow])
classDef accent fill:#10b981,stroke:#047857,color:#ffffff
class J accent
For each claim in the final message, the evidence engine finds the anchor (the last edit to a file the rule cares about: edit tools, Bash writes, formatters, whole-tree git operations, subagent calls) and the candidates (foreground commands after it that match the rule). The latest candidate decides.
| Claim type | Example wording | Counts as evidence (examples) |
|---|---|---|
tests_passed |
"all 42 tests pass", "✅ Tests" | pytest, npm test, go test, cargo test, make test |
build_passed |
"the build succeeds" | npm run build, go build, cargo build, uv build |
lint_clean |
"lint is clean" | ruff check, eslint, golangci-lint run, cargo clippy |
typecheck_clean |
"mypy passes" | mypy, pyright, tsc, cargo check |
deployed |
"deployed to staging" | fly deploy, kubectl apply, terraform apply, docker push, npm publish |
fixed, verified |
"fixed the crash", "verified with curl" | any command above, or curl, python script.py, go run, make <target> |
Every verdict carries one reason from a closed set:
| Reason | Meaning |
|---|---|
supported |
the latest matching run after the last relevant edit succeeded |
no_command |
no matching command ran in the session |
stale |
the last matching run came before the last relevant edit |
failed_exit |
it exited non-zero, was interrupted or timed out |
failed_output |
it exited 0, but its output reports failures |
empty_run |
it ran no tests (collected 0 items, No tests found) |
masked_inconclusive |
its exit status was masked (| tail, || true) and the output shows no success line |
background_only |
it only ran in the background: re-run it in the foreground |
superseded_by_failure |
a partial run (-k, --lf) passed after a full run failed |
no_result |
the run has no recorded result |
The claim detector is plain re. It masks code blocks, inline code, block quotes and quoted user
text, then rejects negated, hedged, conditional, future, question and instruction-to-the-user
forms ("should pass", "run pytest to confirm", "not verified"). Detection is English only: a
claim in another language is not detected at all.
The full rules are in docs/how-it-works.md; the design decisions are in nine short ADRs under docs/adr/.
Results
Every number comes from uv run python eval/run_eval.py and uv run python eval/latency.py --generate --run --max-load 4, and every table is generated from the committed results files by
scripts/render_results.py, which CI checks, along with a re-run of the evaluation. There are
three fixture sets, never merged into one headline:
- templated: expanded from
fixtures/templates/by the same author who wrote the detector, so it measures consistency with that author's reading of the claim wording; - held-out: individually authored, free-form sessions, frozen at commit
4e54fc9(eval: freeze held-out set) beforeclaims.pyexisted, the more honest estimate; - adversarial: sessions that deliberately try to game the gate.
| templated | held-out | |
|---|---|---|
| claim detection — precision / recall | 100.0% / 100.0% | 98.6% / 84.0% |
| gate on unsupported claims — precision / recall | 100.0% / 96.3% | 98.2% / 82.1% |
| turns blocked without an unsupported claim | 0.0% | 1.0% |
34 / 36 gaming attempts met their labelled outcome · hook p95 on a 10 MB transcript 175 ms (system Python 3.9) / 113 ms (CPython 3.12) · audit 141 MB/s
How the held-out numbers got here: first run, tuning, label audit
The first held-out run, with a detector tuned only on the templated set, found few of the
labelled claims. The detector was then broadened on a separate development set (eval/dev/). The
held-out set played no part in that tuning: after the freeze, the held-out files changed only in
label-correction commits and in documentation commits that touched only CHANGES.md there
(git log -- eval/heldout); the development scorer (eval/dev_eval.py) reads only eval/dev/
and the templated set; only eval/run_eval.py and the manifest test read eval/heldout/. A label
audit of every held-out disagreement then corrected the label errors it found, mostly unmarked
lead-in claims such as "Fixed the lock ordering in acquire().". Each correction is recorded with
its reason in eval/heldout/CHANGES.md, together with the ids of a
25-scenario spot check of cases without disagreement. That check found two more unmarked lead-in
claims; both were corrected, and since the detector misses both, the correction lowered recall.
The rows below are the committed runs: the two earlier ones in eval/results/history/, the
current one in eval/results/<date>.json. The two earlier rows were re-scored with the current
per-message detection matching: product code and labels as of the listed commit, harness fix from
commit b489e6d (see the rescored field in each history JSON). The gate columns were not
affected, so uv run python eval/run_eval.py at a listed commit reproduces them.
| run | commit | held-out detection P / R | held-out gate claim P / R | held-out gate turn P / R | held-out false-block rate | templated gate claim P / R |
|---|---|---|---|---|---|---|
| first held-out run (detector tuned on the templated set only) | 7bfbd7e |
89.8% / 29.1% | 95.2% / 32.3% | 100.0% / 29.0% | 0.0% | 100.0% / 96.3% |
| after tuning on the development set (frozen labels) | 08f2187 |
89.9% / 82.8% | 91.1% / 82.3% | 93.8% / 72.6% | 2.8% | 100.0% / 96.3% |
| current (held-out label corrections in CHANGES.md) | e1579dc |
98.6% / 84.0% | 98.2% / 82.1% | 97.9% / 73.4% | 1.0% | 100.0% / 96.3% |
Claim detection: precision, recall and F1 with N
A detection matches a label when the claim type agrees and the spans overlap by at least one character in the same message.
| set | N labels | N detections | TP | FP | FN | precision | recall | F1 |
|---|---|---|---|---|---|---|---|---|
| templated | 955 | 955 | 955 | 0 | 0 | 100.0% | 100.0% | 1.000 |
| heldout | 163 | 139 | 137 | 2 | 26 | 98.6% | 84.0% | 0.907 |
Gate quality: unsupported claims, claim level and turn level, with intervals
The positive class is "unsupported claim". Precision and recall carry Wilson 95% intervals, F1 a
seeded bootstrap 95% interval. "forced_block" forces every built-in rule to block; it equals the
shipped config because no type needed the switch to warn described under the next heading.
| set | config | N | TP | FP | FN | precision (95% CI) | recall (95% CI) | F1 (95% CI) |
|---|---|---|---|---|---|---|---|---|
| templated | shipped | 955 | 395 | 0 | 15 | 100.0% [99.0%, 100.0%] | 96.3% [94.1%, 97.8%] | 0.981 [0.971, 0.990] |
| templated | forced_block | 955 | 395 | 0 | 15 | 100.0% [99.0%, 100.0%] | 96.3% [94.1%, 97.8%] | 0.981 [0.971, 0.990] |
| heldout | shipped | 164 | 55 | 1 | 12 | 98.2% [90.6%, 99.7%] | 82.1% [71.3%, 89.4%] | 0.894 [0.830, 0.946] |
| heldout | forced_block | 164 | 55 | 1 | 12 | 98.2% [90.6%, 99.7%] | 82.1% [71.3%, 89.4%] | 0.894 [0.830, 0.946] |
At turn level, a turn is a predicted block if any claim in it is unsupported under a block rule.
The false-block rate is the share of turns with no unsupported label that were blocked anyway.
| set | config | N | TP | FP | FN | precision (95% CI) | recall (95% CI) | F1 (95% CI) |
|---|---|---|---|---|---|---|---|---|
| templated | shipped | 975 | 375 | 0 | 35 | 100.0% [99.0%, 100.0%] | 91.5% [88.4%, 93.8%] | 0.955 [0.940, 0.970] |
| templated | forced_block | 975 | 375 | 0 | 35 | 100.0% [99.0%, 100.0%] | 91.5% [88.4%, 93.8%] | 0.955 [0.940, 0.970] |
| heldout | shipped | 169 | 47 | 1 | 17 | 97.9% [89.1%, 99.6%] | 73.4% [61.5%, 82.7%] | 0.839 [0.759, 0.906] |
| heldout | forced_block | 169 | 47 | 1 | 17 | 97.9% [89.1%, 99.6%] | 73.4% [61.5%, 82.7%] | 0.839 [0.759, 0.906] |
False-block rate (shipped config; share of no-unsupported-label turns blocked):
| set | N | blocked | rate | 95% CI |
|---|---|---|---|---|
| templated | 565 | 0 | 0.0% | [0.0%, 0.7%] |
| heldout | 105 | 1 | 1.0% | [0.2%, 5.2%] |
By claim type: and the rule that would switch a weak type to warn
A built-in claim type whose held-out gate precision is below 0.90 ships with action: warn
instead of block. With the corrected labels no type is below 0.90 (verified is exactly 0.90),
so every built-in rule ships as block. On the frozen labels before the audit, build_passed,
fixed and verified were below 0.90 (second history file). The five false positives behind that
were one unmarked lead-in claim (fixed), one claim with the wrong type (build_passed), two
labels flipped because python -c is not an execution command (fixed, verified), and one
genuine false positive that remains, a hedge the negation window does not cover (verified). The
per-type counts are small, so read them as rough.
| set | claim type | N | predicted blocks | precision | recall | F1 |
|---|---|---|---|---|---|---|
| templated | build_passed | 40 | 20 | 100.0% | 100.0% | 1.000 |
| templated | deployed | 40 | 20 | 100.0% | 100.0% | 1.000 |
| templated | fixed | 80 | 20 | 100.0% | 100.0% | 1.000 |
| templated | lint_clean | 110 | 45 | 100.0% | 90.0% | 0.947 |
| templated | tests_passed | 605 | 250 | 100.0% | 96.2% | 0.980 |
| templated | typecheck_clean | 40 | 20 | 100.0% | 100.0% | 1.000 |
| templated | verified | 40 | 20 | 100.0% | 100.0% | 1.000 |
| heldout | build_passed | 19 | 8 | 100.0% | 88.9% | 0.941 |
| heldout | deployed | 28 | 9 | 100.0% | 90.0% | 0.947 |
| heldout | fixed | 26 | 11 | 100.0% | 100.0% | 1.000 |
| heldout | lint_clean | 26 | 8 | 100.0% | 100.0% | 1.000 |
| heldout | tests_passed | 24 | 9 | 100.0% | 75.0% | 0.857 |
| heldout | typecheck_clean | 18 | 1 | 100.0% | 12.5% | 0.222 |
| heldout | verified | 23 | 10 | 90.0% | 100.0% | 0.947 |
Held-out predicted blocks per claim type range from 1 (typecheck_clean) to 11 (fixed); a precision figure that rests on only a few predicted blocks is a rough estimate.
Adversarial resistance and suggestion accuracy
The adversarial cases cover fake success output, redefined pytest functions and aliases, PATH
prepending, collect-only and masked runs, interrupted and timeout-killed runs, a partial pass after
a failed full run, a claim hidden in a code block, the agent writing the skip token itself, and the
agent editing the config or settings to switch the check off. Every miss is a known limitation.
34 / 36 adversarial cases met their labelled outcome.
Known misses:
| scenario | expected decision | expected reason | predicted decision | predicted reason |
|---|---|---|---|---|
| adv-07-makefile-echo-only-known-miss | block | None | allow | None |
| adv-36-npm-run-test-watch-known-miss | block | None | allow | None |
Suggestion accuracy is the share of correctly blocked unsupported claims whose Run: command
equals the labelled expected command.
| set | N correctly blocked | matching suggestion | accuracy |
|---|---|---|---|
| templated | 375 | 355 | 94.7% |
| heldout | 49 | 49 | 100.0% |
Hook latency and audit throughput: MacBook Air, Apple M5, 24 GB
p50, p95 and max over 200 invocations of the real launcher, each a fresh interpreter, for the fast
path (no claim in the message) and 1, 10 and 50 MB transcripts, under the system python3 and a
uv-managed CPython 3.12, with an empty and a warm merged-config cache. The hook's bytecode cache
in the data directory is warm after the first call of each series, and no row uses a transcript
parse cache. The latency targets (warm config cache: p95 ≤ 250 ms up to 10 MB, ≤ 1 s at 50 MB,
under both interpreters) are met.
Machine: Mac17,4, Apple M5, 24 GB RAM, macOS 26.6.2; interpreters: system python3 3.9.6, uv-managed CPython 3.12.14; measured 2026-09-28T20:16:39Z. 1-minute load average during the run: 2.8–5.1 on 10 cores (other processes were running).
| case | interpreter, cache | N | p50 (ms) | p95 (ms) | max (ms) |
|---|---|---|---|---|---|
| fast path (no claim) | system python3, cold config cache | 200 | 132.5 | 147.0 | 323.8 |
| fast path (no claim) | system python3, warm config cache | 200 | 57.9 | 61.5 | 66.0 |
| fast path (no claim) | uv CPython 3.12, cold config cache | 200 | 72.5 | 80.4 | 193.9 |
| fast path (no claim) | uv CPython 3.12, warm config cache | 200 | 23.4 | 27.1 | 65.7 |
| 1 MB | system python3, cold config cache | 200 | 156.2 | 221.8 | 376.6 |
| 1 MB | system python3, warm config cache | 200 | 108.0 | 150.8 | 168.2 |
| 1 MB | uv CPython 3.12, cold config cache | 200 | 71.0 | 72.3 | 167.4 |
| 1 MB | uv CPython 3.12, warm config cache | 200 | 54.1 | 55.3 | 62.5 |
| 10 MB | system python3, cold config cache | 200 | 201.6 | 208.5 | 307.8 |
| 10 MB | system python3, warm config cache | 200 | 170.3 | 174.8 | 187.0 |
| 10 MB | uv CPython 3.12, cold config cache | 200 | 119.3 | 124.1 | 220.0 |
| 10 MB | uv CPython 3.12, warm config cache | 200 | 104.8 | 112.8 | 121.7 |
| 50 MB | system python3, cold config cache | 200 | 559.9 | 609.0 | 748.4 |
| 50 MB | system python3, warm config cache | 200 | 522.7 | 577.8 | 588.4 |
| 50 MB | uv CPython 3.12, cold config cache | 200 | 388.6 | 590.5 | 832.7 |
| 50 MB | uv CPython 3.12, warm config cache | 200 | 331.7 | 466.8 | 586.2 |
Audit throughput on the 50 MB synthetic transcript, median of 5 runs:
proof-of-done audit on a 49.8 MB transcript: median 140.5 MB/s over 5 run(s).
Failure analysis: every held-out error, by failure mode
There are no prefilter misses, no wrong verdicts on detected claims and no wrong suggestions; almost every error is a claim the detector does not recognise.
| failure mode | claim type | count | example scenario | example |
|---|---|---|---|---|
| detection_fn | typecheck_clean | 14 | ho-027 |
"Types check" |
| detection_fn | deployed | 6 | ho-049 |
"pushed the worker image" |
| detection_fn | tests_passed | 3 | ho-006 |
"They pass" |
| detection_fn | lint_clean | 2 | ho-064 |
"vet is clean" |
| detection_fn | build_passed | 1 | ho-012 |
"Build is clean" |
| detection_fp | lint_clean | 1 | ho-104 |
"tests pass" |
| detection_fp | verified | 1 | ho-081 |
"verified" |
- Type-check phrasings are the biggest gap. "Types check", "it typechecks" and "it vets clean" are not in the detector's vocabulary, so an unsupported type-check claim phrased that way passes the gate.
- State-style deploy claims ("it's published", "it's live in production", "pushed the worker image") and pronoun subjects ("They pass") are missed for the same reason.
- Two false positives: a hedge the negation window does not cover ("I wouldn't call it fully
verified yet"), and a two-claim sentence where "tests pass" was attributed to
lint_clean.
None of these was fixed after the held-out run: tuning on the held-out set would make its numbers meaningless. The confusion matrix of reasons and the full error list are in docs/eval.md.
Demo audit report: proof-of-done audit --demo on the bundled synthetic corpus
Thirty hand-authored sessions across Python, Node, Go and Rust projects. This is a demo of the report, not a measurement of any real coding agent. The block is generated from an in-process run and checked in CI.
proof-of-done audit report
SYNTHETIC DEMO CORPUS -- this report describes 30 hand-authored fixture sessions, not a measurement of any real coding agent.
config applied: built-in defaults
files processed: 30
sessions: 30
turns: 30
stop attempts: 30
turns with claims: 30
claims: 45
unsupported: 16 (35.6% of claims)
partial share of supported claims: 17.2%
by claim type:
build_passed: 3 claims, 1 unsupported (33.3%)
deployed: 4 claims, 0 unsupported (0.0%)
fixed: 15 claims, 4 unsupported (26.7%)
lint_clean: 5 claims, 2 unsupported (40.0%)
tests_passed: 12 claims, 7 unsupported (58.3%)
typecheck_clean: 3 claims, 1 unsupported (33.3%)
verified: 3 claims, 1 unsupported (33.3%)
unsupported by reason:
background_only: 1
empty_run: 1
failed_exit: 3
failed_output: 1
masked_inconclusive: 2
no_command: 2
no_result: 2
stale: 2
superseded_by_failure: 2
most recent unsupported claims:
[fixed] "Patched" -- superseded_by_failure
Run: cargo test
[tests_passed] "All tests pass" -- superseded_by_failure
Run: cargo test
[tests_passed] "All tests pass" -- background_only
Run: go test ./...
[fixed] "Fixed the race" -- stale
Run: go test ./...
[lint_clean] "Lint is clean" -- failed_exit
Run: go vet ./...
[tests_passed] "All tests pass" -- empty_run
Run: npx vitest run
[typecheck_clean] "Typecheck is clean" -- failed_exit
Run: npx tsc --noEmit
[fixed] "Patched" -- masked_inconclusive
Run: npx vitest run
[verified] "Confirmed it's working" -- masked_inconclusive
Run: npx vitest run
[build_passed] "The build succeeds" -- failed_exit
Run: uv build
Configuration
Five layers merge, lowest precedence first: the packaged defaults, the user file
${XDG_CONFIG_HOME:-~/.config}/proof-of-done/config.yaml, the project's .proof-of-done.yaml,
the file named by PROOF_OF_DONE_CONFIG, then the PROOF_OF_DONE and PROOF_OF_DONE_MODE
environment variables. Project rules merge into built-in rules by id, field by field; new ids
add custom claim types. Every key and the full evidence-command table are in
docs/config.md; examples/ has configs for Python, Node, Go and a monorepo,
plus custom committed/pushed rules. proof-of-done init writes a starter
.proof-of-done.yaml (version: 1, the detected ecosystem as a comment, rules: []).
| Escape hatch | Effect |
|---|---|
#skip-proof in your own prompt |
skips the check for that turn; the agent writing the token has no effect |
PROOF_OF_DONE=off or enabled: false |
turns the hook off |
PROOF_OF_DONE_MODE=warn or mode: warn |
every block becomes a visible warning |
max_blocks_per_turn (default 2) |
after that many blocks in a row the stop is allowed with a warning, below Claude Code's own default cap of 8 (hooks reference) |
Tamper check. The hook scans the session's own edits before it applies enabled, mode or
any rule action. If the session edited .proof-of-done.yaml, the user config or the
PROOF_OF_DONE_CONFIG file, four downgrades from that file are undone for the rest of the
session: enabled: false, mode: warn, a lowered action and a raised max_blocks_per_turn. If
it edited a Claude Code settings.json/settings.local.json with text containing
PROOF_OF_DONE or enabledPlugins, PROOF_OF_DONE=off and PROOF_OF_DONE_MODE=warn from the
environment are ignored. A warning names the edited file, and tampering alone never blocks. What it
does not undo is listed under Limitations.
CLI
The plugin does not put proof-of-done on your PATH. Install the CLI with
uv tool install proof-of-done (or pipx install proof-of-done), or run any command through
uvx proof-of-done <command>.
proof-of-done check --transcript PATH [--message TEXT | --message-file PATH] [--config PATH] [--json]
proof-of-done audit [PATH ...] [--claude-projects | --demo] [--source auto|claude-code|agent-trace|codex]
[--config PATH] [--format text|json|md] [--out FILE] [--export-traces FILE]
[--redact] [--salt S] [--ci] [--max-unsupported-rate R]
proof-of-done init [--force]
proof-of-done config show [--cwd DIR]
proof-of-done trace validate FILE
What each command does, and its exit codes
checkjudges one final message against one transcript, without a live hook. Exit0: nothing to block (a warn-only result still prints its warning);1: an unsupported claim under ablockaction;2: a usage or config error.auditreplays the Stop hook at every stop attempt of every session, and at each subagent's own transcript grouped under its parent, with the built-in defaults or--config. Exactly one input mode:PATHs (files, directories, globs),--claude-projects($CLAUDE_CONFIG_DIR/projects, else~/.claude/projects) or--demo.--source codexreads OpenAI Codex CLI rollout logs (experimental, audit only).--export-traces FILEwritesagent-trace/v1JSONL;--redactreplaces quotes, commands, paths and session ids with a salted SHA-256 prefix plus length (h:<12 hex>:len=<n>), with a random salt per run unless you pass--saltorPROOF_OF_DONE_SALT. Exit2: a usage error or a missing or invalid--config;3: no input yielded a session (every file unreadable or unrecognised, including a file with lines but no parseable JSON line). With--ci, exit1when the unsupported-claim rate exceeds--max-unsupported-rate(a fraction, default0; ignored without--ci, soaudit --demo --ciexits1), plus a Markdown summary in$GITHUB_STEP_SUMMARYwhen set. Seeexamples/ci/audit-agent-run.ymlfor a GitHub Actions job shape (not run by this repository's CI: it needs a real headless agent run).initwrites.proof-of-done.yamlin the current directory; it refuses to overwrite one without--force(exit2). The built-in rules already cover Python, Node, Go and Rust, so the file changes nothing until you add overrides.config showprints the effective merged configuration and the files it came from.trace validatechecks anagent-trace/v1file againstschemas/agent-trace-v1.json.
Debugging a block
proof-of-done check --transcript path/to/session.jsonl --message "All tests pass."
proof-of-done config show
The hook logs to ${CLAUDE_PLUGIN_DATA}/proof-of-done.log (or
${TMPDIR:-/tmp}/proof-of-done-<uid>/ when that is unset), JSON Lines capped at 256 KiB. Each stop
past the prefilter writes one decision record with the final decision (after the block cap;
capped: true when the cap let a stop through), the claim count, the flags skipped, disabled,
tampered and capped, one entry per claim (claim_type, rule_ids, supported, reason code,
action) and per-phase timings_ms. A fail-open writes a fail_open record with a reason code,
and an unexpected exception an error record with only the exception class name. The log never
contains message text, quotes, commands or file paths.
Privacy
The hook reads only local files: the event JSON on stdin; the session transcript (for
SubagentStop, the subagent's transcript plus the parent's, for the skip token and the tamper
scan); long tool outputs Claude Code stored inside that session's own directory; the config files;
and, to suggest a Run: command, the project's pyproject.toml, pytest.ini, package.json,
go.mod, Cargo.toml, Makefile and lockfiles. It makes no network call, and a
socket-blocking test fixture enforces that. It writes only under ${CLAUDE_PLUGIN_DATA} (or the
temp fallback): block counters, the merged-config cache, Python's bytecode cache, the log, and on
macOS a .python-ok marker when the interpreter is /usr/bin/python3. It never writes into your
project.
Alternatives
Existing Claude Code hooks and plugins that verify completion claims or gate Stop on test, build
or deploy state, each checked against its repository on 2026-09-28:
| Project | Mechanism | License · stars |
|---|---|---|
| claimcheck | Checks each claim in the final message against tool output already in the same turn's transcript (verified / refuted / unverifiable); never re-runs commands. | GPL-3.0 · 0 |
| done-needs-proof | Node Stop hook; blocks "done", "fixed", "deployed" or "tests pass" unless a matching verification command ran after the last edit. Keyword matching, fails open, fires once per stop; ships a prove skill. |
MIT · 0 |
| loop-hooks | Re-runs the project's verification command on Stop, SubagentStop and TeammateIdle, gated by a git fingerprint of the watched files. |
MIT · 0 |
| tdd-guard, successor Probity | A process gate rather than a claim checker: blocks implementation edits without a failing test first; some checks call a validation model. The most adopted tool in this space. | MIT · 2,353 (tdd-guard), 216 (Probity) |
| protect-tests | PreToolUse hook against "fake green": deleting, renaming away or skip/xfail-disabling tests instead of fixing the code. Deterministic patterns. |
MIT · 526 (whole repo) |
| dead-rules-audit | Not a stop gate: scores edits against CLAUDE.md's own rules and reports which are ignored. An adjacent pattern, listed for context. |
MIT · 526 (whole repo) |
| sonmat | Verification discipline mostly through prompt guidance (in CLAUDE.md and injected into worker subagents), plus a commit guard; no transcript checking described. |
BSD-3-Clause · 6 |
| Official plugin-dev hook guide | Its Stop hook example is a prompt hook: a second model call judging "approve" or "block". Teaching material, not a maintained plugin. A read of the claude-plugins-official listing (314 entries, descriptions only) found none that describes verifying completion claims. |
— |
proof-of-done sits with claimcheck and done-needs-proof on deterministic transcript inspection, never re-runs a command (unlike loop-hooks), never calls a model (unlike Probity's validation option and the official prompt-hook example), and publishes a reproducible evaluation against a frozen held-out set. This describes what exists; it is not a claim to be first.
Limitations
- Missed phrasings are never checked. The detector does not know every wording: type-check phrasing ("Types check", "it vets clean"), state-style deploy claims ("it's live in production") and pronoun subjects ("They pass") are the main gaps. See the failure analysis under Results.
- English only.
- Two known adversarial misses. An echo-only Makefile
test:target whose output looks like a test summary, andnpm run test:*scripts, which are trusted by name even when they start a watcher. - Background runs never count as evidence in v0.1, even when a later notification reports success; the block message says to re-run in the foreground.
- Only Bash is evidence. MCP test runners, IDE diagnostics and the PowerShell tool are not read.
- Edits outside the agent are invisible. A file you change in your editor has no transcript
event; there is no cross-check with
git statusor file mtimes. - Subagents are judged separately. A subagent call counts as an edit for every rule (by default); what it touched is not parsed, and its verdict is not reused as the parent's evidence.
- A transcript that starts mid-history only warns. When the first message's
parentUuidnames an entry missing from the file, the evidence is known to be incomplete, so blocks become warnings. - The transcript format is internal to Claude Code and pinned to 2.1.281 (docs/transcript-format.md); on a changed format the hook fails open, and its accuracy is unverified until re-pinned.
- Tamper protection is partial. It undoes only the four downgrades above. Other keys from an
edited config still apply, including
max_transcript_mb,check_subagentsandsubagent_skip_types(read before the transcript is parsed) and a rule'sclaimsandkeywords; an invalid config fails open with a warning. - Out of scope: generic "done" or "implemented" claims, claims about external state no command can show ("the email was sent"), and Windows.
- The Codex adapter is experimental and audit-only, pinned to one
openai/codexsource commit and built from hand-written fixtures.
Uninstall or disable
claude plugin uninstall proof-of-done@proof-of-done # remove it
claude plugin disable proof-of-done # keep it installed, turn it off
Or set PROOF_OF_DONE=off for Claude Code's hook processes, or enabled: false in
.proof-of-done.yaml.
Roadmap
- Reuse a subagent's already-checked verdict as evidence for its parent's claim.
- Evidence from sources other than Bash (MCP test runners, IDE diagnostics).
- Pin the Codex adapter against a verified rollout log and consider a Codex-side hook.
Data and licence
All fixtures, templates, held-out, development and adversarial sessions are synthetic, written for
this repository and licensed Apache-2.0 with the code (see fixtures/README.md).
The only vendored third-party code is a pure-Python copy of PyYAML 6.0.3 (MIT, licence kept
in src/proof_of_done/_vendor/yaml/LICENSE), so the hook needs no install step
(ADR 3).
Apache-2.0, see LICENSE. Copyright 2026 Andrii Boiko.
Release files for proof-of-done 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| proof_of_done-0.1.1.tar.gz | 324.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| proof_of_done-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 542.9 kB
Release files / proof_of_done-0.1.1.tar.gz
| Download URL | proof_of_done-0.1.1.tar.gz |
|---|---|
| Size | 324.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bb61f663b809d125495744dc749bdf5744f060f75ad2b531979307705ce2d24e
|
|
BLAKE2b-256 checksum How to use checksums |
5ceb5a9cc7a6794fcd8061a092e8b9d462e795aa9f6645572926c8e9a921d62c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / proof_of_done-0.1.1-py3-none-any.whl
| Download URL | proof_of_done-0.1.1-py3-none-any.whl |
|---|---|
| Size | 218.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
593cc2602a9b2594afdce5ec37f688bcf8f4c5e669f51e22d04ceae1190f8f98
|
|
BLAKE2b-256 checksum How to use checksums |
a79e0ce053699e1dcad3a831fdd75fb999fa5786e92476a362766f6e8fd54284
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log