assurance-budget
What did a coding-agent session actually do — and which limits did a run log never test?
Install
pip install assurance-budget
# or: pip install assurance # every tool
Quick start
Session audit (Claude Code transcript — the command most people want):
$ assurance audit examples/audit/sample-session.jsonl
Claude Code session demo-8f2 — 13 min in /home/you/my-app
10 tool calls, 3 failed — Bash 6, Edit 2, Grep 1, Read 1
Looped: 3 rounds of Bash `pytest -q tests/test_invoice.py` failing the same way, with nothing new read
After the last edit (14:09): no test or check it recognises; 2 unclassified commands ran after it (make lint-fix, python script)
Not classified: 2 shell commands (make lint-fix, python script), so whether they read, wrote or tested anything is unknown.
Also in the transcript: 1 assistant turn, 1 user turn, 1 bookkeeping record.
Not read: 0 lines.
assurance audit --demo prints this from any folder, and assurance audit --session <id> audits one
session by its id, opening that session's transcript and no other (the Claude Code plugin's
/assurance:audit passes its own). As a Claude Code Stop hook, it runs after
every turn and speaks when something is at stake: untested code pushed, merged, published, deployed or
committed on main (check before proceeding), or a failed test or check after the last code edit, or
Claude saying the tests pass with nothing behind it (review suggested). --nudge also asks Claude to
act (once per turn, and it never fails the session).
assurance hook install adds it after showing you the change; assurance hook remove takes it out:
{ "hooks": { "Stop": [ { "hooks": [ { "type": "command", "command": "uvx --offline assurance@0.1.14 audit --hook --nudge" } ] } ] } }
By hand, run uvx assurance@0.1.14 --version once first: --offline runs the copy uv already has, so
the hook never waits on PyPI.
Run-log budget (JSONL with a per-run id):
$ assurance-budget runs.jsonl
1 of 3 runs hit a limit — 1 was going nowhere first
Limits from: built-in defaults
r-002
Stopped: 3 rounds repeating fetch(url=api/invoices) and failing the same way (timeout) with nothing new read and no part of the goal closer. Continuing would spend the rest of this run's budget on the same result.
r-003
Stopped after 20 frontier calls
Not tested by this log: iterations, retries. The log carries no events of that kind, so this is silence rather than a pass.
As a library (stops a live loop, rather than reporting after the fact):
from assurance_core.run_budget import Budget, ProgressWatch, Progress, Spend
spend = Spend(budget=Budget.allowing(tool_calls=20))
watch = ProgressWatch()
for step in range(100):
stopped = spend.charge_tool_call()
if stopped:
print(stopped.message)
break
stalled = watch.observe(Progress(action="fetch(url=api/invoices)", error="timeout"))
if stalled:
print(stalled.message)
break
assert spend.tool_calls <= 20
What it checks
- Tool calls, failures, and loops in a Claude Code session (
assurance audit) - Whether a test or check ran after the last in-project edit, and whether the last one passed
- Shell commands it could not classify, named by kind (
python -c ×3, curl);Not read:lines name why - Edits with no visible read, in
--json(Claude Code itself refuses those, so the text stays quiet) - Which runs in a JSONL log hit a ceiling or stalled with nothing new read
- Which configured limits the log never exercised (silence, not a pass)
- What the session touched next to what it had: MCP servers used and loaded but never used, skills listed and used, agents, hooks with their runs and failures, and commands typed
The inventory
assurance audit --json carries it as inventory, shape assurance.inventory/1. Rooms draws it as a
page. It is read from the transcript alone, and it counts; whether a server or skill the session
carried and never used helped or got in the way is not something a count can say.
| key | what |
|---|---|
tools |
built-in tools called, with how often |
mcp_servers |
each server: name, title (a readable name for one known only by an id), state (used, not used, failed, needs sign-in, pending), calls, tools_used, tools_available |
skills |
listed to the model, and used, with how often |
agents |
listed types, and used, with how often |
hooks |
each hook: name, command, runs, failed, median_ms |
commands |
slash commands typed, with how often |
not_recorded |
what the transcript does not hold: tools loaded into the prompt from the start are not in it, so an unused one among them cannot be counted |
In CI
| exit | means |
|---|---|
0 |
audited |
1 |
--fail-on-exhausted / --fail-on-loop / --fail-on-unverified found a problem |
2 |
the transcript or log could not be read |
Limits
- Operator ceilings. Built-in defaults in
assurance-core; raise via~/.config/assurance/config.tomlorASSURANCE_MAX_*. Project file<cwd>/.assurance/config.tomlcan only lower. A caller flag can tighten, never raise past the operator. - It reads what the log records. Unlogged spend is invisible.
- No dollar figure. Frontier calls are the cost proxy.
- Stall detection needs three identical rounds with flat progress.
- A chat transcript with no run id is refused (exit 2), not treated as one run.
See the root README and CHANGELOG.md.
Metadata
Release files for assurance-budget 0.2.10
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| assurance_budget-0.2.10.tar.gz | 127.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| assurance_budget-0.2.10-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 207.7 kB
Release files / assurance_budget-0.2.10.tar.gz
| Download URL | assurance_budget-0.2.10.tar.gz |
|---|---|
| Size | 127.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bb0cc27503a5cfaeee54304f9132ceac7cb42c4954e163aa121d8edd64573478
|
|
BLAKE2b-256 checksum How to use checksums |
1dcd9b7fc2f2534f72ce309c2731d16cbbfa8f5d4fe718fcd291839689b636b5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency logRelease files / assurance_budget-0.2.10-py3-none-any.whl
| Download URL | assurance_budget-0.2.10-py3-none-any.whl |
|---|---|
| Size | 80.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
01d2ffcadbe3d99994a567903783dd6acb0d711998f3242bd8242fd4548f698b
|
|
BLAKE2b-256 checksum How to use checksums |
e8568ae26ce1d57bdec95e32a5288725cb8be7cf5100e731adc3b43b36c62351
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency log