Skip to main content

assurance-budget

tests PyPI License

What did a coding-agent session actually do — and which limits did a run log never test?

Install

pip install assurance-budget
# or: pip install assurance   # every tool

Quick start

Session audit (Claude Code transcript — the command most people want):

$ assurance audit examples/audit/sample-session.jsonl
Claude Code session demo-8f2 — 13 min in /home/you/my-app
10 tool calls, 3 failed — Bash 6, Edit 2, Grep 1, Read 1

  Looped: 3 rounds of Bash `pytest -q tests/test_invoice.py` failing the same way, with nothing new read
  After the last edit (14:09): no test or check it recognises; 2 unclassified commands ran after it (make lint-fix, python script)
  Against the last prompt (14:00, "The invoice totals are off by a cent for EUR. Fix it in inv…"):
    invoice.py: changed at 14:04 (src/billing/invoice.py).
    pytest -q tests/test_invoice.py: did not run after the last edit to src/billing/rates.py (14:09); it last failed at 14:08, before that.
    Only what the last prompt names is checked here; whether the work does what it asks is not.
  Not classified: 2 shell commands (make lint-fix, python script), so whether they read, wrote or tested anything is unknown.
  Also in the transcript: 1 assistant turn, 1 user turn, 1 bookkeeping record.
  Not read: 0 lines.

assurance audit --demo prints this from any folder, and assurance audit --session <id> audits one session by its id, opening that session's transcript and no other (the Claude Code plugin's /assurance:audit passes its own). As a Claude Code Stop hook, it runs after every turn and speaks when something is at stake: untested code pushed, merged, published, deployed or committed on main, or shipped while a command the project says must pass or a test the last prompt names had not passed after the edit (check before proceeding); a failed test or check after the last code edit, a change to a path the project protects, or Claude saying the tests pass with nothing behind it, or right after a test that failed (review suggested). It says a finding once, and not again while the edit, the command and the time it rests on are the same. --nudge also asks Claude to act (once per turn, and it never fails the session). assurance hook install adds it after showing you the change; assurance hook remove takes it out:

{ "hooks": { "Stop": [ { "hooks": [ { "type": "command", "command": "uvx --offline assurance@0.1.25 audit --hook --nudge" } ] } ] } }

By hand, run uvx assurance@0.1.25 --version once first: --offline runs the copy uv already has, so the hook never waits on PyPI.

Run-log budget (JSONL with a per-run id):

$ assurance-budget runs.jsonl
1 of 3 runs hit a limit — 1 was going nowhere first

  Limits from: built-in defaults

  r-002
      Stopped: 3 rounds repeating fetch(url=api/invoices) and failing the same way (timeout) with nothing new read and no part of the goal closer. Continuing would spend the rest of this run's budget on the same result.
  r-003
      Stopped after 20 frontier calls

  Not tested by this log: iterations, retries. The log carries no events of that kind, so this is silence rather than a pass.

As a library (stops a live loop, rather than reporting after the fact):

from assurance_core.run_budget import Budget, ProgressWatch, Progress, Spend

spend = Spend(budget=Budget.allowing(tool_calls=20))
watch = ProgressWatch()

for step in range(100):
    stopped = spend.charge_tool_call()
    if stopped:
        print(stopped.message)
        break
    stalled = watch.observe(Progress(action="fetch(url=api/invoices)", error="timeout"))
    if stalled:
        print(stalled.message)
        break

assert spend.tool_calls <= 20

Or let the recorder count, and keep the record as it goes. It stops a run only at a limit someone set: the code's limits=, or the operator's settings below, which the code can lower and never raise. watch(client) does the same for an Anthropic or OpenAI client, on a copy it hands back.

import json, tempfile
from pathlib import Path

from assurance_budget.record import Recorder, RunStopped

path = Path(tempfile.mkdtemp()) / "run.jsonl"
with Recorder(path, limits={"tool_calls": 3}) as rec:
    rec.task("Count the open invoices")
    try:
        for page in range(10):
            with rec.tool("fetch", input={"page": page}) as call:
                call.output = f"page {page}: 25 open"
    except RunStopped as stop:
        print(stop.reason)  # Stopped after 3 tool calls — this run's limit is 3. ...
    rec.claim("75 open invoices")

lines = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines()]
assert [line["type"] for line in lines].count("tool") == 3

What it checks

  • Tool calls, failures, and loops in a Claude Code session (assurance audit)
  • Whether a test or check ran after the last in-project edit, and whether the last one passed, each read from where the command ran: work in another repository, or one nested in the folder, is not the project's
  • Shell commands it could not classify, named by kind (python -c ×3, curl); Not read: lines name why
  • Edits with no recorded read, in --json as edits_with_no_recorded_read, with what that rests on: Claude Code refuses an edit to a file the session has not read or written, so there each is a read or a write this reader did not see, and edited_without_read stays empty and the text quiet
  • Which runs in a JSONL log hit a ceiling or stalled with nothing new read
  • Which configured limits the log never exercised (silence, not a pass)
  • What the session touched next to what it had: MCP servers used and loaded but never used, skills listed and used, agents, hooks with their runs and failures, and commands typed
  • The outcome against what was asked: what happened to the files, tests and commands the last prompt names, and to the project's must_run commands and must_not_touch paths, each with what it could not check
  • The same for an agent you wrote, or code that calls a model, from a run record its code writes (assurance.run/1, in the root README): its task, model calls, tools, edits, commands, the checks it made, its last word, and each gate's decision held against what the step then did
  • Recorder writes that record from the agent's own code, records each Anthropic or OpenAI SDK call, and stops the run at the limits set for it (assurance_budget.record, below)
  • The same for any agent that sends OpenTelemetry traces, from OTLP JSON or what the Python SDK's console exporter prints, read by the GenAI, OpenInference and OpenLLMetry conventions; FileExporter writes one from the tracer an agent already has (assurance_budget.otel, in the root README)
  • A run's last word held against what failed: --fail-on-claim exits 1 when a run says it is done and a check did not hold, or a step's last run failed
  • assurance serve: a local endpoint any agent sends OTLP traces (protobuf or JSON) or run record lines to, and any workflow asks for a run's audit and verdict (assurance_budget.serve, in the root README)

The outcome

assurance audit --json carries it as outcome, shape assurance.outcome/1. Each check is a typed decision: a question, an answer from a fixed set, the evidence, and why when the answer is unknown. Whether the work does what the prompt asks is not one of them; not_checked says so.

key what
prompt the person's last prompt: at, an excerpt, and how many images came with it; null when there is none
checks each check: from (prompt, must_run, must_not_touch), kind, subject, question, answer, evidence, unknown_because; a prompt's command says whether it was asked for, and a rule says where it was declared_in
not_checked what it could not look at: a prompt that names nothing, an image, a file outside the project, and whether the work does what was asked

A run record adds run (its id, the other runs in the file, the task's rules and expected outputs, and its last word with what goes against it), model_calls, decisions (shape assurance.decisions/1: each gate's verdict, and held, failed, not checked, did not run, ran anyway or ran before it), and not_recorded, what the record left out that a check needed.

kind answers
file changed, read, not opened, unknown
command passed, failed, unknown, not run, no code edited
test passed, failed, unknown, not run
paths untouched, changed, unknown

The inventory

assurance audit --json carries it as inventory, shape assurance.inventory/1. Rooms draws it as a page. It is read from the transcript alone, and it counts; whether a server or skill the session carried and never used helped or got in the way is not something a count can say.

key what
tools built-in tools called, with how often
mcp_servers each server: name, title (a readable name for one known only by an id), state (used, not used, failed, needs sign-in, pending), calls, tools_used, tools_available
skills listed to the model, and used, with how often
agents listed types, and used, with how often
hooks each hook: name, command, runs, failed, median_ms
commands slash commands typed, with how often
not_recorded what the transcript does not hold: tools loaded into the prompt from the start are not in it, so an unused one among them cannot be counted

In CI

exit means
0 audited
1 --fail-on-exhausted / --fail-on-loop / --fail-on-unverified / --fail-on-outcome / --fail-on-claim found a problem
2 the transcript or log could not be read

Limits

  • Operator ceilings. Built-in defaults in assurance-core; raise via ~/.config/assurance/config.toml or ASSURANCE_MAX_*. Project file <cwd>/.assurance/config.toml can only lower. A caller flag can tighten, never raise past the operator.
  • The recorder stops a run only at a limit someone set, in its limits= or the operator's settings; with neither, it records and never stops. It counts tool calls, model calls and seconds, and a limit of 20 lets 20 run. Settings it cannot read are not guessed at: the built-in defaults apply, and it warns.
  • It reads what the log records. Unlogged spend is invisible.
  • No dollar figure. Frontier calls are the cost proxy.
  • Stall detection needs three identical rounds with flat progress.
  • A chat transcript with no run id is refused (exit 2), not treated as one run.

See the root README and CHANGELOG.md.

Metadata

Release files for assurance-budget 0.2.19

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for assurance-budget 0.2.19
File Size Uploaded
assurance_budget-0.2.19.tar.gz 249.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for assurance-budget 0.2.19
File Interpreter ABI Platform
assurance_budget-0.2.19-py3-none-any.whl Python 3 none any Details

Total release size: 405.7 kB

Release files / assurance_budget-0.2.19.tar.gz

Download URL assurance_budget-0.2.19.tar.gz
Size 249.4 kB
Tags Source
SHA-256 checksum
How to use checksums
ad4cbe185451319fc6d1afb5bfa0fb269df556e8b138f0fe50eadd476dcb375c
BLAKE2b-256 checksum
How to use checksums
6b148730ee32f35e3403efaa66d454a7918ba075e460d49cf640298c9d53f0f1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / assurance_budget-0.2.19-py3-none-any.whl

Download URL assurance_budget-0.2.19-py3-none-any.whl
Size 156.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b2b3f0104a8372b1e1e35839ab6057d65c1985e55c34aca2a8f6ed51d7963b8a
BLAKE2b-256 checksum
How to use checksums
a331c31b177101eaf078cdab54cb6af514116bb97a0ed05b29e1e51ada196c49
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.19 This release

2 release files

0.2.12

2 release files

0.2.11

2 release files

0.2.10

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page