assurance-budget
Your agent didn't fail. It just kept going.
The expensive runs are rarely the ones that crash. They are the ones that retried the same failing call fourteen times, or spent nineteen model calls summarising something nobody read, and finished with a plausible answer and a bill.
Point this at a run log and find out which ones did that.
pip install assurance-budget
assurance-budget runs.jsonl
On a system Python you may hit
error: externally-managed-environment(PEP 668). That is your OS protecting its packages, not this failing:python3 -m venv .venv && .venv/bin/pip install assurance-budget
1 of 4 runs hit a limit — 1 was going nowhere first — 1 ran past the clock
r-002
Stopped: 3 rounds repeating fetch(url=api/invoices) and failing the same way (timeout)
with nothing new read and no part of the goal closer.
r-003
Stopped after 20 frontier calls
r-004
Ran 660s against a 600s cap
Not tested by this log: iterations, retries. The log carries no events of that kind, so this
is silence rather than a pass.
That last paragraph is the part most tools leave out. A limit your log cannot exercise has not passed — it has not been tested, and reporting those the same way is how a green check comes to mean nothing.
What last night's runs actually did
Real output. Three runs, and only one of them looked like a failure from the outside.
$ assurance-budget runs.jsonl
1 of 3 runs hit a limit — 1 was going nowhere first
nightly-218
Stopped: 3 rounds repeating fetch(url=vendor/api/rates) and failing the same way (429)
with nothing new read and no part of the goal closer.
nightly-219
Stopped after 20 frontier calls
Not tested by this log: iterations, retries. The log carries no events of that kind, so
this is silence rather than a pass.
nightly-218 was rate-limited and kept asking. Sixteen identical calls, same error each time.
That is an agent bug: budget remained and spending it was the mistake.
nightly-219 hit the frontier-call ceiling. Different problem, different fix — that run was
doing real work and there was too much of it for one run.
The last paragraph is the part most tools leave out. This log carries no retry or iteration events, so those limits were not tested. Reporting them as passed would make a green check meaningless, so it says which questions it could not answer.
As a gate
assurance-budget runs.jsonl --fail-on-exhausted # exit 1 if any run hit a limit or stalled
When it will refuse your logs — and it probably will
$ assurance-budget transcript.jsonl
Cannot audit: line 1: no run identifier. Looked for run, run_id, runId, session, session_id,
trace_id. Without one, every event would be attributed to a single run and the per-run limits
would be meaningless. [exit 2]
An outside tester pointed this at his own agent transcripts and got exactly that. His logs were
role/message pairs — a chat log, not a run log — and a search of his whole machine found
nothing already in the shape this reads. He was right to stop there, and we recorded it rather than
asking him to rewrite production logs to suit us.
So be honest with yourself before installing: does anything you emit carry a per-run identity? If your agent writes one line per turn with no run id, this cannot help you yet, and no flag changes that.
This is not for you if
- Your logs have no per-run identity. See above. It is the most common reason this is useless.
- You want a dollar figure. Frontier calls are the cost proxy. Prices change per model, per provider and per week, and a number that goes stale silently is worse than a count that does not pretend to be money.
- You want it to stop a run. This reads a log of a run that already finished. Stopping one is the library, called from inside your own loop.
A limit a caller can raise is a suggestion
The ceilings are constants in the library, and every budget is clamped to them at construction. Ask for more and you do not get more:
assurance-budget runs.jsonl --tool-calls 5000 --json | grep tool_calls
# "tool_calls": 40
It clamps rather than erroring, on purpose. A caller asking for 5000 is expressing a preference the runtime declines — that is not a reason to abort somebody's task. Lower values pass straight through, because a caller may always be more conservative: that is how a cheap plan or an untrusted worker gets a shorter leash.
Raising a ceiling is a deliberate edit to a source file, which is the point.
The log
JSONL, one event per line. Field names are matched loosely — run/run_id/session,
action/tool/name, and so on — so most existing logs work without being rewritten.
{"run": "r-002", "action": "fetch(url=api/invoices)", "error": "timeout", "kind": "tool", "ts": 41.0}
kind is one of tool, frontier, retry, iteration and defaults to tool. Anything else is
refused rather than counted as something it isn't.
As a library — the half that prevents rather than reports
The audit tells you it already happened. This stops it happening:
from assurance_core.run_budget import Budget, ProgressWatch, Progress, Spend
spend = Spend(budget=Budget.allowing(tool_calls=20))
watch = ProgressWatch()
for step in range(100):
stopped = spend.charge_tool_call()
if stopped:
print(stopped.message)
break
stalled = watch.observe(Progress(action="fetch(url=api/invoices)", error="timeout"))
if stalled:
print(stalled.message)
break
assert spend.tool_calls <= 20
Two different stops, and the difference matters. Exhausted means the budget ran out.
Stalled means budget remains and spending it is the mistake — three rounds repeating the same
action, failing the same way, with nothing new read and no part of the goal closer.
In a pipeline
assurance-budget runs.jsonl --fail-on-exhausted
| exit | means |
|---|---|
0 |
audited, and nothing hit a limit or stalled |
1 |
audited, and --fail-on-exhausted found a run that did |
2 |
refused — the log could not be read, so there is no audit |
Honest limits
- It reads what your log records. A run that burned money in a way the log does not mention is invisible here, and no amount of analysis fixes that.
- There is no dollar limit. Frontier calls are the cost proxy. Prices change per model, per provider and per week; a number that goes stale silently is worse than a count that does not pretend to be money.
- Stall detection needs three rounds and both halves — identical action, error and result, and flat progress. A repeated action while evidence accumulates is a loop doing work, and stopping that would be the bug.
As an agent skill
skills/run-budget/ — drop it in and an agent reads your run logs the
right way. Its real content is that a limit the log cannot exercise has not passed, and that a
run which stalled is an agent bug while one which was exhausted may just be a job too big.
Where the rules live
assurance_core.run_budget, in assurance-core — pure
Python, no dependencies, no model involved in any of it. This package reads logs and calls it.
Licence
Apache-2.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file assurance_budget-0.1.2.tar.gz.
File metadata
- Download URL: assurance_budget-0.1.2.tar.gz
- Upload date:
- Size: 19.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4e8c2ba96f6a14a563d6518e578342fc9e73662b444a57e1abd8bc2a2781706e
|
|
| MD5 |
a914c469abc510f9b45a37ebe15b91ce
|
|
| BLAKE2b-256 |
292146d6b98c000b7325240faa8cd72569ac8c1b7de7ff94f55e5822aa7fe12a
|
Provenance
The following attestation bundles were made for assurance_budget-0.1.2.tar.gz:
Publisher:
publish.yml on i-ops-hq/assurance-budget
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
assurance_budget-0.1.2.tar.gz -
Subject digest:
4e8c2ba96f6a14a563d6518e578342fc9e73662b444a57e1abd8bc2a2781706e - Sigstore transparency entry: 2720919749
- Sigstore integration time:
-
Permalink:
i-ops-hq/assurance-budget@ca4fef6f9d74a87c666ac392c7f300e25d9cd12a -
Branch / Tag:
refs/tags/v0.1.2 - Owner: https://github.com/i-ops-hq
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ca4fef6f9d74a87c666ac392c7f300e25d9cd12a -
Trigger Event:
push
-
Statement type:
File details
Details for the file assurance_budget-0.1.2-py3-none-any.whl.
File metadata
- Download URL: assurance_budget-0.1.2-py3-none-any.whl
- Upload date:
- Size: 15.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d939878d04462898108e0d77692e2ff1ec7767a86cc56a5f93b177fc3ec47649
|
|
| MD5 |
0b18b3f6c09c273476cac0acd97d2b90
|
|
| BLAKE2b-256 |
d341d76de532c7b5fae9c36a7a17af5f08c32bb44b00f23552fb5d0ed45d09ab
|
Provenance
The following attestation bundles were made for assurance_budget-0.1.2-py3-none-any.whl:
Publisher:
publish.yml on i-ops-hq/assurance-budget
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
assurance_budget-0.1.2-py3-none-any.whl -
Subject digest:
d939878d04462898108e0d77692e2ff1ec7767a86cc56a5f93b177fc3ec47649 - Sigstore transparency entry: 2720919910
- Sigstore integration time:
-
Permalink:
i-ops-hq/assurance-budget@ca4fef6f9d74a87c666ac392c7f300e25d9cd12a -
Branch / Tag:
refs/tags/v0.1.2 - Owner: https://github.com/i-ops-hq
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ca4fef6f9d74a87c666ac392c7f300e25d9cd12a -
Trigger Event:
push
-
Statement type: