Skip to main content

agent-guard

Stop your agent before it burns $400 overnight
budget caps, loop detection and circuit breakers for AI agents
zero dependencies · no proxy · nothing leaves your process

A terminal running examples/basic.py: an agent with a $0.05 budget is stopped on its third call, and the report shows the budget bar at 114%

An agent that stops making progress does not stop running. It calls the same tool with the same arguments, or bounces between two tools forever, buying a fresh context window on every pass. agent-guard is the part of your loop that notices — and the receipt that tells you what it cost.

pip install agent-budget-guard-py
from agentguard import Guard

guard = Guard(max_usd=1.00, max_steps=25)

with guard:
    while True:
        with guard.step() as step:
            response = client.chat.completions.create(...)
            step.record(response)                     # tokens and cost, extracted
            with step.tool("search", {"q": query}):   # fingerprinted for loop detection
                results = search(query)

print(guard.report())

Three things happen on their own: step.record(response) pulls tokens and the model out of any SDK response, step.tool(...) fingerprints the call so a repeat is caught before it runs a fourth time, and guard.report() prints the receipt. Python 3.10+, no runtime dependencies — not even a provider SDK.

Jump to: See it work · What it catches · Wired into your stack · Models it doesn't know · Running it in CI · Honest answers · API reference · All the details


The problem

A budget check that runs after the call can only tell you what you already spent. By the time a cap notices a stuck agent, the money is gone — and the usual failure is not a crash, it is a loop that looks busy.

Two things stop it, and they stop different things:

Catches When it fires
A ceiling — max_usd, max_tokens, max_steps, max_seconds spend that keeps climbing after the call that crossed it
A loop detector — four of them, on by default an agent repeating itself before the repeat runs

agent-guard does both, in your process, using the standard library. No proxy, no server, no account, and nothing is sent anywhere.

See it work

A research agent asked the same question three times. The third identical search_web call is where it stopped — not because the budget ran out, but because the call was identical to the two before it. This is the whole library in one run:

guard = Guard(max_usd=5.00, name="research-agent")
with guard.step(tag="search") as step:
    step.record("gpt-4o", input_tokens=8_000, output_tokens=400)
    with step.tool("search_web", {"query": "weather in oslo"}):
        results = search("weather in oslo")        # 3rd time: stopped here
turn 1: searched, spent $0.0240
turn 2: searched, spent $0.0480

LoopDetected: Loop detected [repeat]: the same call appeared 3 times in the
last 3 steps: search_web({"query":"weather in oslo"}).
agentguard  research-agent
================================================================
  wall time   0ms             steps     3
  llm calls   3               tokens    25,200  (in 24,000 / out 1,200)

  limits
    budget   $0.072 / $5                   1.4%  [................]

  by model
    gpt-4o   3 calls      $0.072      24,000 in / 1,200 out

  tripped: loop [repeat] the same call appeared 3 times in the last 3 steps: ...

The difference between that and an overnight bill is $0.07. Both output blocks are real — the exception line is wrapped to fit here and the report is truncated the same way. Two examples run offline and show the same machinery:

python examples/loop_detection.py   # all four detectors, plus a healthy run that must not trip
python examples/basic.py            # a budget cap instead: stopped at $0.057 of a $0.05 limit

What it catches

Every limit is optional and independent. A guard with no limits still detects loops and still produces a report.

Limit Trips when
max_usd=1.00 Known spend exceeds one dollar
max_tokens=500_000 Input + output tokens exceed the allowance
max_steps=25 A 26th guard.step() is opened
max_seconds=300 Wall-clock time since the guard was created
scoped_budgets={"tool:search": 1.0} One tool, or one tag, exceeds its own share

A run-level ceiling can only tell you the run got expensive. It cannot tell you which part of it did, so one tool in a retry storm can spend the whole allowance before max_usd notices — and the answer to "what ate the budget?" arrives in a report you read afterwards. A scoped budget answers it while it is happening, by naming the tool:

  limits
    budget     $0.0125 / $1                  1.2%  [................]
  ! tool:fetch $0.0125 / $0.012            104.2%  [################]

The four detectors, and what each one is actually for:

Detector Catches Fires on
RepeatDetector The same call, identical arguments 3rd identical call
CycleDetector A, B, A, B — two tools bouncing forever 2nd full pass
SimilarityDetector Paraphrasing: search("python asyncio") → search("python asyncio ") 4th near-identical call
NoProgressDetector A progress marker that never moves 6th unchanged marker

Every verdict is a sentence, because "why did you kill my agent?" is the first question anyone asks. These are real, from examples/loop_detection.py:

exact repeat                     -> repeat
                                    the same call appeared 3 times in the last 3 steps: search_web({"query":"weather in oslo"})
two-step ping-pong               -> cycle
                                    a 2-step pattern repeated 2 times: read_file({"path":"app.py"}) -> write_file(...)
paraphrased calls                -> similarity
                                    4 near-identical calls (>= 95% similar) in the last 4 steps: search_web(...)
no progress                      -> no-progress
                                    the progress marker did not change for 6 consecutive observations: {"rows_written":0}

healthy varied work              -> clean, as it should be

That last line matters as much as the others: a detector that fires on healthy work is a detector you turn off. Each threshold is tuned so its scenario is reachable and nothing else is, and every one is adjustable per guard.

Wired into your stack

Three depths. Nothing is required beyond the first.

1. Manual — works with anything, including a while loop you wrote by hand:

guard = Guard(max_usd=1.0)
guard.record("gpt-4o", input_tokens=1200, output_tokens=300)
guard.check()

2. Structured — one step per iteration, tools fingerprinted for you. This is the Quickstart above.

3. Wrapped client — every call accounted with no call-site changes:

from openai import OpenAI
from agentguard.adapters.openai import guard_openai

client = guard_openai(OpenAI(), max_usd=1.0, max_steps=25)
response = client.chat.completions.create(...)   # recorded automatically

Adapters are pure duck-typing — they never import a provider SDK — so the same wrapper covers Anthropic, LiteLLM, OpenRouter, vLLM, Together, Groq and Azure OpenAI:

from agentguard.adapters.anthropic import guard_anthropic

client = guard_anthropic(Anthropic(), max_usd=2.0)

Streaming works. With stream=True the response is wrapped, chunks pass through untouched, and usage is recorded once when the stream drains. Anthropic's messages.stream() and async clients (async for) are covered the same way.

LangGraph takes one callback handler:

from agentguard.integrations.langgraph import guard_langgraph

handler = guard_langgraph(Guard(max_usd=1.0, max_steps=25))
graph.invoke(inputs, config={"callbacks": [handler]})

Or a decorator, for a request handler rather than a loop:

from agentguard import current_guard, guarded

@guarded(max_usd=0.50, max_steps=20)
def summarise(url: str) -> str:
    guard = current_guard()
    ...

@guarded(max_usd=...) creates a fresh guard per call — the right default when one caller exhausting a budget must not stop the next. Pass guard= a shared guard when spend should accumulate across calls.

The decorator covers async def too: the guard stays entered across every await rather than only until the coroutine object is created, so current_guard() is live inside an async body — and inside an async generator for the whole iteration.

Fanning work out across threads? No thread inherits a contextvars context, so a call made on a pool worker is still counted — the money is never lost — but belongs to no step and no tool, and a scoped_budgets cap on that tool never fires. Two ways to carry the attribution over, both captured where you call them:

with guard.step(tag="fan-out"), guard.tool("fetch"):
    results = list(pool.map(guard.bind(fetch), urls))   # per call
    # or: with guard.context():  ...                    # per block, inside the worker

Refuse a call before paying for it. A post-hoc check can only report overspend; preflight() refuses a call whose worst case will not fit in what is left:

client = guard_openai(OpenAI(), max_usd=0.05, preflight=True)
# BudgetExceeded: refused before spending: a gpt-4o call could reach $0.6,
# over the $0.05 limit (already spent $0)

Resume without refilling the budget. An agent that checkpoints its own state can checkpoint its spend too, so a restart does not hand it a fresh cap:

write_checkpoint({"cursor": 41, "guard": guard.snapshot()})

# later, in a new process
guard = Guard.from_snapshot(read_checkpoint()["guard"], max_usd=5.0)
guard.remaining_usd      # what is actually left, not the full budget

Three decisions worth knowing, all in the details: detector windows survive a checkpoint (a loop that spans one is still a loop); wall-clock time does not (max_seconds caps this process, and restoring elapsed time would trip on time you never spent); and a snapshot it cannot read exactly is refused rather than half-applied.

Models it doesn't know

The bundled table carries 119 models — current families from OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral and Qwen, dated 2026-09 in PRICING_AS_OF. It cannot know your fine-tune, your gateway's aliases, or your negotiated rate:

agentguard config set my-finetune-v3 3 12 --cached 0.3
agentguard config alias acme/fast claude-3-5-haiku
agentguard config disable gpt-4          # do not trust this bundled price

A model it has never heard of is never billed at a guessed rate. It is counted as unpriced, kept out of the budget arithmetic, and reported loudly — because a safety tool that silently assumes $0.00 is worse than no safety tool at all:

  by model
    gpt-4o            1 call       $2.52   1,000,000 in / 2,000 out
    acme-rerank-v3    1 call    unpriced      40,000 in / 0 out

  ! 1 call(s) had no known price and are excluded from the budget:
      acme-rerank-v3
    Price them with `agentguard config set <model> <input> <output>`,
    or pass Guard(pricing={...}) in code.

A model disabled by disable becomes unpriced the same way, so the excluded spend is visible instead of quietly billed at a number you rejected. Everything is overridable in code, and code always beats a file.

Running it in CI

guard.save() in the worker, agentguard report in CI — the reading process needs nothing installed:

agentguard report run.json           # render a report saved by guard.save(...)
agentguard report run.json --json
agentguard pricing                   # effective table, with a source per model
agentguard pricing gpt-4o
agentguard pricing --update          # refresh prices from a public catalogue
agentguard pricing --status          # is a snapshot in effect, and from where
agentguard config path               # where config is read from, what is ignored

Honest answers

Zero dependencies, and it stays that way. No pydantic, no httpx, no provider SDK — standard library only. CI fails the build if a runtime dependency ever appears in the wheel, so it can be dropped into a stack that vendors its dependencies, pinned to an old Python, or shipped inside a Lambda without touching a lockfile.

A price table goes stale. This one is dated, and a provider can change a rate or retire a model the day after. Retired models keep their last published price rather than being dropped, because removing a name silently turns every call to it unpriced. Anything you bill on should be verified, and anything you depend on should be configured. When you would rather not wait for a release, refresh it on purpose:

agentguard pricing --update     # one opt-in download, checksummed and cached

That is the only command in this library that touches the network. Nothing fetches prices at import, on a timer, or in the background — and the result merges underneath your config, so a public catalogue can never override a rate you set deliberately. A snapshot that has been edited or truncated fails its checksum and is refused rather than billed. See the details.

What it is not. Not an observability platform (nothing is sent anywhere, there is no server and no background thread), not a proxy (it cannot see traffic it was not told about), not a tokenizer (pre-flight estimates are heuristic), and not a substitute for provider-side spend limits — use both. agent-guard stops your loop; the provider's limit is what saves you when your process dies with a request already in flight.

Limitations, stated plainly. A response that reports no usage cannot be priced: it warns once and counts as unpriced rather than inventing a number. A stream is priced when drained, so for OpenAI-compatible streams pass stream_options={"include_usage": True}. Pre-flight input counts are estimated from the serialised prompt.

Documentation

docs/API.md Every public name: Guard and Step, recording, loop detection, accounting, pricing, config, exceptions, adapters, decorators, the CLI
docs/DETAILS.md The reasoning: every detector and its tuning, the pricing config's trust model, the checkpoint format, design principles
examples/ Ten runnable programs — budget cap, all four detectors, wrapped client, streaming, LangGraph, custom models, checkpointing, scoped budgets, a price refresh, and a realistic report
CHANGELOG.md Release history
ROADMAP.md What is planned next
CONTRIBUTING.md The four constraints the library is built to

Every docstring example in the package runs as a test, so the documentation cannot drift from the behaviour — and the API reference is checked against the library the same way: its signature blocks, its export table and its exception tree are all asserted in tests/test_api_reference.py. python -m unittest discover -s tests -t . runs the whole suite — 746 tests, no network, no fixtures.

License

MIT — see LICENSE.

Metadata

Release files for agent-budget-guard-py 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-budget-guard-py 0.4.0
File Size Uploaded
agent_budget_guard_py-0.4.0.tar.gz 224.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-budget-guard-py 0.4.0
File Interpreter ABI Platform
agent_budget_guard_py-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 328.0 kB

Release files / agent_budget_guard_py-0.4.0.tar.gz

Download URL agent_budget_guard_py-0.4.0.tar.gz
Size 224.0 kB
Tags Source
SHA-256 checksum
How to use checksums
897f7eccd001684fcf423b15152493aafd9f89d1facc763fd2ea02bc5c4b8863
BLAKE2b-256 checksum
How to use checksums
570ed3282722b419eab4d070d6980e9b67c59073380b13262b8b9c89b7f1eed0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / agent_budget_guard_py-0.4.0-py3-none-any.whl

Download URL agent_budget_guard_py-0.4.0-py3-none-any.whl
Size 104.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d11f80c25d1b435eb18c515f6fd185be115872338d4d6c4c92c8e0808655e648
BLAKE2b-256 checksum
How to use checksums
c93a183bd689df47730350dda280a6302ccf7a79c27bc3adc4ed1cda3519b966
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.2

2 release files

0.3.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page