budget-guard-llm
Stop one user, one feature or one runaway loop from burning through your OpenAI or Anthropic bill.
budget-guard-llm wraps the official SDK clients, prices every call from the tokens the response actually
used, and enforces daily spend limits per user, per feature and in total. It also catches the same
request being sent over and over in a short time, which is what a stuck retry loop or agent looks like.
pip install budget-guard-llm
The package installs as budget-guard-llm and imports as budget_guard.
Two lines
from budget_guard import guard
client = guard(OpenAI(), user="u1", feature="chat") # or Anthropic(), AsyncOpenAI(), AsyncAnthropic()
The rest of your code stays the same: client.chat.completions.create(...),
client.messages.create(...), streaming, async, all as before.
Limits
Set them once at startup:
import budget_guard
from budget_guard import Limits, LoopDetection
budget_guard.configure(
limits=Limits(
per_user=1.00, # USD per user per day
users={"vip": 20}, # overrides for specific users
features={"chat": 5}, # per feature (per_feature= sets a default)
daily_total=50, # everything together
),
loop_detection=LoopDetection(max_repeats=10, window_seconds=60), # the default
on_exceeded="raise", # or "warn": log it and let the call through
on_violation=alert_me, # optional: called with the error in either mode
)
When a total has reached its cap, the next call raises BudgetExceeded before anything is sent to the
provider. A loop raises LoopDetected. Both subclass BudgetError:
from budget_guard import BudgetError
try:
reply = client.chat.completions.create(model="gpt-6-luna", messages=messages)
except BudgetError as err:
return "You've hit today's limit, try again tomorrow."
Limits and LoopDetection each accept action="raise" or "warn" to override on_exceeded, so you
can, for example, block loops but only warn on spend.
How the checks behave:
- Limits are daily and reset at midnight UTC.
- The call that crosses a cap is allowed, because its cost is only known once it returns.
- Calls still running count toward each cap at an estimated cost, so a burst of parallel calls can't all
get through before the first one is recorded. The estimate is the request's text at about 3 characters
a token, plus its
max_tokens(orreserve_output_tokens=, 1024 by default, when it sets none). Settingmax_tokenson your calls makes this protection tighter. - A limit of
0blocks that user or feature entirely. - A loop is the same request (model, messages and options) for the same user and feature, sent more than
max_repeatstimes insidewindow_seconds. Two users asking the same question never count together. Passloop_detection=Noneto turn it off.
Who a call is for
Set defaults when wrapping, then override per request. Ids can be strings or numbers; they're compared
as text, so user=42 and users={"42": 5} match. budget_context works across threads and asyncio
tasks, so it fits request middleware:
from budget_guard import budget_context
with budget_context(user=request.user.id):
client.chat.completions.create(...)
client.with_context(feature="summary").chat.completions.create(...) # a re-scoped copy
Reading the totals
budget_guard.spent(user="u1") # Decimal USD today
budget_guard.spent(feature="chat")
budget_guard.spent() # daily total
budget_guard.remaining(user="u1") # left under the limit, or None if no limit applies
To log every priced call, pass on_record= to configure(). It receives a CallRecord with the
provider, model, user, feature, token usage and cost.
What is tracked
| Provider | Methods |
|---|---|
| OpenAI | chat.completions.create/parse/stream, responses.create/parse/stream, completions.create |
| Anthropic | messages.create/parse/stream, beta.messages.create/parse/stream |
Streams are counted when they finish, are closed, or are dropped part-way. For OpenAI chat streams the
guard turns on stream_options.include_usage and hides the extra usage chunk unless you asked for it
yourself. OpenAI only reports usage at the end of a stream, so a stream cut short is counted at its
estimated cost (see above) and a warning is logged. client.with_options(...) and Anthropic's client.with_middleware(...) stay guarded.
with_raw_response and with_streaming_response calls go through the limit and loop checks, but their
cost isn't counted yet. Not tracked at all yet: batches, Anthropic's tool_runner, embeddings, images and
audio.
Prices
Prices ship in budget_guard/pricing.json (USD per 1M tokens, with cached-input and long-context rates)
and can be changed at runtime:
from budget_guard import default_pricing, PricingTable
default_pricing().set("openai", "my-finetune", {"input": "1.00", "output": "4.00"})
budget_guard.configure(pricing=PricingTable.from_file("my_prices.json"))
Dated model names such as gpt-6-luna-2026-03-01 use the base model's price. A call to a model with no
price goes through uncounted with a warning; configure(on_unknown_model="raise") blocks it instead.
Several servers
Totals are kept in memory, so each process counts on its own and totals reset on restart. To share
them across servers, implement the four methods of budget_guard.Storage (add, get, hit, clear)
on Redis or a database and pass it as configure(storage=...). A Redis backend is planned.
Development
pip install -e ".[dev]"
python -m pytest
The tests run the real OpenAI and Anthropic SDKs against a fake HTTP transport, so nothing is sent and nothing is billed. Release steps are in RELEASING.md.
License
MIT
Metadata
Release files for budget-guard-llm 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| budget_guard_llm-0.1.2.tar.gz | 32.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| budget_guard_llm-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 59.2 kB
Release files / budget_guard_llm-0.1.2.tar.gz
| Download URL | budget_guard_llm-0.1.2.tar.gz |
|---|---|
| Size | 32.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4ec90c59f2f5b489f780c05d8a730164b252ec150d37298b32069cbf212b0878
|
|
BLAKE2b-256 checksum How to use checksums |
b103af3c98fe1e926e0fcba917630db9d0c1b916ce846ace926cbbd2065055b3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / budget_guard_llm-0.1.2-py3-none-any.whl
| Download URL | budget_guard_llm-0.1.2-py3-none-any.whl |
|---|---|
| Size | 26.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f28965c0c9c86e5f6b8d1b620794403c999370dd3cac0bd8f2d734e4805378ce
|
|
BLAKE2b-256 checksum How to use checksums |
f8db6080d92528849926a02d0d7379acd68667e699de0c983c625fcdf5435c97
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log