tokengov
Treat LLM spend like cloud spend.
tokengov is a lightweight governance toolkit for LLM workflows: track
tokens per request, enforce per-workflow dollar budgets, cache responses
(exact + semantic), route easy prompts to cheap models, and lint your system
prompts for bloat. It gives an LLM feature the same cost discipline a
well-run platform team applies to EC2 — because token spend is compute
spend.
No servers, no SaaS, no API keys required to run the demo. Two dependencies
(pydantic, tiktoken), pure Python 3.10+.
The governance loop
flowchart LR
A[Prompt] --> R[Router<br/>complexity score]
R -->|easy| C1[Cheap model]
R -->|hard| C2[Frontier model]
C1 --> X{Cache hit?}
C2 --> X
X -->|yes| L[Ledger<br/>$0]
X -->|no| G[Generate]
G --> T[TokenTracker<br/>counts + prices]
T --> B{Budget OK?}
B -->|over limit| E[BudgetExceeded<br/>workflow stops]
B -->|under limit| L
L --> D[Report<br/>per-model / per-workflow]
D --> F[Lint prompts<br/>cut bloat]
F --> A
Install
pip install . # from a checkout
pip install ".[dev]" # with pytest + ruff
Usage
1. Track tokens and cost per request
from tokengov import track, build_report
with track("gpt-4o", workflow="support-bot") as t:
t.count_input(prompt) # or an int from your client's usage field
completion = client.chat(...)
t.count_output(completion)
print(build_report().render())
2. Enforce per-workflow budgets
from tokengov import Budget, BudgetExceeded, budget_guard, track
budget = Budget(limit=5.00, label="nightly-etl", warn_at=0.8)
try:
with budget_guard(budget), track("gpt-4o-mini", workflow="nightly-etl"):
rows = etl_step()
except BudgetExceeded:
alert("nightly-etl hit its $5 ceiling — investigate before retrying")
Or as a decorator for whole functions:
from tokengov import budgeted
@budgeted("gpt-4o-mini", limit=0.10, workflow="summarize")
def summarize(tracker, text: str) -> str:
tracker.count_input(text)
return client.chat(text)
3. Cache responses (exact + semantic)
from tokengov import TwoTierCache
cache = TwoTierCache(ttl_seconds=600, max_entries=1000, similarity_threshold=0.95)
cache.set(prompt, model, response)
answer = cache.get(prompt, model) # exact or near-duplicate hit
print(cache.stats.hit_rate)
The semantic tier uses a deterministic hashing-vector embedding (no network, no weights). Swap in a real embedder any time:
TwoTierCache(embed_fn=my_embedding_model.encode)
4. Route easy prompts to cheap models
from tokengov import Router
router = Router(cheap_model="gpt-4o-mini", frontier_model="gpt-4o")
choice = router.route("Summarize this paragraph in one sentence.")
print(choice.model) # gpt-4o-mini
print(choice.rationale) # [cheap] no complexity signals fired (score 0.00)
hard = router.route(
"Derive the trade-offs between two-phase commit and Raft, "
"and output a JSON table of failure modes."
)
print(hard.model) # gpt-4o
print(hard.rationale) # [frontier] score 0.75 >= 0.50: reasoning-demand, ...
The classifier is an explainable weighted heuristic — length, reasoning keywords, structured-output demands, domain terms. No opaque ML. Replace it later without touching anything else:
Router(cheap_model=..., frontier_model=..., classifier_fn=my_real_classifier)
5. Report spend
tokengov prices # pricing table
tokengov report # render this process's ledger
tokengov report --input ledger.json # render an exported ledger
tokengov report --csv > spend.csv # feed your finops dashboard
6. Lint prompts for bloat
tokengov lint prompts/system.txt
prompts/system.txt:WARN filler:politeness-padding (line 3): 2x filler phrase ... [~2 tokens]
prompts/system.txt:ERROR duplicated-instruction (line 9): line duplicates line 4 ... [~11 tokens]
prompts/system.txt:ERROR oversized-system-prompt (line 1): prompt is 1,203 tokens, 691 over the 512-token ceiling (~$0.000104/call at gpt-4o-mini input rates) [~691 tokens]
3 finding(s), ~704 tokens recoverable.
7. Run the end-to-end demo (zero API keys)
python examples/governed_chat.py
A mock LLM client drives the full stack — routing, cache, budget, tracker, report — and prints the final cost table.
Why I built this
I've spent 13+ years running engineering teams, and the most expensive sentence in any architecture review is "we'll optimize the LLM bill later." Later is where budgets go to die. The teams I've led that treated token spend as a first-class operational metric — measured per workflow, budgeted like cloud spend, and reviewed in the same forum as AWS bills — routinely cut inference costs by 40–70% without degrading user-facing quality. Token efficiency isn't a prompt trick; it's an engineering culture artifact: you get what you instrument, you keep what you enforce in CI.
tokengov is my reference implementation of that culture in a box:
- Measurement before optimization — per-workflow ledgers make cost a debuggable signal, not a monthly surprise.
- Guardrails, not guidelines — budgets that fail the call (like a CI check) beat wiki pages that nobody reads.
- Cheap by default, expensive by justification — routing with a written rationale turns "why is this on the frontier model?" into a reviewable artifact.
- Explainability over cleverness — the router ships as weighted rules you can defend in a design review; a learned classifier can plug in later.
Design docs
- High-Level Design — the governance loop and module map
- Low-Level Design — data structures, cache internals, ledger schema
- ADR 0001: Explainable heuristics over opaque ML for routing
- ADR 0002: Contextvar-scoped tracking
Roadmap
- Async-first
TokenTrackerwith provider adapters (OpenAI, Anthropic) that auto-report usage fields - Budget backends: Redis for multi-process enforcement, file-based for CI
- Router signal calibration CLI — score prompts from production traffic and re-fit weights
-
tokengov lint --cimode with a configurable token budget gate for pre-commit / PR checks - Prometheus exporter for the ledger
- Cost allocation tags propagated from request headers into reports
License
MIT — see LICENSE.
Metadata
Release files for tokengov-svkmsr6 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokengov_svkmsr6-0.1.0.tar.gz | 29.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokengov_svkmsr6-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.1 kB
Release files / tokengov_svkmsr6-0.1.0.tar.gz
| Download URL | tokengov_svkmsr6-0.1.0.tar.gz |
|---|---|
| Size | 29.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f5fa6c19670465d25e138d483d3b6bd80cc11ee49c17571be4c318aaa30d7c37
|
|
BLAKE2b-256 checksum How to use checksums |
a7f2cd422f82c552c910ff0aa84ce78b34cb39eedbdd637c1d8bea1c23973ac0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / tokengov_svkmsr6-0.1.0-py3-none-any.whl
| Download URL | tokengov_svkmsr6-0.1.0-py3-none-any.whl |
|---|---|
| Size | 23.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9d8963fb4915159401cd67bb357638f147d295510ab325f9889c15d2998f0092
|
|
BLAKE2b-256 checksum How to use checksums |
996943378f9ff6a8521322866defee348d5553c945fe8a121c8069281abe52a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|