Track LLM token usage, enforce limits, and trigger alerts — across OpenAI, Groq, OpenRouter, AWS Bedrock, and custom providers.
Project description
TokenGuard
Production-ready token tracking, policy evaluation engines, budget limits, and alerts for LLM applications.
⚡ Quick Start
# Sync exact tracking with a Sliding Window Policy
from token_guard import TokenGuard, SlidingWindowPolicy
policy = SlidingWindowPolicy(limit=50_000, window=3600)
guard = TokenGuard(policy=policy)
result = guard.track_usage("alice", input_tokens=42, output_tokens=15)
print(result.total_tokens) # 57
print(result.limit_exceeded) # False
print(result.cumulative_usage.total_tokens) # 57
# Async exact tracking with Token Bucket Policy
from token_guard import AsyncTokenGuard, AsyncTokenBucketPolicy
async_policy = AsyncTokenBucketPolicy(capacity=10_000, refill_rate=100.0)
async_guard = AsyncTokenGuard(policy=async_policy)
result = await async_guard.track_usage("bob", input_tokens=42, output_tokens=15)
print(result.total_tokens) # 57
print(result.limit_exceeded) # False
Table of Contents
- Why TokenGuard?
- Architecture
- Features
- Installation
- Documentation Guides
- Provider Compatibility
- Project Structure
- Examples
- Running Tests
- Roadmap
- Contributing
- License
🧠 Why TokenGuard?
LLM calls are billed per token (inputs + outputs). Unchecked application consumption can quickly lead to unexpected cost spikes, upstream API rate-limiting blockages, or abuse.
TokenGuard acts as a lightweight, thread-safe, and event-loop-safe middleware layer. Use it to:
- Evaluate Flexible Policies: Enforce Sliding Window, Token Bucket, Fixed Window, Leaky Bucket, Cost, Quota, or Role-based policies.
- Prevent Cost Spikes: Set and enforce strict token usage budgets per user, model, or session.
- Unify Tracking: Track consumption across OpenAI, Groq, OpenRouter, and AWS Bedrock under a single API.
- Flexible Storage: Keep track in-memory (dev) or plug in Redis or SQLite (prod) with a single config change.
- Proactive Alerts: Fire warnings and webhook notifications (Slack, console) the instant thresholds are crossed.
🏗️ Architecture
TokenGuard splits concerns into clean interfaces:
- Counters: Tokenizer logic (tiktoken, HuggingFace transformers, or Bedrock count API).
- Policy Engine: Evaluates request rules (Sliding Window, Token Bucket, Fixed Window, Cost, Quota, Role).
- Storage: Persists cumulative records (Memory, Redis, SQLite).
- Alerts: Triggers limit-exceeded warning dispatches.
graph TD
App[Application] --> TG[TokenGuard / AsyncTokenGuard]
TG --> Counter[Token Counter]
TG --> Policy[Policy Evaluator]
Policy --> Storage[Storage Backend]
Policy --> Alerts[Alert Manager]
TG --> Result[TrackResult]
✨ Features
| Feature | Detail |
|---|---|
| Policy Engine | Sliding Window, Token Bucket, Fixed Window, Leaky Bucket, Cost, Quota, Role |
| Multi-Provider Counting | OpenAI (exact local), Groq, OpenRouter, AWS Bedrock |
| Exact Tracking | track_usage() records exact token metrics directly from API payloads |
| Pluggable Storage | Seamlessly swap backends (InMemory, Redis, SQLite) with one config line |
| Budget Enforcement | Track usage against configurable limits per user_id |
| Extensible Alerts | Console, Slack, webhooks, or custom handlers |
| Auto-Detect Backend | Auto-detect model tokens based on model name strings |
| FastAPI & Async Ready | Full async entry points and async-native database integrations |
| Robust Test Suite | 180 offline unit and integration tests |
📦 Installation
# Core package (includes OpenAI/tiktoken local counting, policies, and memory storage)
pip install llm-token-guard
# Install optional backends & providers
pip install "llm-token-guard[redis]" # Redis storage support
pip install "llm-token-guard[sqlite-async]" # Async SQLite (aiosqlite) support
pip install "llm-token-guard[groq]" # Groq HuggingFace tokenizers
pip install "llm-token-guard[bedrock]" # AWS Bedrock boto3 exact counts
pip install "llm-token-guard[all]" # All optional dependencies
📖 Documentation Guides
Advanced configuration, setup patterns, and code integrations are organized into individual guides:
1. Policy Engine
- Sliding Window, Token Bucket, Fixed Window, and Leaky Bucket policy configurations.
- Cost Limits ($/day), Quota Caps (tokens/day), and Role-based limit evaluation.
- Combining multiple policies in
TokenGuard&AsyncTokenGuard. - Extending custom policies via
BasePolicyorAsyncBasePolicy.
2. Token Counting & Providers
- Tiktoken exact token counts for OpenAI and OpenRouter.
- HuggingFace tokenizer integrations for Groq models.
- AWS API-driven exact counting for AWS Bedrock.
- Provider accuracy comparison table.
3. Storage Backends
- Using default
InMemoryStorage. - Setting up connection pools, keys namespaces, and TTLs in
RedisStorage. - Configuring persistent file storage via
SQLiteStorage. - Initializing via Environment Variables or Configuration Dictionaries.
4. FastAPI Integration
- Adding
AsyncTokenGuardto standard web applications. - Managing exact API counts inside async routes without blocking.
- Guide to API commands (
curl) for tracking, checking, and resetting.
5. Async Support
- Writing non-blocking async codebases with
AsyncTokenGuard. - Selecting async storage backends (
AsyncInMemoryStorage,AsyncRedisStorage,AsyncSQLiteStorage). - Configuring mixed sync/async alert triggers.
6. Custom Backends
- Subclassing
BaseTokenCounterand registering withCounterFactory. - Subclassing
BaseStorageand registering withStorageFactoryfor databases (e.g. Postgres, DynamoDB).
📊 Provider Compatibility
| Provider | Accuracy | Counting Method | Async Compatible | API Dependency |
|---|---|---|---|---|
| OpenAI | 100% (Exact) | Local tiktoken |
Yes | None |
| Groq (Default) | ~95% | Local tiktoken cl100k |
Yes | None |
| Groq (Transformers) | 100% (Exact) | Local AutoTokenizer |
Yes | transformers |
| AWS Bedrock (Local) | ~85% - ~95% | Local estimator | Yes | None |
| AWS Bedrock (API) | 100% (Exact) | AWS CountTokens API | Yes | boto3 |
| OpenRouter | ~85% - 100% | Local estimator | Yes | None |
| Direct Tracking | 100% (Exact) | track_usage(input, output) |
Yes | None |
🏗️ Project Structure
token_guard/
├── docs/ # Detailed guides and reference docs
├── token_guard/ # Core library source code
│ ├── counters/ # Token counters (OpenAI, Groq, etc.)
│ ├── engine/ # Policy evaluators and execution pipelines
│ ├── policies/ # Rate limiting, cost, quota, and role policies
│ └── storage/ # Storage backends (Memory, Redis, SQLite)
├── tests/ # Test suite (sync & async)
├── example_fastapi.py # FastAPI integration demo
└── pyproject.toml # Build configuration and dependencies
🚀 Examples
Ready-to-run examples demonstrating different configuration patterns:
- FastAPI Integration: Async token limits and route handling.
- Multi-Provider Demo: Basic usage mapping different counter and storage backends.
🧪 Running Tests
pip install -e ".[dev]"
# Run all offline sync and async tests (no API keys required)
pytest tests/ -v
# Run integration tests (requires GROQ_API_KEY env var)
export GROQ_API_KEY=gsk_...
pytest tests/test_groq_integration.py -v -s
🗺️ Roadmap
- Multi-provider token counting — OpenAI, Groq, OpenRouter, Bedrock ✅
- Auto-detect provider —
CounterFactory.auto()✅ - Pluggable storage — Memory, Redis, SQLite ✅
-
StorageFactory—from_env(),from_url(),from_config()✅ - Redis connection pooling + TTL +
from_url()+ping()✅ - GitHub Actions CI/CD — auto-publish on version tag ✅
- Exact token tracking —
track_usage()with API-reported counts ✅ - Async support —
async def track(...)for async frameworks ✅ - Policy Engine (v0.5.0) — Sliding Window, Token Bucket, Fixed Window, Leaky Bucket, Cost, Quota, Role policies ✅
- Budget warnings — alert at configurable % (e.g. 80%) before hard limit
- Prometheus metrics — expose
token_guard_tokens_totalcounter - Vertex AI / Cohere — dedicated exact-count backends
- PostgreSQL / DynamoDB — built-in storage backends
🤝 Contributing
Contributions are welcome! Please follow these basic guidelines:
- Fork the repository and create a feature branch.
- Ensure the full test suite passes locally before submitting your PR:
pytest tests/ -v
- Follow PEP 8 style standards.
📄 License
MIT ©Abhijit Gunjal — see LICENSE for details.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_token_guard-0.5.0.tar.gz.
File metadata
- Download URL: llm_token_guard-0.5.0.tar.gz
- Upload date:
- Size: 46.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6268f21ee4054aefc1bdfee09117ef86918f6a8a475c9aee325fbc2d17c91492
|
|
| MD5 |
72059d3f57f1d5bade627e508706de71
|
|
| BLAKE2b-256 |
79bb9ba9ffaefcd627be32ea94871fc84856cdd5fe73d9f1d9bc4ce660a1a0fc
|
Provenance
The following attestation bundles were made for llm_token_guard-0.5.0.tar.gz:
Publisher:
publish.yml on abhijitgunjal/token_guard
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_token_guard-0.5.0.tar.gz -
Subject digest:
6268f21ee4054aefc1bdfee09117ef86918f6a8a475c9aee325fbc2d17c91492 - Sigstore transparency entry: 2212091046
- Sigstore integration time:
-
Permalink:
abhijitgunjal/token_guard@bfd3eb6aa64c81f5084870aabd357c1f275daf33 -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/abhijitgunjal
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@bfd3eb6aa64c81f5084870aabd357c1f275daf33 -
Trigger Event:
push
-
Statement type:
File details
Details for the file llm_token_guard-0.5.0-py3-none-any.whl.
File metadata
- Download URL: llm_token_guard-0.5.0-py3-none-any.whl
- Upload date:
- Size: 50.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
533b5310f918a559a5dd1b3aae0caf92de70581f3d750c6d80b71e90ff2a577c
|
|
| MD5 |
bb20a0b846b914fd74af863548fb9528
|
|
| BLAKE2b-256 |
5ff90127badb768ee6df793189f67611a75f5129b589fbd82c50a642467a78ce
|
Provenance
The following attestation bundles were made for llm_token_guard-0.5.0-py3-none-any.whl:
Publisher:
publish.yml on abhijitgunjal/token_guard
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_token_guard-0.5.0-py3-none-any.whl -
Subject digest:
533b5310f918a559a5dd1b3aae0caf92de70581f3d750c6d80b71e90ff2a577c - Sigstore transparency entry: 2212091127
- Sigstore integration time:
-
Permalink:
abhijitgunjal/token_guard@bfd3eb6aa64c81f5084870aabd357c1f275daf33 -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/abhijitgunjal
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@bfd3eb6aa64c81f5084870aabd357c1f275daf33 -
Trigger Event:
push
-
Statement type: