Skip to main content
TokenShield

🛡️ TokenShield

Real-time token cost monitoring, budget enforcement, and optimization for LLM applications.

Python 3.11+ License: MIT Tests Coverage

Stop burning money on LLM API calls. TokenShield gives you per-request cost tracking, budget gates, and automatic optimization — before the invoice arrives.


The Problem

Month 1:  $50    "This is cheap!"
Month 2:  $200   "Growth is normal"
Month 3:  $3,400 "WHAT HAPPENED?!"

LLM costs are invisible until the bill arrives. A single misconfigured loop, a verbose system prompt, or an unbound tool list can 10x your spend overnight.

The Solution

from tokenshield import Shield, BudgetPolicy

shield = Shield(
    model="gpt-4o",
    policy=BudgetPolicy(
        max_cost_per_request=0.05,     # $0.05 per request
        max_cost_per_hour=2.00,        # $2/hour
        max_cost_per_day=20.00,        # $20/day
        alert_threshold_pct=80,        # Alert at 80% of any limit
    )
)

# Wrap any LLM call
result = shield.call(
    messages=[{"role": "user", "content": "Summarize this order"}],
    tools=tool_schemas,
)

print(shield.report())
# ┌─────────────────────────────────┐
# │ Requests today:     142         │
# │ Tokens (in/out):    89K / 12K   │
# │ Cost today:         $4.23       │
# │ Budget remaining:   $15.77      │
# │ Avg cost/request:   $0.030      │
# │ Most expensive:     search (48%)│
# └─────────────────────────────────┘

Features

Feature Description
Cost Tracking Per-request, per-hour, per-day cost accumulation with model-aware pricing
Budget Gates Hard limits that reject calls before they execute (no surprise bills)
Alert Hooks Webhook/callback when approaching budget thresholds
Token Estimation Pre-flight token count estimation before calling the API
Model Pricing DB Built-in pricing for GPT-4o, Claude, Gemini, Mistral, and custom models
Optimization Tips Automatic suggestions: "Your system prompt is 4,200 tokens — consider trimming"
Dashboard Export JSON/CSV export for cost dashboards and observability tools
Async Support Full async/await support for high-throughput applications

Architecture

┌──────────────────────────────────────────────────────────┐
│                     Your Application                      │
├──────────────────────────────────────────────────────────┤
│                                                          │
│  ┌──────────┐   ┌──────────┐   ┌──────────────────────┐ │
│  │ shield   │──→│ estimator│──→│ budget_gate          │ │
│  │ .call()  │   │ (tokens) │   │ (allow / reject)     │ │
│  └──────────┘   └──────────┘   └──────────┬───────────┘ │
│       │                                     │            │
│       │         ┌──────────┐   ┌───────────▼──────────┐ │
│       │         │ tracker  │←──│ LLM API call         │ │
│       │         │ (costs)  │   │ (litellm / openai)   │ │
│       │         └────┬─────┘   └──────────────────────┘ │
│       │              │                                   │
│  ┌────▼──────────────▼─────┐   ┌──────────────────────┐ │
│  │ reporter                │   │ alert_hooks          │ │
│  │ (dashboard / export)    │   │ (webhook / callback) │ │
│  └─────────────────────────┘   └──────────────────────┘ │
│                                                          │
└──────────────────────────────────────────────────────────┘

Quick Start

pip install tokenshield

Basic Usage

from tokenshield import Shield

shield = Shield(model="gpt-4o")

# Track a call (wrap your existing LLM call)
result = shield.call(messages=[...])

# Check current spend
print(f"Today: ${shield.tracker.cost_today:.2f}")

Budget Enforcement

from tokenshield import Shield, BudgetPolicy

shield = Shield(
    model="gpt-4o",
    policy=BudgetPolicy(max_cost_per_request=0.10)
)

try:
    result = shield.call(messages=huge_prompt)
except shield.BudgetExceeded as e:
    print(f"Blocked! Estimated cost ${e.estimated_cost:.3f} exceeds limit")

Alert Hooks

shield = Shield(
    model="gpt-4o",
    policy=BudgetPolicy(max_cost_per_day=20.00, alert_threshold_pct=80),
    on_alert=lambda msg: slack.post(channel="#llm-costs", text=msg),
)

Optimization Suggestions

tips = shield.optimize(messages, tools)
# [
#   "System prompt is 3,800 tokens (63% of input). Consider compressing.",
#   "18 tools bound but only 3 used. Use dynamic tool binding to save ~2,250 tokens.",
#   "History has 45 messages. Consider windowing to last 20.",
# ]

Pricing Database

Built-in pricing (updated monthly):

Model Input ($/1M) Output ($/1M) Context
gpt-4o $2.50 $10.00 128K
gpt-4o-mini $0.15 $0.60 128K
claude-3.5-sonnet $3.00 $15.00 200K
claude-3-haiku $0.25 $1.25 200K
gemini-1.5-pro $1.25 $5.00 1M
mistral-large $2.00 $6.00 128K

Add custom models:

shield.pricing.add("my-finetuned-model", input=5.00, output=15.00)

Documentation

License

MIT — see LICENSE

Release files for tokenshield-ai 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenshield-ai 2.0.0
File Size Uploaded
tokenshield_ai-2.0.0.tar.gz 12.5 kB Details

Release files / tokenshield_ai-2.0.0.tar.gz

Download URL tokenshield_ai-2.0.0.tar.gz
Size 12.5 kB
Tags Source
SHA-256 checksum
How to use checksums
7db181e421fd98d1682a8cc9c9e221138f364e8a792320941dd7ca522a1a80f4
BLAKE2b-256 checksum
How to use checksums
06d9b1d0db13be39e0e498679532bf7395c947c1e6ac8019a2662fd9e49ec6a9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

2.0.0 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page