Zero-config AI usage tracking — wrap any OpenAI/Anthropic/Gemini client and log to your TokenGauge dashboard
Project description
TokenGauge SDK
Zero-config AI usage tracking + model recommendations. Wrap your existing OpenAI, Anthropic, or Google Gemini client with one line — every call is automatically logged to your TokenGauge dashboard. Or ask the SDK which model to use before you even make a call.
Your API keys stay with you. The SDK only reads token counts from API responses and sends them to TokenGauge. Nothing is proxied.
Install
pip install tokengauge
Quick start
-
Sign up at tokengauge.onrender.com and copy your SDK token from Settings.
-
Wrap your client:
from tokengauge import TokenGauge
import openai
tw = TokenGauge(token="your-sdk-token")
client = tw.wrap(openai.OpenAI(api_key="sk-..."))
# Use exactly as before — usage appears on your dashboard automatically
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Anthropic
from tokengauge import TokenGauge
import anthropic
tw = TokenGauge(token="your-sdk-token")
client = tw.wrap(anthropic.Anthropic(api_key="sk-ant-..."))
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=256,
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.content[0].text)
Google Gemini
from tokengauge import TokenGauge
from google import genai
tw = TokenGauge(token="your-sdk-token")
client = tw.wrap(genai.Client(api_key="your-gemini-key"))
response = client.models.generate_content(
model="gemini-1.5-flash",
contents="Hello!",
)
print(response.text)
Async clients
from tokengauge import TokenGauge
import openai, asyncio
tw = TokenGauge(token="your-sdk-token")
client = tw.wrap(openai.AsyncOpenAI(api_key="sk-..."))
async def main():
response = await client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
asyncio.run(main())
Model recommendations
Not sure which model to use? recommend_model() classifies your prompt locally, estimates the cost, and scores every model by success probability — no API call required.
tw = TokenGauge(token="your-sdk-token")
rec = tw.recommend_model(
messages=[{"role": "user", "content": "Refactor this Python class to use dataclasses..."}],
provider="anthropic", # optional: also return best within this provider
budget_usd=0.05, # optional: exclude models above this cost estimate
)
print(rec["prompt_type"]) # "code"
print(rec["complexity"]) # 1–10 score
print(rec["best_overall"]["model"]) # e.g. "claude-opus-4.6"
print(rec["best_overall"]["success_probability"]) # e.g. 1.0
print(rec["best_overall"]["estimated_cost_usd"]) # e.g. 0.00047
print(rec["within_provider"]["model"]) # best Anthropic model
What's returned
{
"prompt_type": "code", # classified category
"complexity": 4, # estimated complexity 1–10
"estimated_tokens_in": 312, # rough token estimate
"best_overall": { # best model across all providers
"model": "claude-opus-4.6",
"provider": "anthropic",
"quality_score": 10,
"estimated_tokens_in": 312,
"estimated_tokens_out": 124,
"estimated_cost_usd": 0.00469,
"success_probability": 1.0
},
"within_provider": { ... } # best model within your preferred provider
}
How it works
| Step | What happens |
|---|---|
| Classify | Prompt is matched against 8 categories: code, chat, summarization, analysis, creative, extraction, translation, other |
| Estimate | Token count estimated from text length (~4 chars/token); output estimated at 40% of input |
| Score | Each model gets a success probability based on its type-specific quality score minus a penalty for complexity above its ceiling |
| Rank | Models sorted by success probability, then cheapest on tie; budget filter applied if set |
Prompt text is never sent anywhere — classification runs entirely on your machine.
Supported models in the registry
| Provider | Models |
|---|---|
| OpenAI | gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini, o3, o3-mini, o4-mini |
| Anthropic | claude-opus-4.6, claude-sonnet-4.6, claude-haiku-4-5, claude-3-7-sonnet, claude-3-5-sonnet, claude-3-5-haiku |
| gemini-2.5-pro, gemini-2.5-flash, gemini-2.0-flash, gemini-2.0-flash-lite |
Tag calls by feature
summarizer = tw.wrap(openai.OpenAI(api_key="sk-..."), app_tag="summarizer")
chatbot = tw.wrap(openai.OpenAI(api_key="sk-..."), app_tag="chatbot")
Login instead of pasting a token
tw = TokenGauge.login(email="you@example.com", password="your-password")
What gets tracked
| Field | Description |
|---|---|
| Provider | openai / anthropic / google |
| Model | e.g. gpt-4o-mini, claude-3-5-sonnet |
| Tokens in | Prompt token count |
| Tokens out | Completion token count |
| Cost (USD) | Calculated from current model pricing |
| Latency | End-to-end request time in ms |
| Prompt type | Auto-classified category (code, chat, analysis, etc.) |
| Complexity | Estimated complexity score 1–10 |
| App tag | Optional label you set per-client |
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tokengauge-0.3.2.tar.gz.
File metadata
- Download URL: tokengauge-0.3.2.tar.gz
- Upload date:
- Size: 15.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9231e703176137e689d2482cde7aa78ff729dc1384186dcf97ee9af842b561a0
|
|
| MD5 |
c9b34ff5f3770e92dd5a755057eef50e
|
|
| BLAKE2b-256 |
400d5b0b8969073d3e35a5a96c7bd87b638a65b2409304f5b9d653a9906249f8
|
File details
Details for the file tokengauge-0.3.2-py3-none-any.whl.
File metadata
- Download URL: tokengauge-0.3.2-py3-none-any.whl
- Upload date:
- Size: 11.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5dbe6fe0bf906354938d98d03d675f99e563fe93b86bce655e0ea13549cb69b5
|
|
| MD5 |
0ae3891b70cb142968f2f82999a36f38
|
|
| BLAKE2b-256 |
a1272b76cc7c634f97bee10b7395bd308ce33300cf772b875f5a75333f3e28b7
|