Skip to main content

tokmon

PyPI version Python License: MIT Tests

Know exactly what your LLM calls cost. One decorator. Zero config. Zero dependencies.

pip install tokmon-ai

🎬 Demo

tokmon demo — track LLM token costs per request
Tracking token usage and costs across an agent pipeline — decorator, budget alerts, session dashboard


The Problem

You're building AI agents and you have no idea what they cost per request. Is it $0.01 or $0.50? Which tool call is the expensive one? You'll find out at the end of the month when the invoice arrives.

The Solution

import tokmon

@tokmon.track("search-agent")
def search_and_summarize(query: str) -> str:
    results = llm("Search for: " + query)         # tracked
    summary = llm("Summarize: " + results)         # tracked
    return summary

result = search_and_summarize("latest AI news")

# After the call:
print(tokmon.last_report())
# ┌─────────────────────────────────────────────────┐
# │ search-agent                                     │
# │ Calls: 2 | Tokens: 1,847 | Cost: $0.0042        │
# │ ├─ Call 1: 823 tok ($0.0018) — gpt-4o-mini      │
# │ └─ Call 2: 1024 tok ($0.0024) — gpt-4o-mini     │
# └─────────────────────────────────────────────────┘

That's it. One line added, full visibility.

Features

🎯 Drop-in Decorator

@tokmon.track("my-feature")
def any_function():
    # All LLM calls inside are automatically tracked
    ...

💰 Budget Alerts

@tokmon.budget("expensive-agent", max_usd=1.00)
def expensive_agent(query):
    # Raises tokmon.BudgetExceeded if cost exceeds $1.00
    ...

# Or soft limit (warns but doesn't fail):
@tokmon.budget("agent", max_usd=0.50, hard=False)
def agent(query):
    ...

📊 Session Tracking

# Track costs across an entire session
with tokmon.session("user-123") as s:
    agent.run("question 1")
    agent.run("question 2")
    agent.run("question 3")

print(s.total_cost_usd)     # $0.047
print(s.total_tokens)       # 12,483
print(s.call_count)         # 9
print(s.cost_per_call_usd)  # $0.0052

📈 Export & Reporting

# JSON export for dashboards
tokmon.export_json("costs.json")

# CSV for spreadsheets
tokmon.export_csv("costs.csv")

# Print summary table
tokmon.print_report()
# ┌──────────────────┬───────┬──────────┬──────────┐
# │ Feature          │ Calls │ Tokens   │ Cost     │
# ├──────────────────┼───────┼──────────┼──────────┤
# │ search-agent     │ 142   │ 284,100  │ $0.89    │
# │ summarizer       │ 89    │ 156,200  │ $0.52    │
# │ classifier       │ 1,204 │ 120,400  │ $0.18    │
# └──────────────────┴───────┴──────────┴──────────┘

🖥️ CLI Dashboard

# Watch costs in real-time (requires: pip install tokmon[rich])
tokmon dashboard

# Show historical report
tokmon report --last 7d

# Set global budget alert
tokmon budget --daily 10.00 --alert slack

Supported Providers

Provider Auto-Patch Manual
OpenAI SDK
Anthropic SDK
LiteLLM
Any HTTP API

Auto-patching (zero code changes)

import tokmon
tokmon.auto_patch()  # Patches openai, anthropic, litellm automatically

# All subsequent LLM calls are tracked without any other changes

Manual recording

# If you use a custom client:
tokmon.record(
    feature="my-agent",
    model="gpt-4o",
    prompt_tokens=500,
    completion_tokens=200,
)

How It Works

┌─────────────────────────────────────────────────────┐
│                    Your Code                         │
│                                                     │
│   @tokmon.track("feature")                          │
│   def my_function():                                │
│       llm_call(...)  ←─── intercepted               │
│                                                     │
├─────────────────────────────────────────────────────┤
│                  tokmon Core                         │
│                                                     │
│   Interceptor → Counter → Store → Reporter          │
│       │              │        │         │           │
│   patches SDK    sums tokens  writes   formats      │
│                  + pricing    to log    output       │
└─────────────────────────────────────────────────────┘

Configuration

import tokmon

# Set custom pricing (override defaults)
tokmon.set_pricing("my-fine-tuned-model", prompt=5.00, completion=15.00)

# Set storage backend
tokmon.configure(storage="sqlite:///costs.db")  # or "memory", "json:costs.json"

# Set alert callback
tokmon.on_budget_exceeded(lambda report: slack.post(f"⚠️ {report}"))

Zero Dependencies

Core tokmon has zero dependencies. Optional extras:

  • tokmon[rich] — terminal dashboard with live updates
  • tokmon[openai] — auto-patches OpenAI SDK
  • tokmon[litellm] — auto-patches LiteLLM
  • tokmon[all] — everything

Contributing

git clone https://github.com/naveenkumarbaskaran/tokmon.git
cd tokmon
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest

License

MIT

Release files for tokmon-ai 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokmon-ai 2.0.0
File Size Uploaded
tokmon_ai-2.0.0.tar.gz 9.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokmon-ai 2.0.0
File Interpreter ABI Platform
tokmon_ai-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 22.7 kB

Release files / tokmon_ai-2.0.0.tar.gz

Download URL tokmon_ai-2.0.0.tar.gz
Size 9.8 kB
Tags Source
SHA-256 checksum
How to use checksums
7a54971fa12c269cbb17e46bdc7d7cad420f2a72706910beec2043346aa4e5e2
BLAKE2b-256 checksum
How to use checksums
356cf81097da2f18579c1a965d220a5b57495e0f7b2979e009f8ebd1a16158ee
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release files / tokmon_ai-2.0.0-py3-none-any.whl

Download URL tokmon_ai-2.0.0-py3-none-any.whl
Size 13.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ca946e4789ffe87470b2564bce687122ec929e1138fd875c2015600e7a52dbe
BLAKE2b-256 checksum
How to use checksums
fdde5ed55731c01b09f90c312bcc63261eb6d057ad78e4debbf43a110c54f0c1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page