tokenfin — Python SDK
Track LLM token usage and cost in TokenFin.
The sync client has no dependencies. anthropic, openai and aiohttp are all optional.
pip install tokenfin
pip install "tokenfin[async]" # optional aiohttp transport for AsyncTokenFinClient
Auto-instrument Anthropic / OpenAI
from anthropic import Anthropic
from openai import OpenAI
from tokenfin import TokenFinClient, wrap_anthropic, wrap_openai
tf = TokenFinClient(api_key="tfk_...")
anthropic = wrap_anthropic(Anthropic(), tf, user_email="dev@acme.com", tags={"feature": "chat"})
openai = wrap_openai(OpenAI(), tf)
Every messages.create / messages.stream / chat.completions.create / responses.create
call is now recorded. This works for sync and async clients, with and without streaming.
The wrappers record the model, input and output tokens, cache read/write tokens, and latency.
- The wrappers do not change return values, and SDK errors propagate unchanged.
- Prompt capture is off by default. Turn it on with
capture_prompts=True. - For OpenAI streaming,
stream_options.include_usageis set for you and the usage-only chunk is hidden.
Manual tracking
tf.track("claude-sonnet-4-6", 1200, 380,
cache_read_tokens=9000, cache_write_tokens=400,
project_id="…uuid…", user_email="dev@acme.com", session_id="run-42",
latency_ms=812, idempotency_key="req_123", tags={"feature": "chat"})
tf.shutdown() # drain before exit (also runs from atexit; no signal handlers are installed)
flush() and shutdown() return FlushResult(sent, dropped), and tf.stats() returns
lifetime counters. A non-retryable 4xx counts as dropped. The SDK retries 408, 429, 5xx
and network errors, and honours Retry-After.
Configuration
| Parameter | Default | Description |
|---|---|---|
api_key |
required | tfk_… key with the ingest or write scope |
base_url |
https://tokenfin.curiousdevs.com |
self-hosted URL |
batch_size |
100 |
events per request (server max 500) |
flush_interval |
1.0 |
seconds; 0 = manual flush only |
max_queue_size |
10000 |
when full, the oldest event is dropped and counted |
max_retries |
3 |
retries for 408 / 429 / 5xx / network errors |
max_retry_after |
30.0 |
cap (seconds) on honouring Retry-After |
timeout |
5.0 |
per-request timeout (seconds) |
flush_on_exit |
True |
drain the queue from atexit |
debug |
False |
debug logging on the tokenfin logger |
policy_ttl |
60.0 |
seconds between refreshes of the org policy (model routes / blocks) |
Model routes and limits
wrap_anthropic / wrap_openai read the org policy (GET /api/v1/policy, same key) on a
background thread and fail open. Active model routes — including automatic switches when a
per-model limit is reached — rewrite model before the call and record metadata.routed_from.
Pass enforce_policy=True to raise TokenFinPolicyError for models a limit blocked;
route_models=False disables routing; policy_wait_ms (default 200) caps how long the first
call waits for the first fetch (async clients wait without blocking the event loop).
from tokenfin import TokenFinPolicyError
client = wrap_anthropic(Anthropic(), tf, enforce_policy=True)
try:
client.messages.create(model="claude-opus-4-8", max_tokens=512, messages=msgs)
except TokenFinPolicyError as e:
... # e.model is blocked for this org
See the full SDK guide for everything else.
License
MIT
Metadata
Release files for tokenfin 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokenfin-0.2.1.tar.gz | 24.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokenfin-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 47.0 kB
Release files / tokenfin-0.2.1.tar.gz
| Download URL | tokenfin-0.2.1.tar.gz |
|---|---|
| Size | 24.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
88cd7691971c32ccd30fdbc6b9c62a2c7abfdb8107797233f59731e8f8cf7ecd
|
|
BLAKE2b-256 checksum How to use checksums |
997b2c1aa14fd28a996341ebde40ea095caa700786d1e51135924fb230badd2a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tokenfin-0.2.1-py3-none-any.whl
| Download URL | tokenfin-0.2.1-py3-none-any.whl |
|---|---|
| Size | 22.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e8c20f126f93e92ec4a7ef91eb606654530f5439f77b77ca4c15ab1bcee7d88f
|
|
BLAKE2b-256 checksum How to use checksums |
8a34414e4a2abe9c7bd8918a06fb1af9cc58aacb8b7730f6bd5836e58ccf2e22
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|