Skip to main content

tokenfin — Python SDK

Track LLM token usage and cost in TokenFin. The sync client has no dependencies. anthropic, openai and aiohttp are all optional.

pip install tokenfin
pip install "tokenfin[async]"   # optional aiohttp transport for AsyncTokenFinClient

Auto-instrument Anthropic / OpenAI

from anthropic import Anthropic
from openai import OpenAI
from tokenfin import TokenFinClient, wrap_anthropic, wrap_openai

tf = TokenFinClient(api_key="tfk_...")
anthropic = wrap_anthropic(Anthropic(), tf, user_email="dev@acme.com", tags={"feature": "chat"})
openai = wrap_openai(OpenAI(), tf)

Every messages.create / messages.stream / chat.completions.create / responses.create call is now recorded. This works for sync and async clients, with and without streaming. The wrappers record the model, input and output tokens, cache read/write tokens, and latency.

  • The wrappers do not change return values, and SDK errors propagate unchanged.
  • Prompt capture is off by default. Turn it on with capture_prompts=True.
  • For OpenAI streaming, stream_options.include_usage is set for you and the usage-only chunk is hidden.

Manual tracking

tf.track("claude-sonnet-4-6", 1200, 380,
         cache_read_tokens=9000, cache_write_tokens=400,
         project_id="…uuid…", user_email="dev@acme.com", session_id="run-42",
         latency_ms=812, idempotency_key="req_123", tags={"feature": "chat"})
tf.shutdown()     # drain before exit (also runs from atexit; no signal handlers are installed)

flush() and shutdown() return FlushResult(sent, dropped), and tf.stats() returns lifetime counters. A non-retryable 4xx counts as dropped. The SDK retries 408, 429, 5xx and network errors, and honours Retry-After.

Configuration

Parameter Default Description
api_key required tfk_… key with the ingest or write scope
base_url https://tokenfin.curiousdevs.com self-hosted URL
batch_size 100 events per request (server max 500)
flush_interval 1.0 seconds; 0 = manual flush only
max_queue_size 10000 when full, the oldest event is dropped and counted
max_retries 3 retries for 408 / 429 / 5xx / network errors
max_retry_after 30.0 cap (seconds) on honouring Retry-After
timeout 5.0 per-request timeout (seconds)
flush_on_exit True drain the queue from atexit
debug False debug logging on the tokenfin logger
policy_ttl 60.0 seconds between refreshes of the org policy (model routes / blocks)

Model routes and limits

wrap_anthropic / wrap_openai read the org policy (GET /api/v1/policy, same key) on a background thread and fail open. Active model routes — including automatic switches when a per-model limit is reached — rewrite model before the call and record metadata.routed_from. Pass enforce_policy=True to raise TokenFinPolicyError for models a limit blocked; route_models=False disables routing; policy_wait_ms (default 200) caps how long the first call waits for the first fetch (async clients wait without blocking the event loop).

from tokenfin import TokenFinPolicyError
client = wrap_anthropic(Anthropic(), tf, enforce_policy=True)
try:
    client.messages.create(model="claude-opus-4-8", max_tokens=512, messages=msgs)
except TokenFinPolicyError as e:
    ...  # e.model is blocked for this org

See the full SDK guide for everything else.

License

MIT

Metadata

Release files for tokenfin 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenfin 0.2.1
File Size Uploaded
tokenfin-0.2.1.tar.gz 24.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokenfin 0.2.1
File Interpreter ABI Platform
tokenfin-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 47.0 kB

Release files / tokenfin-0.2.1.tar.gz

Download URL tokenfin-0.2.1.tar.gz
Size 24.8 kB
Tags Source
SHA-256 checksum
How to use checksums
88cd7691971c32ccd30fdbc6b9c62a2c7abfdb8107797233f59731e8f8cf7ecd
BLAKE2b-256 checksum
How to use checksums
997b2c1aa14fd28a996341ebde40ea095caa700786d1e51135924fb230badd2a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / tokenfin-0.2.1-py3-none-any.whl

Download URL tokenfin-0.2.1-py3-none-any.whl
Size 22.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e8c20f126f93e92ec4a7ef91eb606654530f5439f77b77ca4c15ab1bcee7d88f
BLAKE2b-256 checksum
How to use checksums
8a34414e4a2abe9c7bd8918a06fb1af9cc58aacb8b7730f6bd5836e58ccf2e22
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page