Skip to main content

tokenfin — Python SDK

Track LLM token usage and cost in TokenFin. The sync client has no dependencies. anthropic, openai and aiohttp are all optional.

pip install tokenfin
pip install "tokenfin[async]"   # optional aiohttp transport for AsyncTokenFinClient

Auto-instrument Anthropic / OpenAI

from anthropic import Anthropic
from openai import OpenAI
from tokenfin import TokenFinClient, wrap_anthropic, wrap_openai

tf = TokenFinClient(api_key="tfk_...")
anthropic = wrap_anthropic(Anthropic(), tf, user_email="dev@acme.com", tags={"feature": "chat"})
openai = wrap_openai(OpenAI(), tf)

Every messages.create / messages.stream / chat.completions.create / responses.create call is now recorded. This works for sync and async clients, with and without streaming. The wrappers record the model, input and output tokens, cache read/write tokens, and latency.

  • The wrappers do not change return values, and SDK errors propagate unchanged.
  • Prompt capture is off by default. Turn it on with capture_prompts=True.
  • For OpenAI streaming, stream_options.include_usage is set for you and the usage-only chunk is hidden.

Manual tracking

tf.track("claude-sonnet-4-6", 1200, 380,
         cache_read_tokens=9000, cache_write_tokens=400,
         project_id="…uuid…", user_email="dev@acme.com", session_id="run-42",
         latency_ms=812, idempotency_key="req_123", tags={"feature": "chat"})
tf.shutdown()     # drain before exit (also runs from atexit; no signal handlers are installed)

flush() and shutdown() return FlushResult(sent, dropped), and tf.stats() returns lifetime counters. A non-retryable 4xx counts as dropped. The SDK retries 408, 429, 5xx and network errors, and honours Retry-After.

Configuration

Parameter Default Description
api_key required tfk_… key with the ingest or write scope
base_url https://tokenfin.curiousdevs.com self-hosted URL
batch_size 100 events per request (server max 500)
flush_interval 1.0 seconds; 0 = manual flush only
max_queue_size 10000 when full, the oldest event is dropped and counted
max_retries 3 retries for 408 / 429 / 5xx / network errors
max_retry_after 30.0 cap (seconds) on honouring Retry-After
timeout 5.0 per-request timeout (seconds)
flush_on_exit True drain the queue from atexit
debug False debug logging on the tokenfin logger
policy_ttl 60.0 seconds between refreshes of the org policy (model routes / blocks)

Model routes and limits

wrap_anthropic / wrap_openai read the org policy (GET /api/v1/policy, same key) on a background thread and fail open. Active model routes — including automatic switches when a per-model limit is reached — rewrite model before the call and record metadata.routed_from. Pass enforce_policy=True to raise TokenFinPolicyError for models a limit blocked; route_models=False disables routing; policy_wait_ms (default 200) caps how long the first call waits for the first fetch (async clients wait without blocking the event loop).

from tokenfin import TokenFinPolicyError
client = wrap_anthropic(Anthropic(), tf, enforce_policy=True)
try:
    client.messages.create(model="claude-opus-4-8", max_tokens=512, messages=msgs)
except TokenFinPolicyError as e:
    ...  # e.model is blocked for this org

See sdk/README.md for the full guide.

License

MIT

Metadata

Release files for tokenfin 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenfin 0.2.0
File Size Uploaded
tokenfin-0.2.0.tar.gz 24.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokenfin 0.2.0
File Interpreter ABI Platform
tokenfin-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 47.0 kB

Release files / tokenfin-0.2.0.tar.gz

Download URL tokenfin-0.2.0.tar.gz
Size 24.8 kB
Tags Source
SHA-256 checksum
How to use checksums
60b7412ac2b060a63445642021f1ea5cd593693ed751f682bd647a9801e3cebd
BLAKE2b-256 checksum
How to use checksums
ec1a0b97c369a277a6f4e5d98e853708e1dbb0cee665dcaa49d93b07dd21f828
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / tokenfin-0.2.0-py3-none-any.whl

Download URL tokenfin-0.2.0-py3-none-any.whl
Size 22.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5ea9e29ded2e18c33cb2763c194ef4232a3ab515d128c7b78ece2b9c92e6a59e
BLAKE2b-256 checksum
How to use checksums
d231bbe96cb7046a0a7f2b1a9ce5626d08af7c53240d42a332d2b158e98768c7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page