Skip to main content

tokeven

Per prompt cost visibility for the Claude, OpenAI, and Google SDKs, without Tokeven ever holding a provider key or seeing a prompt.

tokeven wraps the official provider SDKs. Your call goes to the provider exactly as before, and the wrapper reads the usage block off the response and pushes a small metrics event to your Tokeven dashboard. It also answers the question that saves the most money, which is what a call is about to cost before you send it.

Thirty seconds, no install, no account

uvx tokeven teardown

Reads the Claude Code session logs already on your machine, prices them against live rates, and prints what you have actually been spending. It makes zero network calls and asks for nothing. Paste the output at https://tokeven.com/teardown for the breakdown by model, or keep it local.

What it does

  • Estimates cost and confidence across all three providers before you send a call, so model choice is an informed decision rather than a default.
  • Records what each call actually cost, with cache aware pricing recomputed server side.
  • Runs as an MCP server, so Claude Code can ask for an estimate directly.
  • Ships a live cost widget for your shell prompt or status line, and a self hosted sidecar for capturing usage without an SDK wrapper.
  • Suggests cheaper models and prompt level efficiency improvements, and always leaves the decision to you. Nothing is rerouted silently.

Privacy

  • Tokeven receives token counts and metadata such as model, latency, and a character count computed locally. Prompt and completion content never reaches Tokeven through this pipeline.
  • Every outgoing event passes a strict key allowlist, so content cannot be transmitted by accident.
  • Authentication to Tokeven uses a write only ingest token, which cannot read your data or manage your account. It is not a provider key, and Tokeven never asks for one.
  • Two features are opt in exceptions that do send prompt text you explicitly hand them, because rewriting a prompt requires reading it: the prompt revision tool on the website and advisor.improve(). Both are disclosed at the point of use. Details at https://tokeven.com/privacy
  • If any part of the instrumentation fails, your provider call is returned unchanged. Tokeven is designed not to break your application.

Install

pip install tokeven                  # wrapper, advisor, CLI, widget
pip install "tokeven[anthropic]"     # with the Anthropic SDK
pip install "tokeven[openai]"        # with the OpenAI SDK
pip install "tokeven[google-genai]"  # with the Google SDK
pip install "tokeven[xai]"           # for xAI / Grok (uses the OpenAI SDK)
pip install "tokeven[mcp]"           # with the MCP server
pip install "tokeven[proxy]"         # with the self hosted sidecar
pip install "tokeven[all]"           # everything

For the command line tools on their own, without touching a project environment:

pipx install tokeven

Quickstart

Wrap your existing provider client. Configuration is read from the environment: TOKEVEN_INGEST_TOKEN, TOKEVEN_BASE_URL, TOKEVEN_PROJECT, TOKEVEN_MODE.

from anthropic import Anthropic
from tokeven import TokevenAnthropic

client = TokevenAnthropic(Anthropic(), project="billing-agent")

resp = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.content)   # unchanged Anthropic response

client.flush()   # optional: force-push buffered events before exit

TokevenOpenAI and TokevenGemini wrap the OpenAI and Google SDKs the same way.

For xAI / Grok, use TokevenXAI. Grok is served over the OpenAI SDK against api.x.ai, so there is no separate xAI package to install:

from tokeven import TokevenXAI

client = TokevenXAI(project="research-agent")   # reads XAI_API_KEY

resp = client.chat.completions.create(
    model="grok-4.6",
    messages=[{"role": "user", "content": "Hello"}],
)

wrap() recognises Grok too, including a client you built yourself:

from openai import OpenAI
from tokeven import wrap

client = wrap(OpenAI(base_url="https://api.x.ai/v1", api_key="..."))

Two Grok pricing quirks worth knowing, both handled for you: cached input is priced per model rather than as one family-wide discount, and every Grok model charges double — input and output — once a prompt reaches 200,000 tokens.

First run

Set your write only ingest token:

export TOKEVEN_INGEST_TOKEN="tkv_ing_..."
export TOKEVEN_BASE_URL="https://api.tokeven.com"

Create a token at https://tokeven.com/settings

Estimate the cost of a call before you send it:

tokeven estimate --model claude-opus-4-8 "your prompt here"

The estimate runs entirely on your machine and needs no account at all.

Use it from Claude Code (MCP)

The advisor runs as an MCP server, so Claude can price a call before it makes one. It is listed in the MCP registry as com.tokeven/tokeven, or add it by hand:

claude mcp add tokeven -- uvx --with 'mcp>=2.0,<3.0' tokeven mcp

tokeven mcp and tokeven-mcp are the same server; the subcommand form is what registry installs launch.

The MCP server estimates and advises. It does not track: nothing it does emits a usage event, because wrap() and the sidecar are the only capture paths. tokeven status reports which of those are actually live on this machine.

Is it actually tracking? (tokeven status)

Instrumentation that is configured but inert looks exactly like instrumentation that works, so there is one command that says which it is:

tokeven status

It reports which capture paths are live (and that the MCP server is an estimator, not a tracker), the resolved base_url and where that value came from, the ingest token prefix and its source, whether base_url is reachable, how many events are sitting undelivered in the local durable queue, and a plain verdict on whether anything is being tracked right now. It needs no login, works offline (--offline skips the reachability probes), never prints your full token, and exits non zero when it finds problems so you can use it in a script. --json gives the same report machine readably.

The sidecar (optional)

If you would rather not wrap an SDK, run the self hosted sidecar and point your provider base URL at it. It forwards every request to the provider verbatim and captures usage on a fail open side channel. Provider keys and prompt content never reach Tokeven.

pip install "tokeven[proxy]"
tokeven-proxy

Point each SDK at the proxy in place of its provider:

Provider Base URL
Anthropic http://localhost:7779
OpenAI http://localhost:7779/v1
Google http://localhost:7779
xAI / Grok http://localhost:7779/xai/v1

Grok is namespaced because xAI serves OpenAI's exact paths, so the proxy has no other way to tell the two apart. It costs nothing: reaching api.x.ai already requires a custom base URL. The /xai segment is stripped before forwarding, so xAI sees the request your SDK built.

A container image and an example docker-compose.yml ship alongside this package.

The live cost widget (optional)

A one line spend gauge for your shell prompt, tmux, or editor status line. It reuses your dashboard login, never an ingest token or a provider key.

tokeven-widget            # print one status line and exit
tokeven-widget watch      # poll and reprint on an interval

What is and is not tracked

Tokeven sees the traffic you instrument. Calls made through the provider consoles, through the Anthropic Workbench, or from applications you have not wrapped will not appear on the dashboard.

Free and paid

Everything this package does on your machine is free, and there is no trial clock on it. That includes pre send cost estimates, the model picker, efficiency suggestions, the MCP server, the sidecar, and the live cost widget. A free Tokeven account adds a dashboard covering 1,000 tracked calls per month.

The paid tier is priced as gainshare. Tokeven earns 10 percent of verified savings, measured per call against live token prices rather than a frozen baseline, and you keep the other 90 percent. If you save nothing, you pay nothing, and the fee falls automatically when model prices fall. There is no per seat charge.

Paid adds:

  • Unlimited tracked calls
  • Prompt revision suggestions
  • Organization wide savings dashboard and per person savings tracking
  • Analytics across all three providers
  • Embeddable widgets
  • n8n and webhook integrations
  • CSV, JSON, and API export

Enterprise terms, including capped arrangements for teams that budget fixed line items, SSO, and custom retention, are handled case by case.

If you are weighing the paid tier or want a hand sizing what you would actually save, email team@tokeven.com. We are happy to look at your numbers with you, and there is no obligation attached to asking. Full pricing detail is at https://tokeven.com/pricing

License

Proprietary. The full terms ship in the LICENSE file bundled with this package. Questions: team@tokeven.com

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tokeven-0.3.0.tar.gz (205.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tokeven-0.3.0-py3-none-any.whl (154.3 kB view details)

Uploaded Python 3

File details

Details for the file tokeven-0.3.0.tar.gz.

File metadata

  • Download URL: tokeven-0.3.0.tar.gz
  • Upload date:
  • Size: 205.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tokeven-0.3.0.tar.gz
Algorithm Hash digest
SHA256 03d70e70ceadab35f6758ac54edae53f6783519f8c0ef495ea94101f74ce3258
MD5 7df9c2145c160fdf1ed6a2bfe8e03592
BLAKE2b-256 f81f9026fdb9c663fdd765fa93e76e600a0f9cc71f350af874edc49ba72ae425

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokeven-0.3.0.tar.gz:

Publisher: release-pypi.yml on harrhall/tokeven

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tokeven-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: tokeven-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 154.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tokeven-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5c962a32fa2bd7919cd7fa90adcf8f67420c70854b539f30830fb185ebf9f725
MD5 b17c5a17b4b07be483ea9dd16dc4b64b
BLAKE2b-256 9d24a292e6f27f5636abb1bb880d071894b8a5c63c705daaa72095cd90db84df

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokeven-0.3.0-py3-none-any.whl:

Publisher: release-pypi.yml on harrhall/tokeven

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.4.1

2 files

0.4.0

2 files

This release

0.3.0 This release

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page