litellm-cost
Standalone LLM cost estimation: model pricing data + token cost math, with zero litellm dependency.
Why this package exists
LiteLLM is a 100+ provider LLM gateway — but its cost
estimation logic is trapped inside a 9,700-line god-module (litellm/utils.py) and paid for
with a heavy dependency bill (pydantic, httpx, openai, jinja2, …). Two design choices make it
painful to reuse in isolation:
- Import-time network fetch —
import litellmpulls the pricing JSON from GitHub by default, which fails or blocks in offline / air-gapped environments and adds a supply-chain surface. - Dependency bloat — the cost path only needs a data dictionary and arithmetic, yet
installing
litellmdrags in the full gateway stack.
litellm-cost extracts exactly the reusable core — the pricing data dictionary and the
token-to-cost math — into a zero-dependency package that:
- never touches the network on import (or at all, unless you explicitly ask it to refresh),
- exposes a small, typed, documented API,
- ships an embedded pricing snapshot with its own
data_versionfor traceability, - keeps the upstream O(1) case-insensitive model lookup and CJK-aware token heuristics.
Installation
pip install litellm-cost
Optional extras:
pip install "litellm-cost[tiktoken]" # tiktoken accelerator for token estimation
pip install "litellm-cost[dev]" # pytest + pytest-cov for development
Python 3.9 – 3.12 supported. No third-party dependencies in the core package.
Quick start
Estimate the cost of a call
from litellm_cost import get_cost
# gpt-4o: $2.5e-06 / input token, $1e-05 / output token
cost = get_cost("gpt-4o", input_tokens=1250, output_tokens=350)
print(f"${cost:.6f}") # $0.006625
# provider-prefixed model keys work as in LiteLLM
cost = get_cost("groq/llama-3.1-8b-instant", input_tokens=1000, output_tokens=500)
# cached tokens are billed at the model's cache-read price
# (defaults to 50% of the input price when the data has no explicit value)
cost = get_cost("gpt-4o", input_tokens=1000, cache_read_tokens=900)
# or pass an OpenAI-style usage block directly
cost = get_cost("gpt-4o", usage={"prompt_tokens": 100, "completion_tokens": 50})
Estimate tokens (no tokenizer dependency)
from litellm_cost import estimate_tokens
estimate_tokens("Hello, world!") # heuristic, CJK-aware
estimate_tokens("你好世界,这是一个测试。") # CJK text handled correctly
estimate_tokens([{"role": "user", "content": "Hi"}]) # OpenAI-style messages
# opt-in tiktoken accelerator (silently falls back to the heuristic
# when tiktoken is not installed)
estimate_tokens("Hello, world!", use_tiktoken=True, model="gpt-4o")
Query pricing data
from litellm_cost import get_pricing, list_providers, data_version
info = get_pricing("gpt-4o")
print(info["input_cost_per_token"]) # 2.5e-06
print(info["max_input_tokens"]) # 128000
providers = list_providers() # full 100+ provider roster
print(data_version()) # e.g. "2025.06.17.1"
Handle errors explicitly
from litellm_cost import LitellmCostError, UnknownModel, MissingPricing, ContextWindowExceededError
try:
get_cost("not-a-real-model", input_tokens=10)
except UnknownModel as exc:
print(exc) # unknown model: 'not-a-real-model'. It is not present in the embedded pricing data...
try:
get_cost("gpt-4o", input_tokens=10**9, check_context_window=True)
except ContextWindowExceededError:
pass # opt-in: off by default, cost estimation should not fail on hypothetical usage
Keep pricing data fresh
The embedded snapshot is versioned and traceable. To update it from LiteLLM's upstream
model_prices_and_context_window.json (stdlib only, no third-party dependencies):
python -m litellm_cost.refresh --check # dry run: report without writing
python -m litellm_cost.refresh # fetch, merge, bump data_version
Or programmatically:
from litellm_cost.refresh import refresh_pricing_data
summary = refresh_pricing_data()
print(summary) # {"data_version": ..., "models": ..., "providers": ..., "written": True}
The refresh pipeline inherits LiteLLM's own supply-chain guards: the fetched payload must be a non-empty JSON object, and the refresh is always an explicit, opt-in action — never an import side effect.
API overview
| Function | Purpose |
|---|---|
get_cost(model, ...) |
Compute the USD cost of a call from token counts |
estimate_tokens(source, ...) |
Estimate tokens for text or OpenAI-style messages |
get_pricing(model) |
Query the pricing-info dict for a model (defensive copy) |
list_providers() |
Sorted list of supported providers (100+) |
cost_per_token(model) |
(input_cost_per_token, output_cost_per_token) tuple |
data_version() |
Version string of the embedded pricing data |
Full signatures, parameter tables, error semantics and examples: docs/API.md.
Relationship to LiteLLM
This package is a decomposition of the cost-estimation logic found in LiteLLM
(litellm.cost_calculator, litellm.utils, litellm_core_utils.token_counter, and the
model_prices_and_context_window.json data dictionary), stripped of runtime dependencies
(network I/O, provider routing, logging, telemetry). The extraction boundary review covers
LiteLLM v1.85.1; the package is versioned independently of upstream.
What is not copied: the router, the proxy server, provider request adapters, the logging/telemetry stack, streaming processors, and the import-time network fetch.
License
MIT — see LICENSE. The LICENSE retains the upstream attribution notice: the pricing data schema and token-to-cost math are derived from LiteLLM (MIT), and the embedded pricing JSON is upstream community-maintained data with its source URL recorded in the bundle.
Development
git clone https://github.com/akn-runtime/litellm-cost
cd litellm-cost
pip install -e ".[dev]"
pytest
To rebuild the distribution artifacts:
pip install build
python -m build # produces sdist + wheel in dist/
Metadata
Release files for litellm-cost 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| litellm_cost-0.1.0.tar.gz | 91.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| litellm_cost-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 108.3 kB
Release files / litellm_cost-0.1.0.tar.gz
| Download URL | litellm_cost-0.1.0.tar.gz |
|---|---|
| Size | 91.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
19ba0602373d76c9b2e7e6734e12d9b7b01efc365e358341cce37af10167c7be
|
|
BLAKE2b-256 checksum How to use checksums |
2da0c4524bc470e585ba564896fda106515c74c0f789a304138a66263b9140d6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / litellm_cost-0.1.0-py3-none-any.whl
| Download URL | litellm_cost-0.1.0-py3-none-any.whl |
|---|---|
| Size | 17.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b8f06a1abaf31d9a05f953b7945e92d3d813b35ef7dfbf01f470837de152fde5
|
|
BLAKE2b-256 checksum How to use checksums |
f0eebe9a342ca4a517a8a47d4f264d8cad1d553b2713047df3bdb390fe647ac4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|