Skip to main content

litellm-cost

Standalone LLM cost estimation: model pricing data + token cost math, with zero litellm dependency.

PyPI - Version PyPI - Python Version License: MIT Tests


Why this package exists

LiteLLM is a 100+ provider LLM gateway — but its cost estimation logic is trapped inside a 9,700-line god-module (litellm/utils.py) and paid for with a heavy dependency bill (pydantic, httpx, openai, jinja2, …). Two design choices make it painful to reuse in isolation:

  1. Import-time network fetch — import litellm pulls the pricing JSON from GitHub by default, which fails or blocks in offline / air-gapped environments and adds a supply-chain surface.
  2. Dependency bloat — the cost path only needs a data dictionary and arithmetic, yet installing litellm drags in the full gateway stack.

litellm-cost extracts exactly the reusable core — the pricing data dictionary and the token-to-cost math — into a zero-dependency package that:

  • never touches the network on import (or at all, unless you explicitly ask it to refresh),
  • exposes a small, typed, documented API,
  • ships an embedded pricing snapshot with its own data_version for traceability,
  • keeps the upstream O(1) case-insensitive model lookup and CJK-aware token heuristics.

Installation

pip install litellm-cost

Optional extras:

pip install "litellm-cost[tiktoken]"   # tiktoken accelerator for token estimation
pip install "litellm-cost[dev]"        # pytest + pytest-cov for development

Python 3.9 – 3.12 supported. No third-party dependencies in the core package.

Quick start

Estimate the cost of a call

from litellm_cost import get_cost

# gpt-4o: $2.5e-06 / input token, $1e-05 / output token
cost = get_cost("gpt-4o", input_tokens=1250, output_tokens=350)
print(f"${cost:.6f}")   # $0.006625

# provider-prefixed model keys work as in LiteLLM
cost = get_cost("groq/llama-3.1-8b-instant", input_tokens=1000, output_tokens=500)

# cached tokens are billed at the model's cache-read price
# (defaults to 50% of the input price when the data has no explicit value)
cost = get_cost("gpt-4o", input_tokens=1000, cache_read_tokens=900)

# or pass an OpenAI-style usage block directly
cost = get_cost("gpt-4o", usage={"prompt_tokens": 100, "completion_tokens": 50})

Estimate tokens (no tokenizer dependency)

from litellm_cost import estimate_tokens

estimate_tokens("Hello, world!")                        # heuristic, CJK-aware
estimate_tokens("你好世界,这是一个测试。")                # CJK text handled correctly
estimate_tokens([{"role": "user", "content": "Hi"}])    # OpenAI-style messages

# opt-in tiktoken accelerator (silently falls back to the heuristic
# when tiktoken is not installed)
estimate_tokens("Hello, world!", use_tiktoken=True, model="gpt-4o")

Query pricing data

from litellm_cost import get_pricing, list_providers, data_version

info = get_pricing("gpt-4o")
print(info["input_cost_per_token"])    # 2.5e-06
print(info["max_input_tokens"])        # 128000

providers = list_providers()           # full 100+ provider roster
print(data_version())                  # e.g. "2025.06.17.1"

Handle errors explicitly

from litellm_cost import LitellmCostError, UnknownModel, MissingPricing, ContextWindowExceededError

try:
    get_cost("not-a-real-model", input_tokens=10)
except UnknownModel as exc:
    print(exc)   # unknown model: 'not-a-real-model'. It is not present in the embedded pricing data...

try:
    get_cost("gpt-4o", input_tokens=10**9, check_context_window=True)
except ContextWindowExceededError:
    pass  # opt-in: off by default, cost estimation should not fail on hypothetical usage

Keep pricing data fresh

The embedded snapshot is versioned and traceable. To update it from LiteLLM's upstream model_prices_and_context_window.json (stdlib only, no third-party dependencies):

python -m litellm_cost.refresh --check          # dry run: report without writing
python -m litellm_cost.refresh                  # fetch, merge, bump data_version

Or programmatically:

from litellm_cost.refresh import refresh_pricing_data

summary = refresh_pricing_data()
print(summary)   # {"data_version": ..., "models": ..., "providers": ..., "written": True}

The refresh pipeline inherits LiteLLM's own supply-chain guards: the fetched payload must be a non-empty JSON object, and the refresh is always an explicit, opt-in action — never an import side effect.

API overview

Function Purpose
get_cost(model, ...) Compute the USD cost of a call from token counts
estimate_tokens(source, ...) Estimate tokens for text or OpenAI-style messages
get_pricing(model) Query the pricing-info dict for a model (defensive copy)
list_providers() Sorted list of supported providers (100+)
cost_per_token(model) (input_cost_per_token, output_cost_per_token) tuple
data_version() Version string of the embedded pricing data

Full signatures, parameter tables, error semantics and examples: docs/API.md.

Relationship to LiteLLM

This package is a decomposition of the cost-estimation logic found in LiteLLM (litellm.cost_calculator, litellm.utils, litellm_core_utils.token_counter, and the model_prices_and_context_window.json data dictionary), stripped of runtime dependencies (network I/O, provider routing, logging, telemetry). The extraction boundary review covers LiteLLM v1.85.1; the package is versioned independently of upstream.

What is not copied: the router, the proxy server, provider request adapters, the logging/telemetry stack, streaming processors, and the import-time network fetch.

License

MIT — see LICENSE. The LICENSE retains the upstream attribution notice: the pricing data schema and token-to-cost math are derived from LiteLLM (MIT), and the embedded pricing JSON is upstream community-maintained data with its source URL recorded in the bundle.

Development

git clone https://github.com/akn-runtime/litellm-cost
cd litellm-cost
pip install -e ".[dev]"
pytest

To rebuild the distribution artifacts:

pip install build
python -m build          # produces sdist + wheel in dist/

Metadata

Release files for litellm-cost 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for litellm-cost 0.1.0
File Size Uploaded
litellm_cost-0.1.0.tar.gz 91.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for litellm-cost 0.1.0
File Interpreter ABI Platform
litellm_cost-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 108.3 kB

Release files / litellm_cost-0.1.0.tar.gz

Download URL litellm_cost-0.1.0.tar.gz
Size 91.3 kB
Tags Source
SHA-256 checksum
How to use checksums
19ba0602373d76c9b2e7e6734e12d9b7b01efc365e358341cce37af10167c7be
BLAKE2b-256 checksum
How to use checksums
2da0c4524bc470e585ba564896fda106515c74c0f789a304138a66263b9140d6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / litellm_cost-0.1.0-py3-none-any.whl

Download URL litellm_cost-0.1.0-py3-none-any.whl
Size 17.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b8f06a1abaf31d9a05f953b7945e92d3d813b35ef7dfbf01f470837de152fde5
BLAKE2b-256 checksum
How to use checksums
f0eebe9a342ca4a517a8a47d4f264d8cad1d553b2713047df3bdb390fe647ac4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page