Skip to main content

thriftllm — route every LLM call to the cheapest model that can actually do it

Most "cost-aware" routers rank models by input_rate + output_rate. That sum assumes a 1:1 input/output ratio, and almost no real workload has one. thriftllm computes the expected cost of this specific request — forecast output length, tiered rates, cache state, reasoning tokens — and ranks on $ per quality point.

from thriftllm.catalog.pricing import Catalog, expected_cost

catalog = Catalog.from_snapshot("prices.json")        # vendored, offline, SHA-pinned
model   = catalog.get("anthropic/claude-sonnet-5.5")  # unknown id raises; never $0

cost = expected_cost(
    model,
    input_tokens=20_000,            # a RAG call: long prompt, short answer
    expected_output_tokens=500,     # forecast, not a YAML constant
    cached_input_tokens=18_000,     # cache state enters the decision
)

print(cost.total_usd)      # 0.0141
print(cost.as_dict())      # every component separately auditable
pip install thriftllm      # one dependency: httpx

Status: pre-release. The pricing core and its tests are written. The router, the benchmark and the measured savings number are not. This README states no savings percentage because none has been measured yet — see PRE-REGISTRATION.md for the claim we have committed to making and the margin we committed to before running anything.


Why cost-aware LLM routing usually gets the cost wrong

Output tokens are typically 60–90% of an LLM bill, and the field does not forecast them. Three failures are common, and all three are in shipped code today:

Failure Consequence
Adding input_rate + output_rate Ranking is wrong for any workload that isn't 1:1
A constant expected_output_tokens in YAML The number that dominates the bill is a guess you typed
Unpriced model treated as $0 The model nobody priced always wins "cheapest"

thriftllm's position is narrow and checkable: be the project whose cost number is right. Not another selection algorithm — a correct denominator.

Does routing break prompt caching?

Usually yes, and this is the strongest argument against routers in general. Prompt caches are model-scoped. Cache reads cost 0.1× base input — and as little as 0.025× on some models — so switching models mid-session can forfeit a 75–97% discount to chase a smaller routing saving.

thriftllm treats this as a first-class cost term rather than a footnote:

  • Cache state is priced at decision time. The forfeited cache discount is subtracted from the candidate's expected saving before it is compared.
  • Minimum-cacheable thresholds are honoured. Below the per-model minimum, providers silently do not cache and charge full rate. thriftllm records that explicitly instead of quietly over-estimating the discount.
  • Session affinity keeps a conversation on its model unless the measured saving exceeds the measured cache loss.

When we publish a benchmark, the baseline runs with caching fully enabled and warmed, and the cache-hit rate of every arm is published next to its cost. A cost claim measured against a cache-disabled baseline is void, and we would rather say that ourselves than have it said in a comment thread.

When thriftllm will not help you

  • Long agentic sessions on one model with a warm cache. Keep the cache. Use cost tracking only.
  • Uniformly hard workloads. Published routers beat random routing by ~14% on knowledge-dense benchmarks while still needing the frontier model for over half of calls. Some traffic is simply not routable.
  • Single-shape workloads. If every request looks the same, choose the model once at build time. You do not need a router.
  • Workloads where output style matters. Switching models changes voice.

You should know this before you install it, not after.

How much does it actually save?

Not measured yet. When it is, the number will be reported as all-in cost (including this library's own overhead and every retried call), per stratum, against the best fixed single model — not against the most expensive model in the pool, which is how savings in this category are usually inflated.

We have pre-committed to publishing the result when routing loses.

Install

pip install thriftllm          # core: httpx only
pip install "thriftllm[server]"  # optional /v1/decision sidecar

Python 3.10+. Fully typed, py.typed shipped.

How it fits with your existing gateway

thriftllm does not reimplement the OpenAI wire format. Over half of the open issues on the largest gateway in this space are request-translation bugs; that is a maintenance burden with no upside for a routing project.

Instead it is a decision layer: a /v1/decision sidecar and plugins for gateways that already own the bytes. We win the decision; your gateway keeps the transport.

Supply chain

The dominant package in this category was compromised on PyPI, and that is a standing cost to everyone shipping here. thriftllm commits to:

  • One runtime dependency (httpx)
  • PyPI Trusted Publishing with PEP 740 attestations — no long-lived token
  • No install-time code execution
  • A vendored, SHA-pinned price snapshot — no network call at import

Prior art and credit

The price catalog is seeded from OpenRouter's public models API, which co-locates pricing with quality indices, and from models.dev. The candidate-filtering and scoring shape is informed by vllm-project/semantic-router (Apache-2.0), which is the strongest open implementation of multi-factor selection. RouteLLM defined the evaluation vocabulary this project is measured in.

License

Apache-2.0. Contributions under DCO (git commit -s). See LICENSE.

Metadata

Release files for thriftllm 0.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for thriftllm 0.0.0
File Size Uploaded
thriftllm-0.0.0.tar.gz 21.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for thriftllm 0.0.0
File Interpreter ABI Platform
thriftllm-0.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 34.6 kB

Release files / thriftllm-0.0.0.tar.gz

Download URL thriftllm-0.0.0.tar.gz
Size 21.9 kB
Tags Source
SHA-256 checksum
How to use checksums
f61c5c99cddd86ead9cb0c0f094c572c8df392f6e39b7ba6c8d605c978f8644f
BLAKE2b-256 checksum
How to use checksums
1d0c267833f56fd474ae90caf4300607200e32b872547a2c2e3066ab03948867
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.9

Release files / thriftllm-0.0.0-py3-none-any.whl

Download URL thriftllm-0.0.0-py3-none-any.whl
Size 12.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
079c6bfe3d3fa669d72425a2fdad49955679ecd303b79d2ed1f6e6acd195f835
BLAKE2b-256 checksum
How to use checksums
a849472dd2a95274d47b7d3aa7ce1fd9dee5463f647361b6335415753bd30e90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.9

Release history Release notifications | RSS feed

This release

0.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page