llm-cache-router
A lightweight, production-ready Python library that combines semantic caching, multi-provider LLM routing, and cost tracking in a single async-first API. Cut your LLM bill, ship faster, and never hardcode a single provider again.
Table of Contents
- Why llm-cache-router
- Comparison
- Features
- Installation
- Try in one command
- Interactive Playground (Colab)
- Quickstart
- Streaming
- Cache Warmup
- Routing Strategies
- Cache Backends
- Cache Invalidation, Versioning & Exact Match
- Multimodal Messages & Cache Keys
- Budget and Cost Tracking
- FastAPI Integration
- Async Context Manager
- Supported Providers
- Architecture
- Development
- Roadmap
- Contributing
- License
Why llm-cache-router
Calling LLMs directly is expensive, slow, and locks you into a single vendor. This library solves all three problems at once:
- Save money — a semantic cache returns answers for near-duplicate queries without re-calling the provider, typically cutting spend by 30–70% on production workloads.
- Stay resilient — swap providers on the fly, use fallback chains, and never take a full outage because one vendor is down.
- Control cost — built-in daily/monthly budget guardrails with Prometheus metrics for every request.
One dependency. Six providers. Three cache backends. Full async support.
Comparison
| llm-cache-router | LiteLLM | Raw provider SDK | |
|---|---|---|---|
| Semantic cache | First-class (memory / Redis / Qdrant) | Optional (redis-semantic, qdrant-semantic, …) |
No |
| Multiprovider | Built-in router + strategies | Yes (100+ providers / proxy) | Single vendor |
| Cost tracking | Built-in budget + savings + metrics | Yes (strong in proxy) | DIY |
| Async-first | Async API by design | Sync + acompletion |
Vendor-dependent |
When to choose:
- llm-cache-router — embeddable Python library: semantic cache, routing, and budget guardrails in one async API, no proxy required.
- LiteLLM — gateway/proxy with the widest provider coverage and ops features (rate limits, virtual keys, admin UI).
- Raw SDK — single vendor, full control; you build cache, routing, and cost tracking yourself.
Features
- Semantic cache — vector-similarity matching via
sentence-transformers, not just exact string hashing. Optionalexact_matchmode, key versioning and manual invalidation for correctness-sensitive workloads. - Multimodal-aware cache keys — images, audio, and video blocks are hashed into the query; cache is scoped per requested
model. - Multi-provider routing across OpenAI, Anthropic, Google Gemini, Ollama, MiniMax, Qwen (Dashscope) and any OpenAI-compatible endpoint (OpenRouter, vLLM, llama.cpp server, LiteLLM proxy, self-hosted inference).
- Three routing strategies:
CHEAPEST_FIRST,FASTEST_FIRST,FALLBACK_CHAIN. - Pluggable cache backends: in-memory (FAISS), Redis, Qdrant.
- Streaming — native async SSE streaming for every provider, transparent to the cache layer.
- Cost tracker with per-model pricing, daily/monthly budget limits and savings accounting.
- Cache warmup with controlled concurrency for pre-production pre-loading.
- FastAPI middleware + Prometheus metrics endpoint out of the box.
- Typed — Pydantic v2 models everywhere, fully typed public API.
- Tested — 15 test modules covering router, cache (incl. multimodal keys, model isolation, invalidation, key versioning), strategies, embeddings, providers, retry, warmup, and HTTP middleware.
Latest: v0.3.0 release notes — openai_compatible provider, cache invalidation, key versioning and exact-match mode.
Installation
pip install llm-cache-router
Optional extras:
pip install "llm-cache-router[redis]" # Redis cache backend
pip install "llm-cache-router[qdrant]" # Qdrant vector cache backend
pip install "llm-cache-router[fastapi]" # FastAPI middleware + Prometheus
pip install "llm-cache-router[all]" # everything above
pip install "llm-cache-router[dev]" # tests, ruff, mypy
Requires Python 3.11+.
Try in one command
Full demo stack: Redis + Qdrant + FastAPI with semantic cache.
cp .env.example .env # set OPENAI_API_KEY
docker compose up --build
curl -X POST http://localhost:8000/chat \
-H 'Content-Type: application/json' \
-d '{"message":"What is a semantic cache?"}'
Switch cache backend in .env: CACHE_BACKEND=redis (default) or CACHE_BACKEND=qdrant.
Details: examples/demo/README.md.
Interactive Playground (Colab)
Notebook: notebooks/playground.ipynb — install, two similar queries to see cache_hit, streaming, and router.stats() with the in-memory backend (no Redis/Qdrant required in Colab).
Quickstart
import asyncio
from llm_cache_router import CacheConfig, LLMRouter, RoutingStrategy
async def main() -> None:
router = LLMRouter(
providers={
"openai": {"api_key": "sk-...", "models": ["gpt-4o-mini"]},
"anthropic": {"api_key": "sk-ant-...", "models": ["claude-3-5-sonnet"]},
"gemini": {"api_key": "AIza...", "models": ["gemini-1.5-flash"]},
"ollama": {"base_url": "http://localhost:11434", "models": ["llama3.2"]},
},
cache=CacheConfig(
backend="memory",
threshold=0.92, # cosine similarity threshold
ttl=3600, # cache TTL in seconds
max_entries=10_000,
),
strategy=RoutingStrategy.CHEAPEST_FIRST,
budget={"daily_usd": 5.0, "monthly_usd": 50.0},
)
response = await router.complete(
messages=[{"role": "user", "content": "What is a semantic cache?"}],
model="gpt-4o-mini",
)
print(response.content)
print(f"cache_hit={response.cache_hit} cost=${response.cost_usd:.6f}")
asyncio.run(main())
Streaming
All providers (OpenAI, Anthropic, Gemini, Ollama, MiniMax, Qwen) support native SSE streaming. The cache layer is transparent: on a cache hit you receive a single final chunk, on a miss — a real streaming response that is also written to the cache once complete.
async for chunk in router.stream(
messages=[{"role": "user", "content": "Explain async/await in Python"}],
model="gpt-4o-mini",
):
print(chunk.delta, end="", flush=True)
if chunk.is_final:
print(f"\nprovider={chunk.provider_used} cost=${chunk.cost_usd:.6f}")
Cache Warmup
Pre-load the cache with known queries before traffic hits production:
from llm_cache_router.models import WarmupEntry
results = await router.warmup(
entries=[
WarmupEntry(
messages=[{"role": "user", "content": "What is RAG?"}],
model="gpt-4o-mini",
),
WarmupEntry(
messages=[{"role": "user", "content": "Explain vector databases"}],
model="gpt-4o-mini",
),
],
concurrency=5,
skip_cached=True,
)
print(results) # {"warmed": 2, "skipped": 0, "failed": 0}
Routing Strategies
| Strategy | Description |
|---|---|
CHEAPEST_FIRST |
Picks the cheapest provider/model by live pricing for each call. |
FASTEST_FIRST |
Picks the provider with the lowest observed latency (EMA). |
FALLBACK_CHAIN |
Tries providers in order, falls back on error/timeout. |
router = LLMRouter(
providers={
"openai": {"api_key": "sk-...", "models": ["gpt-4o"]},
"anthropic": {"api_key": "sk-ant-...", "models": ["claude-3-5-sonnet"]},
},
strategy=RoutingStrategy.FALLBACK_CHAIN,
fallback_chain=["openai/gpt-4o", "anthropic/claude-3-5-sonnet"],
)
Cache Backends
In-memory (FAISS)
Default. Zero dependencies beyond the core install. Best for single-process apps and tests.
cache=CacheConfig(backend="memory", threshold=0.92, ttl=3600, max_entries=10_000)
Redis
Production-grade distributed cache with LRU eviction, configurable timeouts, retry/backoff and bounded candidate set for vector search.
cache=CacheConfig(
backend="redis",
redis_url="redis://localhost:6379/0",
redis_namespace="llm_cache_router_prod",
threshold=0.92,
ttl=3600,
max_entries=50_000,
redis_command_timeout_sec=1.5,
redis_retry_attempts=3,
redis_retry_backoff_sec=0.2,
redis_candidate_k=256,
)
Qdrant
Native vector database for very large caches (millions of entries) and cross-service deployments.
pip install "llm-cache-router[qdrant]"
cache=CacheConfig(
backend="qdrant",
qdrant_url="http://localhost:6333",
qdrant_api_key=None, # optional for Qdrant Cloud
qdrant_collection="llm_cache",
threshold=0.92,
ttl=3600,
max_entries=100_000,
)
Cache Invalidation, Versioning & Exact Match
A semantic cache returns answers for similar queries — which also means it can serve stale or simply wrong answers. Three tools keep this under control:
Choosing a safe threshold
The default threshold=0.92 is deliberately strict, but cosine similarity cannot fully distinguish «how do I reset my password?» from «how do I reset my admin password?». If false positives are expensive in your domain:
- raise the threshold (0.95+), or
- enable
exact_match=True— the cache then only returns byte-identical queries (semantic search is disabled), or - treat the cache as an advisory layer and validate downstream.
Manual invalidation
# Drop cached entries for one model (e.g. after a prompt/model update)
removed = await router.invalidate_cache(model="gpt-4o-mini")
# Drop everything
removed = await router.invalidate_cache()
# Full clear (all models, all key versions)
await router.clear_cache()
Key versioning
Deployed a new system prompt? Bump key_version instead of flushing: old entries become invisible immediately and age out via TTL, while the new version starts with a clean slate.
from llm_cache_router.models import CacheConfig
cache = CacheConfig(key_version="2026-09-04-prompt-v2")
Works identically across memory, Redis and Qdrant backends.
Multimodal Messages & Cache Keys
Messages follow the OpenAI-compatible shape: content can be a string or a list of blocks (text, image_url, Anthropic image, audio, video). The router passes your requested model into the cache layer so different models never share a hit for the same text.
from llm_cache_router.models import Message, WarmupEntry
messages: list[Message] = [
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64,..."},
},
],
}
]
response = await router.complete(messages=messages, model="gpt-4o-mini")
# Second call with the same text + same image → cache_hit=True
# Same text but a different image → cache miss
# Same messages but model="gpt-4o" → cache miss (different model scope)
Warmup supports the same multimodal payloads:
WarmupEntry(
messages=messages,
model="gpt-4o-mini",
)
Binary media is stored in the cache key as a short sha256 fingerprint, not the full base64 payload.
Budget and Cost Tracking
Set per-day and per-month USD limits — requests that would exceed the budget are rejected before hitting the provider.
router = LLMRouter(
providers={...},
budget={"daily_usd": 5.0, "monthly_usd": 50.0},
)
stats = router.stats()
print(stats.total_cost_usd) # total spent since start
print(stats.saved_cost_usd) # saved via cache hits
print(stats.daily_spend_usd)
print(stats.budget_remaining_usd) # None if no limit is set
print(stats.cache_hit_rate) # 0.0–1.0
FastAPI Integration
pip install "llm-cache-router[fastapi]"
from fastapi import FastAPI
from llm_cache_router.middleware.fastapi import (
add_http_metrics_middleware,
mount_metrics_endpoint,
)
app = FastAPI()
add_http_metrics_middleware(app=app)
mount_metrics_endpoint(app=app, router=router, path="/metrics")
Exposed Prometheus metrics:
llm_router_http_requests_total{method,path,status}llm_router_http_request_duration_seconds_*(histogram)llm_router_cache_hits_total,llm_router_cache_misses_totalllm_router_cost_usd_total,llm_router_saved_cost_usd_total
Async Context Manager
async with LLMRouter(providers={...}) as router:
response = await router.complete(messages=[...], model="gpt-4o-mini")
# close() is called automatically — closes provider clients and cache connections
Supported Providers
| Provider | Streaming | Notes |
|---|---|---|
| OpenAI | yes | gpt-4o, gpt-4o-mini, o1-*, etc. |
| OpenAI-compatible | yes | Any endpoint speaking the OpenAI Chat Completions protocol: OpenRouter, vLLM, llama.cpp server, LiteLLM proxy, self-hosted inference |
| Anthropic | yes | Claude 3.5 Sonnet/Haiku, Opus |
| Google Gemini | yes | 1.5 Flash, 1.5 Pro |
| Ollama | yes | Any locally-served model |
| MiniMax | yes | MiniMax-Text-01 and others |
| Qwen (Dashscope) | yes | qwen-plus, qwen-max, etc. |
Any OpenAI-compatible endpoint works with a single provider entry — base_url is required, api_key is optional (local servers often run without one):
router = LLMRouter(
providers={
"openai_compatible": {
"base_url": "http://localhost:8000/v1", # vLLM / llama.cpp / LiteLLM proxy
"api_key": "optional", # omit for keyless local servers
"models": ["qwen2.5-32b-instruct"],
}
}
)
Adding a new provider = subclass LLMProvider, register with @register_provider("name"). See llm_cache_router/providers/base.py.
Architecture
llm_cache_router/
cache/ # memory (FAISS) / redis / qdrant backends
providers/ # openai, anthropic, gemini, ollama, minimax, qwen
strategies/ # cheapest, fastest, fallback
embeddings/ # SentenceEncoder, HashingEncoder
cost/ # CostTracker with daily/monthly budgets
middleware/ # FastAPI middleware
observability/ # Prometheus metrics
models.py # Pydantic models (Message, LLMResponse, CacheEntry, ...)
router.py # LLMRouter — public entrypoint
retry.py # RetryConfig + exponential backoff
warmup.py # async warmup helper
Development
git clone https://github.com/svalench/llm-cache-router.git
cd llm-cache-router
# using uv (recommended)
uv sync --all-extras
uv run pytest
# or plain pip
pip install -e ".[all,dev]"
pytest
Code quality is enforced in CI via:
ruff check(lint) andruff format --check(style)mypy --ignore-missing-imports(type check)pyteston Python 3.11, 3.12, 3.13 with coverage
Roadmap
- v0.3 — Request tracing hooks (OpenTelemetry spans).
- v0.4 — Streaming retry (reconnect on SSE drop); Django helpers and middleware.
- v0.5 — Persistent budget counters (Redis/SQLite) surviving process restarts; shared EMA latency metrics for multi-worker deployments.
- v1.0 — LLM-verified cache hits (cheap re-check of semantic matches on a mini model); pluggable pricing providers.
Contributing
Pull requests are welcome — see CONTRIBUTING.md for the full guide. Quick version:
- Open an issue first for anything larger than a small bug fix.
- Add tests for new behaviour.
- Run
ruff check,ruff format,mypyandpytestbefore pushing.
License
MIT — see LICENSE for details.
🇷🇺 Краткое описание (Russian)
llm-cache-router — лёгкая production-ready Python-библиотека для семантического кэширования LLM-запросов, мульти-провайдер роутинга и контроля бюджета. Экономит 30–70% на LLM-счетах за счёт векторного кэша, переключается между провайдерами (OpenAI, Anthropic, Gemini, Ollama, MiniMax, Qwen и любой OpenAI-compatible endpoint — OpenRouter, vLLM, llama.cpp server, LiteLLM proxy) без изменений в коде приложения, и включает встроенный трекинг стоимости с дневными/месячными лимитами. Поддерживает три бэкенда кэша (in-memory / Redis / Qdrant), инвалидацию и версионирование ключей кэша, режим точного совпадения, нативный стриминг для всех провайдеров и FastAPI-middleware с Prometheus-метриками.
v0.3.0: провайдер openai_compatible (OpenRouter, vLLM, llama.cpp, LiteLLM proxy), инвалидация и версионирование ключей кэша, режим точного совпадения. Release notes.
Установка:
pip install llm-cache-router
# с дополнительными бэкендами
pip install "llm-cache-router[redis]"
pip install "llm-cache-router[qdrant]"
pip install "llm-cache-router[fastapi]"
pip install "llm-cache-router[all]"
Требуется Python 3.11+. Полная документация и примеры — выше (на английском).
Демо: docker compose up --build (Redis + Qdrant + FastAPI) или Colab playground.
Metadata
Release files for llm-cache-router 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_cache_router-0.3.0.tar.gz | 45.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_cache_router-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 97.5 kB
Release files / llm_cache_router-0.3.0.tar.gz
| Download URL | llm_cache_router-0.3.0.tar.gz |
|---|---|
| Size | 45.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0dca8ca0600303106dad1d264ded502e4ac57af158466fc564bd699c932800fd
|
|
BLAKE2b-256 checksum How to use checksums |
4f451cecffb14f464b6f6bf7b9e7ba7739f2a8d3aa5420dc899df0b77c66de9d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency logRelease files / llm_cache_router-0.3.0-py3-none-any.whl
| Download URL | llm_cache_router-0.3.0-py3-none-any.whl |
|---|---|
| Size | 51.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0d1c990b5704545b459900199294eec9e3a8da5d019fb647ce50fd823dab094d
|
|
BLAKE2b-256 checksum How to use checksums |
63ed0da1ba455e75698b7005931d5676df29a413bfd484366ac95773aaf18b17
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency log