llm-router
Spend less of your Claude Pro or Max plan on routine prompts.
llm-router drafts an answer on a free local model first, so the cheap questions can be
answered cheaply. No API keys. No change to how you work.
How the default mode actually works: the draft is injected into Claude's context
as an unverified hint, and Claude decides whether to use it. A draft being produced does
not by itself mean a Claude turn was skipped or quota was saved — Claude still takes the
turn. Only the explicit turn-replacement modes (LLM_ROUTER_ZERO_CLAUDE=1) substitute
the routed answer outright. Measured on 105 real prompts from this author's own sessions,
76% produced a draft and 72% produced one worth relaying.
Install in 30 seconds
pip install llm-routing # installs the `llm-router` command
Works with Claude Code, Codex, and Gemini CLI · No API keys required on Claude Pro/Max
Local-first. No hosted proxy. No account required.
📑 Table of Contents
Why people install this
You are on a Claude Pro or Max plan. You have not spent a cent beyond the subscription. And at 3pm you hit the five-hour limit and stop working.
The cause is not that you asked too much. It is that every prompt went to the premium model — "what does this error mean", "reformat this JSON", "is the service up" — and each one drew down the same quota as the architectural question you actually needed it for.
llm-router runs inside your coding tool's own lifecycle. It reads each prompt
before the model does, sends the routine ones to a local or cheap model, and
leaves your seat for the work that needs it. Same workflow, same commands, same
transcript — the model choice changes underneath.
Why a proxy cannot do this
Most routers in this category are a proxy: you point your agent at a local endpoint and it forwards requests using your API keys. That design has a hard limit — a proxy cannot intercept a session authenticated by a subscription, because there is no key to forward.
If you pay per token, a proxy serves you well and there are good ones. If you pay a flat monthly fee and the thing you run out of is quota, a proxy has nothing to offer, and that is the gap this fills.
| Pays per token | Pays a subscription | |
|---|---|---|
| What runs out | your invoice | your five-hour window |
| Needs API keys | yes | no |
| A proxy can help | yes | no — nothing to intercept |
| llm-router helps | yes | yes |
Two things worth checking before you install
- It works with zero API keys. On a Claude subscription, routing goes through MCP tools and local models. Adding keys widens the pool; nothing requires them.
- The routing quality is measured by someone else. llm-router is scored on RouterArena, a third-party accuracy-versus-cost leaderboard. What was measured, what it cost, and what did not work is written up in docs/ROUTERARENA.md — including the negative results.
On the RouterArena leaderboard
llm-router is benchmarked on RouterArena,
a community leaderboard scoring routers on accuracy versus cost, plus optimality,
robustness and latency.
The claim worth reading is not the badge. docs/ROUTERARENA.md states what was measured, on which split, what it cost to reproduce, and what failed — including that skill-cluster classification never beat simply always picking one model, and that tuning on a proxy split misled by 4.25 points. Rank moves as new routers land; see the live leaderboard for the current standing.
Quick Start
1. Install
pip install llm-routing
llm-router install
2. Add providers (optional)
export OPENAI_API_KEY="sk-..." # GPT-4o, o3
export GEMINI_API_KEY="AIza..." # Gemini Flash/Pro (free tier available)
export OLLAMA_BASE_URL="http://localhost:11434" # Local models (free)
export OPENROUTER_API_KEY="sk-or-v1-…" # 343 OpenRouter models (qwen, deepseek, grok, …)
Works with zero API keys on Claude Code Pro/Max subscriptions — routing uses MCP tools that call external models only when beneficial. Add OPENROUTER_API_KEY to unlock the open-weight workhorse pool used by the cost_aggressive policy.
3. Verify
llm-router health # Check provider connectivity
If you already use Claude Code, Codex, or Gemini CLI, keep your existing workflow and let llm-router choose models underneath it.
Example Routing
| Prompt | Routed to |
|---|---|
| "What does this Python error mean?" | Ollama / Gemini Flash / Codex |
| "Refactor this endpoint" | GPT-4o / Gemini Pro |
| "Design a distributed tracing strategy" | o3 / Claude Opus |
The exact chain depends on your configured providers, budget profile, and routing policy.
Works With
| Tool | Mode | Savings (this host) |
|---|---|---|
| Claude Code | Full auto-routing via hooks | 60–80% |
| Codex CLI | Manual MCP tools · hooks 🔜 | 30–50% |
| Gemini CLI | Full auto-routing via hooks | 50–70% |
| VS Code / Cursor | Manual MCP tools · hooks 🔜 | 30–50% |
| Any MCP client | Manual MCP tools | Varies |
- Full auto-routing means hooks intercept prompts and route automatically with no workflow change.
- Manual MCP tools means routing is available on demand through tools such as
llm_query. - 🔜 means the host supports prompt interception and we have not shipped it yet — not that
it cannot be done. Codex CLI ships
UserPromptSubmit(enabled by default, and itsPreToolUsecan even rewrite arguments); Cursor shipsbeforeSubmitPrompt. Both can block a prompt before the model sees it, which is the same mechanism Claude Code uses today.
The full picture, including what each host genuinely cannot do and which payload fields have been verified against a real run rather than read off a docs page, is in guide/HOST_SUPPORT_MATRIX.md.
llm-router install # Claude Code (default)
llm-router install --host codex # Codex CLI
llm-router install --host gemini-cli # Gemini CLI
llm-router install --host vscode # VS Code
llm-router install --host cursor # Cursor
See guide/HOST_SUPPORT_MATRIX.md for full details on each host.
Protect your Claude Code 5-hour quota
enforce: smart + mode: zero_claude makes prompts either complete externally or stop
before native Claude runs — see
guide/GETTING_STARTED.md.
How It Works
User prompt
│
▼
┌──────────────────────┐
│ Complexity Classifier │ ← Heuristic (free, instant) or Ollama/Flash ($0.0001)
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Free-First Router │ ← Tries cheapest model first, walks up the chain
│ │
│ Ollama (free) │
│ → Codex (prepaid) │
│ → Gemini Flash │
│ → GPT-4o / Claude │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Guards (parallel) │ ← Circuit breaker, budget pressure, quality check
└──────────┬───────────┘
│
▼
Response + cost logged to local SQLite
Classification is free for many tasks (regex heuristics catch ~70%) or near-free for ambiguous prompts when using local Ollama or Gemini Flash.
Features
Beyond "send cheap prompts to cheap models":
- Secrets never leave your machine. A prompt containing an API key, token or private key routes to local models only — fail-closed, so it cannot reach an external provider.
- Cost-inverted subscription routing. Free/local first for simple and moderate
prompts, your one paid seat first for complex ones, and the seat demoted when its quota
is strained. Opt in with
LLM_ROUTER_SUBSCRIPTION_PROVIDER. - Automatic fallback with circuit breakers. A provider that fails or rate-limits is skipped, not retried into the ground.
- You can see it working. A status line, terminal title and OS notification show the last model routed, savings and health — for hosts with no native statusline.
- Session-end summary. Savings vs baseline, tier mix, per-provider cost, latency p50/p95/p99 and top routes.
- Media and pipelines too.
llm_image/llm_video/llm_audio, andllm_orchestratefor multi-step research.
CLI
llm-router install # wire up your host (Claude Code by default)
llm-router health # provider connectivity
llm-router status # savings + quota at a glance
llm-router doctor # diagnose a broken setup
llm-router okf index # index this repo so routed models can see your code
llm-router okf status # what is in the knowledge store, per project
llm-router sessions status # is any session's context unreadable?
okf index is worth running once per repo you work in. Without it the knowledge
store can only learn from answers that were already routed, which is a deadlock —
nothing routes because the model has no context, and the store stays empty because
nothing routed.
Full command reference: guide/GETTING_STARTED.md
Providers
20+ providers, free-first. Ollama (local, free) leads the chain; OpenRouter (hundreds of models behind one key — see the provider reference for the current list) is the biggest single unlock; Gemini and Groq have usable free tiers. Anthropic works via your existing Claude subscription — no API key needed.
Every provider, its models, cost tier and env var: guide/PROVIDERS.md
Routing Policies
A policy sets how eagerly the router routes away from your premium model —
conservative (10–15% savings) through balanced (the default, 35–45%) to
cost_aggressive (70–85%, needs OPENROUTER_API_KEY).
llm-router policy set cost_aggressive
All six policies, thresholds and the YAML schema: guide/POLICIES.md
MCP Tools
60 tools across routing, analysis, code, media, budget and diagnostics — exposed to any
MCP host. The default consolidated surface shows 11 front-door tools; set
LLM_ROUTER_SLIM=full for all 60.
Every tool with its signature: guide/TOOLS.md
Savings: How It Works
Savings are calculated by comparing actual spend against a baseline of routing every task to Claude Sonnet/Opus.
Methodology:
- Each routed task logs: model used, tokens consumed, estimated cost
- A baseline cost is computed as if the same tokens were processed by the most expensive model in the chain
- Savings =
(baseline - actual) / baseline
Assumptions and limitations:
- Baseline assumes you would have used Opus/Sonnet for everything (worst case)
- Token estimates use
len(text) / 4approximation, not exact tokenizer counts - Cost data comes from LiteLLM's pricing tables (may lag provider price changes)
- Savings vary significantly by workload — code-heavy sessions route more to cheap models
- The router itself adds small overhead (classification costs ~$0.0001 per ambiguous task)
On savings figures. Percentages quoted anywhere in this project are a counterfactual —
what the same tokens would have cost at API list price, against what was actually spent. On a
flat-rate subscription that is not money saved; it is quota preserved, and the two are not
interchangeable. The "35–80%" and "87%" figures are single-user observations over particular
development periods, with no stated denominator, and should be read as anecdotes rather than
as a range you can expect. llm-router summary reports what your own usage actually did.
Trust, Privacy, and Local-First Design
llm-router runs entirely on your machine. There is no hosted proxy, nothing is sent to an llm-router service, and no account is required.
It does keep local telemetry, and since 2026-09 that telemetry steers routing: every model
attempt and its outcome is recorded so a chronically slow model can be demoted, and a routed
answer is marked successful only if it is actually usable. All of it stays in
~/.llm-router/ and none of it leaves the machine.
| What | Where | Details |
|---|---|---|
| Your prompts | Sent to configured providers | Exactly like using those providers directly |
| API keys | .env or ~/.llm-router/config.yaml |
Local files, never transmitted |
| Usage logs | ~/.llm-router/usage.db |
Unencrypted SQLite (filesystem permissions) |
| Classification cache | In-memory | Cleared on process restart |
| Hook scripts | ~/.claude/hooks/ |
Local shell scripts, inspectable |
| Model attempts | ~/.llm-router/attempts.jsonl |
Per-attempt outcome and latency; rotated |
| Session context | ~/.llm-router/projects/<id>/ |
Your prompts, tool calls and routed answers, per project |
| Repo knowledge (OKF) | ~/.llm-router/knowledge/ |
Symbols and paths from tracked files, never prose |
What we do:
- Scrub API keys from structured logs
- Detect hook deadlocks before installation
- Store all data locally in
~/.llm-router/ - Respect provider rate limits and TOS
What you should know:
- Prompts are sent to whichever provider the router selects — review your provider's privacy policy
- Usage logs (SQLite) are not encrypted at rest — use full-disk encryption if needed
- The router cannot prevent model jailbreaks or prompt injection at the provider level
LLM_ROUTER_DIRECT_EXECUTION — read this before your first run
This is on by default. When enabled, hooks/auto-route.py tries to answer a prompt
locally before Claude Code sees it. For prompts it classifies as needing file work, it runs
a tool-calling agent loop that hands the local model three tools — write_file, edit_file
and run_command — for up to 15 iterations. Writes do NOT reach disk by default:
LLM_ROUTER_AGENT_WRITES defaults to propose, so the loop returns a patch.
What is actually enforced:
write_file/edit_fileare confined to the project root. This works as described.run_commanddoes not use a shell. It isshlex.split+subprocess.run(argv), so pipes, redirections,;and$(...)are literal arguments, not operators.run_commandpasses through two independent layers:agent_writes.guard_command(an allowlist of inspection programs, plus blocked subcommands) and a regex blocklist of top-level destructive patterns. The allowlist is the stronger of the two and is what blocksgit push --force,npm install,pip installandrm -rf ./src.- Setting
LLM_ROUTER_AGENT_COMMANDS=alldisables the allowlist, leaving only the regex. Nothing sets it for you.
What that blocklist does not stop (measured, not estimated): targeted deletes inside the
project (rm -rf ./src), $HOME deletes via shell expansion, git push --force,
git reset --hard, arbitrary npm/pip install, reads outside the project
(cat ../../.ssh/id_rsa), network exfiltration (curl -X POST … -d @.env), and echoing
API keys. It stops catastrophic system damage — not project damage, credential
disclosure, or exfiltration.
Turn it off:
export LLM_ROUTER_DIRECT_EXECUTION=false
Routing still works with it disabled; you lose only the local pre-answer path.
Since 13.2.0, a draft that reaches you has passed two grounding checks: it may not
cite a file, or call a function, that exists neither in the material it was given nor
in the indexed repo. A draft that does is discarded and the turn falls through to
Claude. This catches the mechanical way a context-fed answer goes wrong — a confident
reference to a test that was never written. It does not verify that the answer is
correct, and it cannot see invented prose; LLM_ROUTER_GROUNDING_CHECK=off and
LLM_ROUTER_SYMBOL_GROUNDING=off disable them.
See SECURITY.md for the full analysis and the responsible disclosure policy.
Configuration
Everything is environment variables — no config file required to start:
export OPENROUTER_API_KEY="sk-or-v1-..." # biggest single unlock
export OLLAMA_BASE_URL="http://localhost:11434" # local, free
export LLM_ROUTER_POLICY="cost_aggressive" # routing policy
export LLM_ROUTER_ENFORCE="smart" # off | advise | smart | hard
export LLM_ROUTER_OLLAMA_TIMEOUT=45 # seconds; 45 clears a real local p50
export LLM_ROUTER_PROJECT_ROOT="$PWD" # scope the knowledge store explicitly
LLM_ROUTER_OLLAMA_TIMEOUT matters more than it looks. It was 4s before 13.2.0, and
no local model can answer in 4s — measured p50s on an M-series machine are 11-28s, so
every local attempt aborted and fell through to Claude. If you run larger models,
raise it further rather than wondering why nothing routes.
Full reference, config file schema and per-host overrides: guide/GETTING_STARTED.md
Documentation
Full index: guide/README.md
| Document | Purpose |
|---|---|
| Quick Start (2 min) | Fastest path to working routing |
| Getting Started | Full setup walkthrough |
| Host Support Matrix | Per-host feature comparison |
| Providers | Provider setup and model recommendations |
| Routing Policies | routing.yaml schema and authoring your own policy |
| Tool Reference | All 60 MCP tools with examples |
| Architecture | Internal design and module structure |
| Troubleshooting | Common issues and fixes |
| Testing the Router | Isolation suite for verifying routing health |
| Benchmarks | Model cost/latency/quality table, regenerated by CI |
| Changelog | Release notes (archive) |
Enterprise
llm-router is built for individual developers and small teams: local cost savings, zero
ops overhead, no hosted anything. If you need team-wide policy enforcement, audit export,
SSO or per-org budgets, that is what Chuzom is for.
Contributing
Contributions welcome. See CONTRIBUTING.md for full guidelines.
git clone https://github.com/ypollak2/llm-router.git
cd llm-router
uv sync --extra dev
uv run pytest tests/ -q # Run tests (1900+)
uv run ruff check src/ tests/ # Lint
| Name | What it is |
|---|---|
llm-routing |
Current PyPI package (pip install llm-routing) |
llm-router |
CLI command and GitHub repo name |
claude-code-llm-router |
Deprecated legacy package (redirects to llm-routing) |
⭐ If llm-router saved you money, star the repo — it helps other developers discover it.
Issues · Discussions · PyPI · Changelog
MIT License
Release files for llm-routing 13.3.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_routing-13.3.2.tar.gz | 2.8 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_routing-13.3.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.5 MB
Release files / llm_routing-13.3.2.tar.gz
| Download URL | llm_routing-13.3.2.tar.gz |
|---|---|
| Size | 2.8 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d4efde707886245cbc2362ff4e97a163eff29a780c8cb1c3d38c32fb932f0fdb
|
|
BLAKE2b-256 checksum How to use checksums |
1f2047c52475efffb90e6a1841d80524fbae82971a223d6927b62698dfb3c0a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / llm_routing-13.3.2-py3-none-any.whl
| Download URL | llm_routing-13.3.2-py3-none-any.whl |
|---|---|
| Size | 2.7 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0c8e6a96d8a3f0657b8d11feb29730ac8a491c32e20caf69499e50c4d079edcd
|
|
BLAKE2b-256 checksum How to use checksums |
2b8197a708367996b61770f9f4f9433cf5570e8b2c280ed92c3d8f648010cbb5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|