claude-router
claude-router is a local prompt router that picks the right Claude model tier and prepends the right scaffold using local embeddings, before you call the API.
Claude teams overspend on Sonnet or Opus because nobody has a fast, repeatable way to decide when Haiku plus structure is enough. claude-router classifies a prompt locally, chooses the right Claude tier, and prepends the right scaffold when scaffolding actually improves quality.
- "We default to Sonnet for everything because nobody trusts routing by hand."
- "Some prompts need structure, but we keep discovering that too late."
- "We know Haiku is cheaper, but we do not know when it is safe."
- "Prompt reviews catch model-choice mistakes after the API bill already happened."
Install
Requires: Python 3.10+, requests, numpy, and Ollama running locally with nomic-embed-text.
pip install claude-router
ollama pull nomic-embed-text
from claude_router import ClaudeRouter
router = ClaudeRouter()
result = router.route("Evaluate this research paper for methodological rigor")
print(result["model"], result["scaffold_key"])
print(result["pricing"]["input_usd_per_mtok"], result["pricing"]["output_usd_per_mtok"])
claude-haiku-4-5 calibrated-scoring
1.0 5.0
When To Use It
Use claude-router when you already call Claude models and want a local, deterministic routing layer for eval, research, content, and review prompts.
When Not To Use It
Do not use claude-router as a general agent framework, as proof that these exact routes transfer to your workload, or if you do not want an Ollama-based local classifier in the loop.
Results
| Task | Best Setup | Run cost (2026-03) | Quality vs. Baseline |
|---|---|---|---|
| Eval/scoring | Haiku + scaffold | $0.06 | MAE 1.0 (vs Sonnet raw: 1.2) |
| Research | Sonnet + scaffold | $0.28 | 8.49/10 (vs Opus raw: 7.45) |
| Content | Haiku + scaffold | $0.06 | 4/5 blind wins vs Sonnet |
| Code review | Sonnet (raw) | $0.28 | Sonnet raw preferred; scaffolds hurt coding |
Run costs are what these benchmark batches cost at the list prices in effect on their 2026-03 run dates. They are historical, not a forecast, and not current pricing — see Pricing.
These runs used the Claude 4.x generation (Haiku 4.5, Sonnet 4.6, Opus 4.6). The router
now returns the current generation for each tier (see Routing table),
so the routing table's evidence is one model generation behind the models it returns. It
has not been re-validated on Sonnet 5 or Opus 5 yet; claude-router --eval (below) checks
classification, not output quality.
Anti-findings
These are the blocker issues. The router handles them automatically:
- Scaffolds break operational tasks (0/9 success). Haiku treats constraints as meta-instructions instead of executing tasks.
- Scaffolds hurt coding (4.9 vs 6.4 raw). Don't scaffold code review, design, or debugging.
- Opus doesn't scaffold. Safety-critical evals need Opus raw (MAE 0.0), not scaffolded.
The routing table avoids these entirely: no scaffolds on operational, coding, safety-critical, or conversation tasks.
Quick start
from claude_router import ClaudeRouter
router = ClaudeRouter()
result = router.route("Evaluate this research paper for methodological rigor")
print(result["model"]) # claude-haiku-4-5
print(result["scaffold_key"]) # calibrated-scoring
print(result["pricing"]) # {'model_id': 'claude-haiku-4-5',
# 'input_usd_per_mtok': 1.0, 'output_usd_per_mtok': 5.0,
# 'input_usd_per_1k': 0.001, 'output_usd_per_1k': 0.005,
# 'basis': 'first_party_uncached_non_batch_global',
# 'as_of': '2026-09-11', 'source': 'https://platform.claude.com/...'}
# Build prompt with scaffold prepended
prompt = router.build_prompt("Evaluate this research paper...")
# → Pass prompt as system message to Anthropic API
Or CLI:
claude-router "Write a blog post about Q2 results"
From a source checkout without installing, python router.py "..." runs the same router.
How it works
- Embed your prompt using nomic-embed-text (~5ms)
- Compare against pre-computed task-category centroids
- Look up routing table: category → model + scaffold
- Return model ID and scaffold text
No LLM calls for routing. All locally in ~10ms. When the classifier is not confident in a category, the router defaults to Opus.
The 5 scaffolds
Each scaffold is validated through blind evaluation. They work by constraining the model's output space to the task structure.
See scaffolds.json for full text and evidence:
- calibrated-scoring: Integer 1-10, cite evidence, not generous/critical
- insight-first: Lead non-obvious, concrete recs, 3-4 sentences
- plan-first: g:goal;c:constraints;s:steps;r:risks prefix
- substance-check: Real gaps not surface, name issue and location
- bug-hunt: Specific bugs, line numbers, severity, one-line fix
Routing table
eval → Haiku + calibrated-scoring
research → Sonnet + insight-first
content → Haiku + insight-first
analytical_review → Haiku + substance-check
search → Haiku + plan-first
coding → Sonnet (raw)
operational → Sonnet (raw)
status_check → Haiku (raw)
conversation → Opus (raw)
safety_critical → Opus (raw)
Low confidence → Opus (safe default).
Tiers resolve to current-generation model IDs: Haiku → claude-haiku-4-5,
Sonnet → claude-sonnet-5, Opus → claude-opus-5. A fourth tier, fable →
claude-fable-5-1, is priced and accepted in custom routing tables, but no default
category routes there: it costs 2x Opus and none of the bundled evidence covers it.
Evaluating the routing table
claude-router --eval [cases.json] routes a labelled prompt set and reports how often the
classifier lands on the expected category and tier, which prompts it misrouted, and what
the routed tiers cost against sending every prompt to Sonnet:
claude-router --eval # shipped set: 24 prompts, 2 per category
claude-router --eval my_cases.json # {"cases": [{"prompt": ..., "category": ...}]}
from claude_router.evaluate import evaluate, load_cases
report = evaluate(ClaudeRouter(), load_cases(), tokens_in=1000, tokens_out=1000)
report["category_accuracy"], report["tier_accuracy"], report["misroutes"]
report["cost"]["routed_over_baseline"] # routed list price / all-Sonnet list price
The cost figure is a list-price estimate at the stated tokens per call, not a measurement. The shipped labels are the intended routes, not benchmark ground truth: a miss means the classifier and the label disagree, and either may be wrong. Routing still needs Ollama; only the scoring is offline.
Pricing
route() returns the routed model's exact list prices, with the date and source they
were read from, so you can do the arithmetic on your own token volumes:
| Tier | Model ID | Input $/MTok | Output $/MTok |
|---|---|---|---|
| Haiku | claude-haiku-4-5 |
$1.00 | $5.00 |
| Sonnet | claude-sonnet-5 |
$2.00 | $10.00 |
| Opus | claude-opus-5 |
$5.00 | $25.00 |
| Fable | claude-fable-5-1 |
$10.00 | $50.00 |
Base (uncached, non-batch, global-inference) first-party Claude API prices as of
2026-09-11, from platform.claude.com/docs/en/about-claude/pricing.
Prompt caching, the Batch API, and inference_geo all apply multipliers this table does
not model. The single maintained copy is src/claude_router/model_pricing.json,
read by both the packaged router and router.py.
This repo publishes no cost-savings total. What routing saves depends on your prompt mix and, critically, on your input:output token ratio — output costs 5x input on every tier above, so a savings figure computed from input price alone is wrong. Multiply your own measured token counts by the two columns above.
result["cost_per_1k"] is still returned for existing consumers. It is deprecated and
input-only (result["cost_per_1k_basis"] == "input_tokens_only"): it is the input price
per 1K tokens and has never included output tokens. Use result["pricing"] for anything
that needs to be right.
Customization
Swap scaffolds, centroids, or routing table:
router = ClaudeRouter(
centroids_path="my_centroids.json",
routing_table_path="my_routing.json",
scaffolds_path="my_scaffolds.json"
)
Limitations
- Requires Ollama locally (for embeddings)
- Centroids trained on one task distribution — test on your workload
- The classifier is not perfect — ambiguous prompts fall to low confidence and default to Opus
- Anti-findings are real: scaffolds on coding/operational make things worse
- The routing table was validated on the 4.x generation; it now returns Sonnet 5 and Opus 5 without a fresh quality benchmark on them
- Prices are a dated snapshot, not a live feed — re-check
model_pricing.jsonagainst the published source before relying on it for billing - No Lite mode (Haiku-first routing): it was planned for v1.1 but did not ship in 1.1.0
Evidence
Benchmarks: benchmarks/ | Raw citations: scaffolds.json | License: MIT
Key experiments: 4-condition code/research crossover, scaffolds-vs-operational stress test, scaffolded Sonnet beats Opus 75% on research (6/8 blind wins, 140 API calls).
Need this calibrated to your pipeline? Open an issue with the task categories and failure cases you want to benchmark.
About Hermes Labs
Hermes Labs is an AI reliability engineering studio for product and engineering teams shipping production agents and LLM applications. We find the structural AI failures standard evals miss, then harden retrieval, memory, agents, and the language layers around production AI systems with runtime controls and defensible evidence.
Browse the open-source catalog or contact roli@hermes-labs.ai.
Metadata
Release files for claude-router 1.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| claude_router-1.1.1.tar.gz | 291.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| claude_router-1.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 401.0 kB
Release files / claude_router-1.1.1.tar.gz
| Download URL | claude_router-1.1.1.tar.gz |
|---|---|
| Size | 291.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
551ddb452f58ea0d425e0f347b5597066206a9db94de3be895fe063622993e8d
|
|
BLAKE2b-256 checksum How to use checksums |
4f58d183d2d653b0df774ed5c56768624663b1af0a47d382a172ff036dc84735
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.6
|
Release files / claude_router-1.1.1-py3-none-any.whl
| Download URL | claude_router-1.1.1-py3-none-any.whl |
|---|---|
| Size | 109.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
97f69d47eafd80a5c257f38c9632a1770811e8219889e9294464c78fa1ecac0a
|
|
BLAKE2b-256 checksum How to use checksums |
928998af8ecc1f486a11bd24f3c9fdb587d7b80048f969b5d99a9da5571c78bd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.6
|