Skip to main content

claude-router

claude-router is a local prompt router that picks the right Claude model tier and prepends the right scaffold using local embeddings, before you call the API.

Claude teams overspend on Sonnet or Opus because nobody has a fast, repeatable way to decide when Haiku plus structure is enough. claude-router classifies a prompt locally, chooses the right Claude tier, and prepends the right scaffold when scaffolding actually improves quality.

  • "We default to Sonnet for everything because nobody trusts routing by hand."
  • "Some prompts need structure, but we keep discovering that too late."
  • "We know Haiku is cheaper, but we do not know when it is safe."
  • "Prompt reviews catch model-choice mistakes after the API bill already happened."

Install

Requires: Python 3.10+, requests, numpy, and Ollama running locally with nomic-embed-text.

pip install claude-router
ollama pull nomic-embed-text
from claude_router import ClaudeRouter

router = ClaudeRouter()
result = router.route("Evaluate this research paper for methodological rigor")
print(result["model"], result["scaffold_key"])
print(result["pricing"]["input_usd_per_mtok"], result["pricing"]["output_usd_per_mtok"])
claude-haiku-4-5 calibrated-scoring
1.0 5.0

When To Use It

Use claude-router when you already call Claude models and want a local, deterministic routing layer for eval, research, content, and review prompts.

When Not To Use It

Do not use claude-router as a general agent framework, as proof that these exact routes transfer to your workload, or if you do not want an Ollama-based local classifier in the loop.

claude-router preview

Results

Task Best Setup Run cost (2026-03) Quality vs. Baseline
Eval/scoring Haiku + scaffold $0.06 MAE 1.0 (vs Sonnet raw: 1.2)
Research Sonnet + scaffold $0.28 8.49/10 (vs Opus raw: 7.45)
Content Haiku + scaffold $0.06 4/5 blind wins vs Sonnet
Code review Sonnet (raw) $0.28 Sonnet raw preferred; scaffolds hurt coding

Run costs are what these benchmark batches cost at the list prices in effect on their 2026-03 run dates. They are historical, not a forecast, and not current pricing — see Pricing.

These runs used the Claude 4.x generation (Haiku 4.5, Sonnet 4.6, Opus 4.6). The router now returns the current generation for each tier (see Routing table), so the routing table's evidence is one model generation behind the models it returns. It has not been re-validated on Sonnet 5 or Opus 5 yet; claude-router --eval (below) checks classification, not output quality.

Anti-findings

These are the blocker issues. The router handles them automatically:

  • Scaffolds break operational tasks (0/9 success). Haiku treats constraints as meta-instructions instead of executing tasks.
  • Scaffolds hurt coding (4.9 vs 6.4 raw). Don't scaffold code review, design, or debugging.
  • Opus doesn't scaffold. Safety-critical evals need Opus raw (MAE 0.0), not scaffolded.

The routing table avoids these entirely: no scaffolds on operational, coding, safety-critical, or conversation tasks.

Quick start

from claude_router import ClaudeRouter

router = ClaudeRouter()
result = router.route("Evaluate this research paper for methodological rigor")

print(result["model"])           # claude-haiku-4-5
print(result["scaffold_key"])    # calibrated-scoring
print(result["pricing"])         # {'model_id': 'claude-haiku-4-5',
                                 #  'input_usd_per_mtok': 1.0, 'output_usd_per_mtok': 5.0,
                                 #  'input_usd_per_1k': 0.001, 'output_usd_per_1k': 0.005,
                                 #  'basis': 'first_party_uncached_non_batch_global',
                                 #  'as_of': '2026-09-11', 'source': 'https://platform.claude.com/...'}

# Build prompt with scaffold prepended
prompt = router.build_prompt("Evaluate this research paper...")
# → Pass prompt as system message to Anthropic API

Or CLI:

claude-router "Write a blog post about Q2 results"

From a source checkout without installing, python router.py "..." runs the same router.

How it works

  1. Embed your prompt using nomic-embed-text (~5ms)
  2. Compare against pre-computed task-category centroids
  3. Look up routing table: category → model + scaffold
  4. Return model ID and scaffold text

No LLM calls for routing. All locally in ~10ms. When the classifier is not confident in a category, the router defaults to Opus.

The 5 scaffolds

Each scaffold is validated through blind evaluation. They work by constraining the model's output space to the task structure.

See scaffolds.json for full text and evidence:

  • calibrated-scoring: Integer 1-10, cite evidence, not generous/critical
  • insight-first: Lead non-obvious, concrete recs, 3-4 sentences
  • plan-first: g:goal;c:constraints;s:steps;r:risks prefix
  • substance-check: Real gaps not surface, name issue and location
  • bug-hunt: Specific bugs, line numbers, severity, one-line fix

Routing table

eval              → Haiku   + calibrated-scoring
research          → Sonnet  + insight-first
content           → Haiku   + insight-first
analytical_review → Haiku   + substance-check
search            → Haiku   + plan-first

coding            → Sonnet  (raw)
operational       → Sonnet  (raw)
status_check      → Haiku   (raw)
conversation      → Opus    (raw)
safety_critical   → Opus    (raw)

Low confidence → Opus (safe default).

Tiers resolve to current-generation model IDs: Haiku → claude-haiku-4-5, Sonnet → claude-sonnet-5, Opus → claude-opus-5. A fourth tier, fable → claude-fable-5-1, is priced and accepted in custom routing tables, but no default category routes there: it costs 2x Opus and none of the bundled evidence covers it.

Evaluating the routing table

claude-router --eval [cases.json] routes a labelled prompt set and reports how often the classifier lands on the expected category and tier, which prompts it misrouted, and what the routed tiers cost against sending every prompt to Sonnet:

claude-router --eval                      # shipped set: 24 prompts, 2 per category
claude-router --eval my_cases.json        # {"cases": [{"prompt": ..., "category": ...}]}
from claude_router.evaluate import evaluate, load_cases

report = evaluate(ClaudeRouter(), load_cases(), tokens_in=1000, tokens_out=1000)
report["category_accuracy"], report["tier_accuracy"], report["misroutes"]
report["cost"]["routed_over_baseline"]   # routed list price / all-Sonnet list price

The cost figure is a list-price estimate at the stated tokens per call, not a measurement. The shipped labels are the intended routes, not benchmark ground truth: a miss means the classifier and the label disagree, and either may be wrong. Routing still needs Ollama; only the scoring is offline.

Pricing

route() returns the routed model's exact list prices, with the date and source they were read from, so you can do the arithmetic on your own token volumes:

Tier Model ID Input $/MTok Output $/MTok
Haiku claude-haiku-4-5 $1.00 $5.00
Sonnet claude-sonnet-5 $2.00 $10.00
Opus claude-opus-5 $5.00 $25.00
Fable claude-fable-5-1 $10.00 $50.00

Base (uncached, non-batch, global-inference) first-party Claude API prices as of 2026-09-11, from platform.claude.com/docs/en/about-claude/pricing. Prompt caching, the Batch API, and inference_geo all apply multipliers this table does not model. The single maintained copy is src/claude_router/model_pricing.json, read by both the packaged router and router.py.

This repo publishes no cost-savings total. What routing saves depends on your prompt mix and, critically, on your input:output token ratio — output costs 5x input on every tier above, so a savings figure computed from input price alone is wrong. Multiply your own measured token counts by the two columns above.

result["cost_per_1k"] is still returned for existing consumers. It is deprecated and input-only (result["cost_per_1k_basis"] == "input_tokens_only"): it is the input price per 1K tokens and has never included output tokens. Use result["pricing"] for anything that needs to be right.

Customization

Swap scaffolds, centroids, or routing table:

router = ClaudeRouter(
    centroids_path="my_centroids.json",
    routing_table_path="my_routing.json",
    scaffolds_path="my_scaffolds.json"
)

Limitations

  • Requires Ollama locally (for embeddings)
  • Centroids trained on one task distribution — test on your workload
  • The classifier is not perfect — ambiguous prompts fall to low confidence and default to Opus
  • Anti-findings are real: scaffolds on coding/operational make things worse
  • The routing table was validated on the 4.x generation; it now returns Sonnet 5 and Opus 5 without a fresh quality benchmark on them
  • Prices are a dated snapshot, not a live feed — re-check model_pricing.json against the published source before relying on it for billing
  • No Lite mode (Haiku-first routing): it was planned for v1.1 but did not ship in 1.1.0

Evidence

Benchmarks: benchmarks/ | Raw citations: scaffolds.json | License: MIT

Key experiments: 4-condition code/research crossover, scaffolds-vs-operational stress test, scaffolded Sonnet beats Opus 75% on research (6/8 blind wins, 140 API calls).

Need this calibrated to your pipeline? Open an issue with the task categories and failure cases you want to benchmark.


About Hermes Labs

Hermes Labs is an AI reliability engineering studio for product and engineering teams shipping production agents and LLM applications. We find the structural AI failures standard evals miss, then harden retrieval, memory, agents, and the language layers around production AI systems with runtime controls and defensible evidence.

Browse the open-source catalog or contact roli@hermes-labs.ai.

Metadata

Release files for claude-router 1.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for claude-router 1.1.1
File Size Uploaded
claude_router-1.1.1.tar.gz 291.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for claude-router 1.1.1
File Interpreter ABI Platform
claude_router-1.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 401.0 kB

Release files / claude_router-1.1.1.tar.gz

Download URL claude_router-1.1.1.tar.gz
Size 291.1 kB
Tags Source
SHA-256 checksum
How to use checksums
551ddb452f58ea0d425e0f347b5597066206a9db94de3be895fe063622993e8d
BLAKE2b-256 checksum
How to use checksums
4f58d183d2d653b0df774ed5c56768624663b1af0a47d382a172ff036dc84735
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.6

Release files / claude_router-1.1.1-py3-none-any.whl

Download URL claude_router-1.1.1-py3-none-any.whl
Size 109.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
97f69d47eafd80a5c257f38c9632a1770811e8219889e9294464c78fa1ecac0a
BLAKE2b-256 checksum
How to use checksums
928998af8ecc1f486a11bd24f3c9fdb587d7b80048f969b5d99a9da5571c78bd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.6

Release history Release notifications | RSS feed

This release

1.1.1 This release

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page