This release is a pre-release and may not be stable for production use.
argrouter — pick the LLM, and how hard it should think, per query
argrouter routes each request to the model and reasoning effort that maximise
P(correct) − λ · expected cost, where the cost is what that model actually bills on
similar requests (hidden reasoning tokens included), not its list price. Two models with
the same price tag can differ 10× in real cost because one of them thinks for 3,000
tokens before answering.
On RouterArena (ICLR 2026, the open router leaderboard) argrouter scores 76.29 in RouterArena's automated evaluation, at $0.54 per 1,000 queries; the submission is under review. Details below.
The cost engine underneath is usable on its own:
from argrouter.catalog.pricing import Catalog, expected_cost
catalog = Catalog.from_snapshot("prices.json") # vendored, offline, SHA-pinned
model = catalog.get("anthropic/claude-sonnet-5.5") # unknown id raises; never $0
cost = expected_cost(
model,
input_tokens=20_000, # a RAG call: long prompt, short answer
expected_output_tokens=500, # forecast, not a YAML constant
cached_input_tokens=18_000, # cache state enters the decision
)
print(cost.total_usd) # 0.0141
print(cost.as_dict()) # every component separately auditable
pip install --pre argrouter # alpha; one dependency: httpx
Status: alpha. The cost engine, provider selection and router inference API are released and tested. Router training and the trained weights are not part of the open-source package.
Why cost-aware LLM routing usually gets the cost wrong
Output tokens are typically 60–90% of an LLM bill, and the field does not forecast them. Three failures are common, and all three are in shipped code today:
| Failure | Consequence |
|---|---|
Adding input_rate + output_rate |
Ranking is wrong for any workload that isn't 1:1 |
A constant expected_output_tokens in YAML |
The number that dominates the bill is a guess you typed |
Unpriced model treated as $0 |
The model nobody priced always wins "cheapest" |
argrouter's position is narrow and checkable: be the project whose cost number is right. Not another selection algorithm — a correct denominator.
Does routing break prompt caching?
Usually yes, and this is the strongest argument against routers in general. Prompt caches are model-scoped. Cache reads cost 0.1× base input — and as little as 0.025× on some models — so switching models mid-session can forfeit a 75–97% discount to chase a smaller routing saving.
argrouter treats this as a first-class cost term rather than a footnote:
- Cache state is priced at decision time. The forfeited cache discount is subtracted from the candidate's expected saving before it is compared.
- Minimum-cacheable thresholds are honoured. Below the per-model minimum, providers silently do not cache and charge full rate. argrouter records that explicitly instead of quietly over-estimating the discount.
- Session affinity keeps a conversation on its model unless the measured saving exceeds the measured cache loss.
When we publish a benchmark, the baseline runs with caching fully enabled and warmed, and the cache-hit rate of every arm is published next to its cost. A cost claim measured against a cache-disabled baseline is void, and we would rather say that ourselves than have it said in a comment thread.
When argrouter will not help you
- Long agentic sessions on one model with a warm cache. Keep the cache. Use cost tracking only.
- Uniformly hard workloads. Published routers beat random routing by ~14% on knowledge-dense benchmarks while still needing the frontier model for over half of calls. Some traffic is simply not routable.
- Single-shape workloads. If every request looks the same, choose the model once at build time. You do not need a router.
- Workloads where output style matters. Switching models changes voice.
You should know this before you install it, not after.
RouterArena
RouterArena (ICLR 2026) scores routers on 8,400 queries from 23 public benchmarks: accuracy, weighted with log-scaled cost (β = 0.1). argrouter on all 8,400 queries, scored by RouterArena's automated evaluation (PR #220; prices: OpenRouter list prices on 2026-10-07):
| router | arena | accuracy | $ / 1K queries |
|---|---|---|---|
| argrouter (submitted, under review) | 76.29 | 79.34% | 0.54 |
| KT-ModelRouter (current #1) | 76.28 | 78.14% | 0.27 |
| Sqwish Router (#2) | 76.21 | 79.76% | 0.70 |
| Divyam (#3) | 75.85 | 78.59% | 0.48 |
| NotDiamond (powers OpenRouter Auto) | 57.29 | 60.83% | 4.10 |
| RouteLLM | 48.07 | 47.04% | 0.27 |
How it was run. The pool is five models; reasoning effort is part of the choice (gemini-3.8-flash at low effort, gemma-4-31b, glm-5.3-flash, gpt-6-luna, mimo-v2.6-flash), chosen by greedy forward selection. The router reads only the question content: the instruction and answer-format paragraphs are dropped with generic rules, and no RouterArena file is read. It was trained on 3,562 held-out items from the same public sources, every RouterArena item excluded, with sources weighted equally.
The pool and λ were fixed on calibration data before RouterArena was routed, and nothing was tuned on RouterArena results. Submission to the official leaderboard is pending.
How much does it actually save?
Model routing: not measured yet. When it is, the number will be reported as all-in cost (including this library's own overhead and every retried call), per stratum, against the best fixed single model — not against the most expensive model in the pool, which is how savings in this category are usually inflated.
Provider selection (same model, many providers): measured, and the honest answer
is "about the same as OpenRouter's own price sort". Two runs, 377 billed calls,
4 open-weight models × 3 workloads, every cost taken from OpenRouter's usage.cost
(raw data and method):
| strategy | billed, both runs | vs thrift |
|---|---|---|
| OpenRouter default routing | $0.03123 | +67% |
OpenRouter sort: price + same quantization floor |
$0.01892 | +1% |
| rank providers by input + output price (LiteLLM/Plano rule) | $0.01885 | +1% |
| argrouter | $0.01874 | — |
What this means: if you call open-weight models through OpenRouter, turning on
any price-aware provider sort saves ~40% over the default, and that is most of
the win. argrouter's mix-aware ranking only changes the pick when two providers'
input/output prices cross (here: gpt-oss-120b, 1–4% cheaper on the output-heavy
workloads, worse on RAG whenever its pick was busy); the rest of the
gap between price-aware strategies is availability noise, not ranking. argrouter
also enforces a quality floor (≥ fp8, ≥ 99% uptime) that the price sort does not.
We have pre-committed to publishing the result when routing loses.
Install
pip install argrouter # core: httpx only
pip install "argrouter[server]" # optional /v1/decision sidecar
Python 3.10+. Fully typed, py.typed shipped.
How it fits with your existing gateway
argrouter does not reimplement the OpenAI wire format. Over half of the open issues on the largest gateway in this space are request-translation bugs; that is a maintenance burden with no upside for a routing project.
Instead it is a decision layer: a /v1/decision sidecar and plugins for gateways
that already own the bytes. We win the decision; your gateway keeps the
transport.
Supply chain
The dominant package in this category was compromised on PyPI, and that is a standing cost to everyone shipping here. argrouter commits to:
- One runtime dependency (
httpx) - PyPI Trusted Publishing with PEP 740 attestations — no long-lived token
- No install-time code execution
- A vendored, SHA-pinned price snapshot — no network call at import
Prior art and credit
The price catalog is seeded from OpenRouter's public models API, which co-locates pricing with quality indices, and from models.dev. The candidate-filtering and scoring shape is informed by vllm-project/semantic-router (Apache-2.0), which is the strongest open implementation of multi-factor selection. RouteLLM defined the evaluation vocabulary this project is measured in, and RouterArena provides the benchmark, scorer and price table the leaderboard result above uses.
This project was briefly published as thriftllm. It was renamed to avoid confusion with
ThriftLLM (Huang et al., 2025), an unrelated paper on
budget-constrained LLM ensemble selection, and with the thriftllm.com gateway.
License
Apache-2.0. Contributions under DCO
(git commit -s). See LICENSE.
Metadata
Release files for argrouter 0.1.0a2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| argrouter-0.1.0a2.tar.gz | 85.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| argrouter-0.1.0a2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 115.8 kB
Release files / argrouter-0.1.0a2.tar.gz
| Download URL | argrouter-0.1.0a2.tar.gz |
|---|---|
| Size | 85.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ef11605c3f0b5b06ef180bcdc6232c7bf108f71c158cabf520fa8cca807af373
|
|
BLAKE2b-256 checksum How to use checksums |
9b8443b5fcd2c19f013491d0ad58318aadc0d14a1c5063b31dca9cba326b19f8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / argrouter-0.1.0a2-py3-none-any.whl
| Download URL | argrouter-0.1.0a2-py3-none-any.whl |
|---|---|
| Size | 30.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
427dd9338537d7d3bc6dcb9621c55beed26b2dfd53ad316fba407242992868ad
|
|
BLAKE2b-256 checksum How to use checksums |
aea328c119748c088d08811bd384c96d134f0f5ce8eaedb30cda490d47a2855c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log