llmroutes
Benchmark-driven router for OpenAI-compatible LLM endpoints. If you juggle several models, gateways, or API keys, you already have a routing table — it's just in your head, or in a stale doc. llmroutes measures your endpoints and routes every request by policy, with automatic fallback when one fails.
pip install llmroutes
# describe your endpoints once
cat > endpoints.toml <<'EOF'
[[endpoint]]
name = "fast-small"
base_url = "https://gateway.example.com/v1"
api_key_env = "GATEWAY_API_KEY"
model = "small-8b"
price_in = 0.20 # USD per 1M tokens
price_out = 0.40
tags = ["parallel"]
[[endpoint]]
name = "big-brain"
base_url = "https://gateway.example.com/v1"
api_key_env = "GATEWAY_API_KEY"
model = "large-400b"
price_in = 3.00
price_out = 9.00
tags = ["serial"]
EOF
llmroutes bench --config endpoints.toml
llmroutes list --config endpoints.toml
llmroutes run --config endpoints.toml --policy fastest --prompt "summarize this thread"
llmroutes run --config endpoints.toml --policy cheapest --tag parallel < draft.md
Zero dependencies, standard library only. Python 3.11+.
Policies
- fastest — lowest benchmarked p50 latency first.
- cheapest — lowest estimated cost first, from your
price_in/price_outand a documented chars/4 token heuristic. Endpoints without prices sort last. - balanced — mean of latency rank and cost rank.
Endpoints that error, time out, or return garbage are skipped with a stderr
note and the next candidate is tried. If everything fails, you get the full
list of errors and a non-zero exit. run prints the response body to stdout
and routing metadata (served by X in Ys, tokens …) to stderr, so it composes
with pipes. --quiet silences the metadata.
bench sends a few tiny ping probes per endpoint and stores p50 latency
plus success rate in results.json next to the config (override with
--results). Re-bench whenever your lineup changes; run refuses to guess
without fresh numbers.
Why not just use LiteLLM?
LiteLLM is a full proxy server you deploy and operate. llmroutes is a 300-line CLI you call per request — no server, no database, no config service. Different tool for a different job: pick this when you want routing as a shell primitive, not infrastructure.
Limitations
- Routing is by measured latency, price, and availability — not by answer
quality. Tag your strong models and select them with
--tagwhen quality is what matters. - Non-streaming
/chat/completionsonly. If you need tokens as they arrive, this isn't your tool.
License
MIT.
Metadata
Release files for llmroutes 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llmroutes-0.1.0.tar.gz | 9.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llmroutes-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 17.2 kB
Release files / llmroutes-0.1.0.tar.gz
| Download URL | llmroutes-0.1.0.tar.gz |
|---|---|
| Size | 9.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9394cbc7a1c332093c64bc3abcc1c0e95273ca740190d78b9970990ceb437921
|
|
BLAKE2b-256 checksum How to use checksums |
e548c9843a5c045cde2cb80e5bf29441230dd80f53e18d99a87564581424bdcd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.
Transparency logRelease files / llmroutes-0.1.0-py3-none-any.whl
| Download URL | llmroutes-0.1.0-py3-none-any.whl |
|---|---|
| Size | 8.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5cf1a871f305027c51095317416e6f54a882b6d4d5f7f0946023f6d5a610082a
|
|
BLAKE2b-256 checksum How to use checksums |
32ae24fe946386eb4506a77b9ee3c450ed74302f1b866e82597fba51fd298062
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.
Transparency log