Skip to main content

llmroutes

ci

Benchmark-driven router for OpenAI-compatible LLM endpoints. If you juggle several models, gateways, or API keys, you already have a routing table — it's just in your head, or in a stale doc. llmroutes measures your endpoints and routes every request by policy, with automatic fallback when one fails.

pip install llmroutes

# describe your endpoints once
cat > endpoints.toml <<'EOF'
[[endpoint]]
name = "fast-small"
base_url = "https://gateway.example.com/v1"
api_key_env = "GATEWAY_API_KEY"
model = "small-8b"
price_in = 0.20    # USD per 1M tokens
price_out = 0.40
tags = ["parallel"]

[[endpoint]]
name = "big-brain"
base_url = "https://gateway.example.com/v1"
api_key_env = "GATEWAY_API_KEY"
model = "large-400b"
price_in = 3.00
price_out = 9.00
tags = ["serial"]
EOF

llmroutes bench --config endpoints.toml
llmroutes list --config endpoints.toml

llmroutes run --config endpoints.toml --policy fastest  --prompt "summarize this thread"
llmroutes run --config endpoints.toml --policy cheapest --tag parallel < draft.md

Zero dependencies, standard library only. Python 3.11+.

Policies

  • fastest — lowest benchmarked p50 latency first.
  • cheapest — lowest estimated cost first, from your price_in/price_out and a documented chars/4 token heuristic. Endpoints without prices sort last.
  • balanced — mean of latency rank and cost rank.

Endpoints that error, time out, or return garbage are skipped with a stderr note and the next candidate is tried. If everything fails, you get the full list of errors and a non-zero exit. run prints the response body to stdout and routing metadata (served by X in Ys, tokens …) to stderr, so it composes with pipes. --quiet silences the metadata.

bench sends a few tiny ping probes per endpoint and stores p50 latency plus success rate in results.json next to the config (override with --results). Re-bench whenever your lineup changes; run refuses to guess without fresh numbers.

Why not just use LiteLLM?

LiteLLM is a full proxy server you deploy and operate. llmroutes is a 300-line CLI you call per request — no server, no database, no config service. Different tool for a different job: pick this when you want routing as a shell primitive, not infrastructure.

Limitations

  • Routing is by measured latency, price, and availability — not by answer quality. Tag your strong models and select them with --tag when quality is what matters.
  • Non-streaming /chat/completions only. If you need tokens as they arrive, this isn't your tool.

License

MIT.

Metadata

Release files for llmroutes 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llmroutes 0.1.0
File Size Uploaded
llmroutes-0.1.0.tar.gz 9.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llmroutes 0.1.0
File Interpreter ABI Platform
llmroutes-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 17.2 kB

Release files / llmroutes-0.1.0.tar.gz

Download URL llmroutes-0.1.0.tar.gz
Size 9.2 kB
Tags Source
SHA-256 checksum
How to use checksums
9394cbc7a1c332093c64bc3abcc1c0e95273ca740190d78b9970990ceb437921
BLAKE2b-256 checksum
How to use checksums
e548c9843a5c045cde2cb80e5bf29441230dd80f53e18d99a87564581424bdcd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release files / llmroutes-0.1.0-py3-none-any.whl

Download URL llmroutes-0.1.0-py3-none-any.whl
Size 8.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5cf1a871f305027c51095317416e6f54a882b6d4d5f7f0946023f6d5a610082a
BLAKE2b-256 checksum
How to use checksums
32ae24fe946386eb4506a77b9ee3c450ed74302f1b866e82597fba51fd298062
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page