Skip to main content

Provider-agnostic LLM routing: free-first fallback chains, host-supplied policy, per-tenant keys. Transport by LiteLLM.

Project description

oq-ai-router

Provider-agnostic LLM routing for Python. Free-first fallback chains, a routing policy supplied by your application, and per-tenant credentials — with the actual transport delegated to LiteLLM.

Extracted from a production application rather than designed in the abstract.

What this is, and what it deliberately is not

The hard-won lesson behind this package: most of an "AI layer" is already solved, and the part that isn't is the part specific to your application.

Concern Owner
The HTTP call, per-provider message shapes LiteLLM
Fallback execution, retries, cooldowns LiteLLM
Exception normalization across providers LiteLLM
Which models to try, in what order, for what kind of work your app, via a RoutingPolicy
Free-tier-first ordering your app — no library knows which providers you have free access to
Reading messy JSON out of free models this package
What to do when there is no key at all your app

So this is a thin policy layer, not a gateway and not a framework. It ships no routing table of its own — a library that hardcodes one app's taxonomy is not reusable, and adopting it would mean editing the library.

Install

pip install oq-ai-router

Use

Define your application's own categories and model preferences, register the policy once at startup, then call:

import oq_ai_router
from oq_ai_router import AiConfig, TablePolicy, complete_json

MY_POLICY = TablePolicy(
    # Free sources first, paid last. This ORDER is the free-first mechanism.
    provider_order=["omniroute", "openrouter", "groq", "anthropic"],
    models={
        "omniroute":  {"chat": ["auto/offline", "auto/cheap"]},
        "openrouter": {"chat": ["openrouter/free"]},
        "groq":       {"chat": ["llama-3.3-70b-versatile"]},
        "anthropic":  {"chat": ["claude-3-5-haiku-latest"]},
    },
)
oq_ai_router.set_default_policy(MY_POLICY)

result = complete_json(
    "chat",
    [{"role": "user", "content": "Reply as JSON: {\"ok\": true}"}],
    config=AiConfig(keys={"groq": "gsk-..."}),
)
result["data"]        # the parsed object
result["model_used"]  # which model actually answered

Categories are yours. One app's are explain | suggest | audit | concierge, another's are vision | json | prose; both route through this package with no edits to it. If table-driven ordering cannot express what you need (latency budgets, freshness windows, live quota telemetry), implement the two-method RoutingPolicy protocol directly.

Per-tenant keys

AiConfig is injected by the host — the package never reads env, a database or app settings on its own. Multi-tenant hosts pass a different AiConfig per request; env_config() is the zero-config path for single-tenant use.

Zero-credential operation

Providers may declare requires_key=False for a local gateway that holds the upstream credentials itself. The bundled omniroute entry does this, so a fresh self-hoster running OmniRoute gets a working AI chain with no API keys at all.

Verified 2026-07-24: a completion through oq_ai_router → LiteLLM Router → a keyless OmniRoute returned successfully.

Do not attempt to route a consumer AI subscription (Claude Pro/Max, ChatGPT Plus, Gemini One) through this or any proxy. Those OAuth flows are scoped to each vendor's own first-party clients; reusing them breaches the consumer terms and risks the account. Use metered API keys (every major provider has a free tier) or a genuinely free lane.

Design notes worth knowing

  • A Router is cached per (chain, credentials). LiteLLM's Router expects a static deployment list, but per-tenant keys resolve per request. Cache keys are hashed, so a credential cannot leak through a dict repr or a traceback.
  • Cooldowns are per Router, therefore shared only between tenants on the same credentials — which is the correct boundary, since separate keys have separate quotas.
  • Per-provider timeouts. A direct paid API should fail fast so the chain moves on; a local gateway draining free quota across upstreams legitimately takes longer. Measured, not guessed.
  • JSON validity is handled here, not by LiteLLM, which has no concept of "returned 200, body would not parse". Free models emit think-tags, code fences and trailing commas; parse_lenient copes, and unparseable output falls through to the next model.

Evals and cost guardrails

Model ids rotate and free models vary wildly, so "which model for which category" is an empirical question. oq_ai_router.evals answers it by measurement: run a corpus across every model in a chain and take the cheapest model that still passes — accuracy, cost and latency scored together.

Two corpora, both first-class. A clean synthesized corpus measures the happy path; the failures that actually matter arrive wearing real client data that can never be committed. So:

Lives Committed Defines
public evals/corpus/public/ yes the contract — what we promise
private outside the repo, via OQ_AI_ROUTER_PRIVATE_CORPUS never reality — what actually broke

Both run through one schema and one scorer, and results are reported separately: a green public score beside a red private score means the contract holds but reality has moved, which is the most actionable signal here. Private cases never expose their text or id in a report — only the failure kind.

Closing the loop is the point. When a private case fails, derive a synthesized twin with synthesize_case, read it yourself, and commit it to the public corpus. Over time the public corpus accumulates the shape of every real failure without absorbing a byte of real data. Redaction handles the identifier formats it knows (PAN, GSTIN, CIN, email, phone) and is a starting point for human review, never a security boundary.

oq_ai_router.budget adds the spend ceilings: a per-request cap (pushed upstream as X-OmniRoute-Budget where the gateway enforces it) and a per-tenant ledger for the slow bleed no single request would trip. Cost is recorded even when a cap is breached — money spent is spent, and dropping it would let violations repeat indefinitely.

Licence

Dual licensed: AGPL v3 or later for community use, plus a proprietary commercial licence for anyone who needs to keep modifications private or embed this in a closed-source product. See LICENSING.md.

Note the Affero clause: running a modified version as a network service obliges you to offer its source to that service's users.

A CLA must be in place before any outside contribution is accepted, because dual licensing only works while the copyright is wholly owned. Until then this repository does not accept contributions.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

oq_ai_router-0.1.0.tar.gz (44.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

oq_ai_router-0.1.0-py3-none-any.whl (41.2 kB view details)

Uploaded Python 3

File details

Details for the file oq_ai_router-0.1.0.tar.gz.

File metadata

  • Download URL: oq_ai_router-0.1.0.tar.gz
  • Upload date:
  • Size: 44.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.4

File hashes

Hashes for oq_ai_router-0.1.0.tar.gz
Algorithm Hash digest
SHA256 5a0ae3e759c0c67f775648b6ab04a542afa8200e42ea1dc568e7d5ffafd10c28
MD5 03b7f1de38aaec7c9c5f92e5c2ed9271
BLAKE2b-256 8685ea309bca12ae5c9d49b545809d000aa5f31c8449a01b59e8c1e15df9d7fb

See more details on using hashes here.

File details

Details for the file oq_ai_router-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: oq_ai_router-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 41.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.4

File hashes

Hashes for oq_ai_router-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 35d0c74ce5c4b1a2c1af991467063a79df008b6b432b7487c1ae80158a401cff
MD5 a5602728e4387e4b455dc9e3d7fda608
BLAKE2b-256 c00f7d8b873b77f654c6448c3e4306461f9c7bde2dbe84886994c9ac308494ba

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page