Skip to main content

ModelDispatcher

A reusable internal Python library that acts as a resilient AI Model Gateway/Router shared across applications.

Status: working library + demo. The core runs end-to-end (routing, fallback, quota, agent loop, onboarding), ships real OpenAI/Anthropic adapters, is covered by a behavioral test suite, and has an interactive FastAPI + React demo. See ARCHITECTURE.md for the design and demo/ to run it in a browser.

What it does

  • Strategy providers — every model backend implements one ModelProvider interface, so providers are hot-swappable.
  • Chain-of-Responsibility fallback — rate limits and exhaustion are intercepted and the request transparently escalates to the next candidate model.
  • Native agent orchestration — a small, dependency-free tool-calling loop with explicit state management (no heavy agent framework).
  • Triage & cost routing — cheap/free models for simple work, premium models reserved for complex reasoning.
  • Token-aware multi-tenant quotas — pre-flight reservation + post-call reconciliation per tenant.
  • Secure proxy perimeter — inbound validation and a credential-precedence chain.
  • Two-stage onboarding — zero-setup free tier by default; when limits are hit, a structured 402/429 handoff payload drives a GUI key wizard.

Install

pip install "model-dispatcher[openai,anthropic,gemini]"

Each provider adapter is an optional extra — install only the ones you key. Not yet published to PyPI (or need a version ahead of the latest tag)? Pin to a git ref instead:

pip install "model-dispatcher[openai] @ git+https://github.com/joka-7/ModelDispatcher@v0.2.0"

The TypeScript client (@joka-7/modeldispatcher-client) is published to GitHub Packages — see clients/typescript. It talks to your own backend, which is what runs the Python gateway above.

For an app with no backend at all — a pure browser app doing bring-your-own-key calls straight to a provider — see clients/browser-agent (@joka-7/modeldispatcher-browser-agent) instead: the same multi-provider, multi-key fallback idea (Gemini/OpenAI/Anthropic/Groq/Ollama, several pooled keys per vendor), running client-side with no server and no vendor SDK required.

Building the settings screen for that in React? See clients/react-ui (@joka-7/modeldispatcher-react-ui) for <ModelPicker> — add one or more providers, a model picked from a curated list per provider, pooled API keys, saving a favorite free AI app, and a link to the interactive docs/ai-glossary.html for first-time users, with nothing in it that navigates — plus <AskExternallyButton>, the separate action that actually opens that favorite from wherever the user is asking a question, and <PasteExternalReply> for apps that need the answer back in a specific structure (parsing it is the app's own job — this just captures the raw pasted text). So every app renders the same picker instead of each one hand-building its own, and a settings screen never redirects on its own.

Adopting either isn't all-or-nothing: resolveDispatcherFeatures from browser-agent gives each app's own developer — never the end user — two flags (ui, dispatch) to opt out per app during rollout instead of switching everything on at once. See docs/USAGE.md.

Quickstart

No API keys needed — this uses the keyless MockProvider:

pip install -e .            # from a clone of this repo
python examples/basic_agent.py
from model_dispatcher import (
    CompletionRequest, Message, ModelGateway, ProviderRegistry,
    Role, TenantContext, TenantId, TenantQuota,
)
from model_dispatcher.providers import MockProvider  # swap for OpenAIProvider, etc.

providers = ProviderRegistry()
providers.register(MockProvider("mock:free"))
gateway = ModelGateway.create(providers)  # build once at startup

tenant = TenantContext(
    tenant_id=TenantId("demo-user"),
    quota=TenantQuota(requests_per_min=20, tokens_per_min=40_000, tokens_per_day=1_000_000),
)
request = CompletionRequest(
    messages=(Message(role=Role.USER, content="Hello!"),),
    tenant=tenant.tenant_id,
)
result = gateway.dispatch(request, tenant)
print(result.final_message.content)

See examples/basic_agent.py for the full version with a tool the agent calls on its own.

Using it from another app

docs/USAGE.md is the integration guide: installing into a Python backend, wiring the TypeScript client to a frontend, mapping gateway errors onto HTTP responses, and pinning versions across multiple consuming repos.

Layout

See ARCHITECTURE.md for the directory layout, class blueprints, and algorithmic flows, and docs/HLD.md / docs/LLD.md for the design docs that stay current when behavior evolves past what's written there.

ModelDispatcher/
├── .github/
├── clients/            # Non-Python integration layers, documented in ARCHITECTURE.md's…
├── demo/               # Interactive end-to-end demo of the gateway
├── docs/
├── examples/
├── src/
├── templates/
├── tests/              # Behavioral test suite (routing, fallback, quota, agent loop, security,…
├── .ai                 # Ogen-ai submodule — the shared source of rules, skills and the ai-sync…
├── .dockerignore
├── .gitignore
├── .gitleaksignore
├── .gitmodules
├── AGENTS.md           # The compiled coding rules every AI assistant reads — generated, do not…
├── ARCHITECTURE.md     # ModelDispatcher — Architecture
├── CLAUDE.md           # Claude Code's copy of AGENTS.md (generated)
├── Dockerfile
├── GEMINI.md           # Gemini CLI's copy of AGENTS.md (generated)
├── LICENSE
├── README.md           # ModelDispatcher
├── ai-config.local.md  # Project-specific rules appended verbatim to the generated AGENTS.md
├── ai-config.toml      # Which rule fragments and target tools ai-sync compiles for this repo
└── pyproject.toml

Full annotated tree, every file: docs/STRUCTURE.md. Generated — regenerate after adding/renaming a file with:

python .ai/skills/repo_tree/gen_tree.py --project . --output docs/STRUCTURE.md
python .ai/skills/repo_tree/gen_tree.py --project . --output README.md --max-depth 1

Development

pip install -e ".[dev]"
ruff check src tests
mypy --strict src
pytest

Requires Python >= 3.11.

Try it in a browser

docker build -t model-dispatcher-demo .
docker run --rm -p 8000:8000 model-dispatcher-demo   # http://localhost:8000

The demo drives the real gateway through keyless mock providers, so you can watch routing, fallback, quota meters, and the key-wizard handoff without any API keys. See demo/README.md for the two-process dev setup.

Release files for model-dispatcher 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for model-dispatcher 0.6.0
File Size Uploaded
model_dispatcher-0.6.0.tar.gz 432.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for model-dispatcher 0.6.0
File Interpreter ABI Platform
model_dispatcher-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 505.3 kB

Release files / model_dispatcher-0.6.0.tar.gz

Download URL model_dispatcher-0.6.0.tar.gz
Size 432.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7a09b908a0ecc48c35098d9f5820e7d40cfc8eda8e2508563f3c998a59c91388
BLAKE2b-256 checksum
How to use checksums
c3174eb978a383c4bb72dc2c96d35e9686596bfbe9858cb5b340fb4b14a5686d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release files / model_dispatcher-0.6.0-py3-none-any.whl

Download URL model_dispatcher-0.6.0-py3-none-any.whl
Size 73.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c11bdeff8d35cbc80dd039540571a3a317b00006a7c33b2a71b25f1daa0781af
BLAKE2b-256 checksum
How to use checksums
8177d475ea3f202922a76be0bbd2d9d94f79a5f04fb1515889a9d2d64a838448
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

This release

0.6.0 This release

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page