Router-Maestro
Router-Maestro is a local or self-hosted proxy that lets OpenAI-, Anthropic-, and Gemini-compatible clients use models from GitHub Copilot, OpenAI, Anthropic, and custom providers — with priority-based selection and automatic fallback.
TL;DR
Use every model in your GitHub Copilot catalog—including Claude, GPT, Gemini, Grok, and MAI—from Claude Code, OpenAI Codex, Gemini CLI, or any compatible client.
Router-Maestro acts as a proxy that gives you access to models from multiple providers through a unified API. Authenticate once with GitHub Copilot, and use its models anywhere that supports OpenAI or Anthropic APIs.
Features
Core
- Multi-provider support: GitHub Copilot (OAuth), OpenAI, Anthropic, and custom OpenAI-compatible endpoints
- Four ingress protocols: Anthropic Messages, OpenAI Chat Completions, OpenAI Responses, and Gemini generation all use the same dispatcher
- Full GitHub Copilot catalog across clients: Claude Code, OpenAI Codex, and Gemini CLI can select every model exposed by the live Copilot catalog; Router-Maestro chooses the compatible Messages, Chat, or Responses transport
- Gemini API compatibility: Gemini REST API format (
/api/gemini/v1beta/...) for Gemini CLI/SDK - Lazy cross-protocol translation: Native protocol pairs use a copy-on-write identity path; only cross-protocol attempts materialize the shared semantic IR
- Intelligent routing: Priority-based model selection with automatic fallback on failure
- Deterministic model matching: Public model IDs are provider-qualified, while convenient bare aliases (for example,
opus-4-6) are matched by score and routed according to the configured priorities - CLI management: Full command-line interface for configuration and server control
- Visual configuration: A loopback-only BIOS-style web portal for contexts, model catalogs, context windows, trusted projects, and Claude Code/Codex/Gemini configuration with preview, backup, and one-click apply
- Docker ready: Production-ready Docker images with Traefik integration
- Configuration hot-reload: Auto-reload config files every 5 minutes without server restart
Advanced
- 1M context support:
config claude-codeshows the live model catalog in a searchable selector, including every server-advertised context size, followed by a context-window selector. Choosing an advertised 1M tier adds Claude Code's client-side[1m]hint and raisesCLAUDE_CODE_AUTO_COMPACT_WINDOWfor the main model. Router-Maestro does not inject synthetic catalog entries or rewrite the selection to a dedicated upstream-1mmodel. - Capability-aware reasoning tiers:
reasoning_effortand Anthropic thinking controls stay on the selected base model. The ordered effort ladder isminimal < low < medium < high < xhigh < max. An advertised exact tier is passed through; otherwise Router-Maestro may substitute only the highest supported tier no greater than the request, or reject it when no such tier exists. Unknown tiers are rejected.minimalhas no implicit token-budget equivalent, so a smallthinking.budget_tokensvalue is not guessed to meanminimal. Copilot's known catalog-onlynonesentinel is preserved as model capability metadata, but it is not a client request tier, budget mapping, or downward-substitution target. Effort does not route through-highor-xhighmodel suffixes. - Anthropic adaptive effort passthrough:
output_config.effortis preserved across the stable route and its beta compatibility alias. Explicit effort replaces an adaptive thinking budget, while manualthinking.type="enabled"retains its protocol-requiredbudget_tokenswhenbudget_tokens < max_tokens. Omittingbudget_tokensuses the configured server default. The Messages runtime rejects present token limits unless they are positive, non-boolean integers. If an exact reasoning tier is unavailable, Router-Maestro may use the highest advertised tier that does not exceed the request. It never silently raises reasoning effort, cost, or latency; a request with no valid lower tier is rejected.
Table of Contents
Quick Start
Get a local server running in 3 steps. The server (started locally or via Docker with ~/.config/router-maestro mounted) auto-creates a local context with a generated API key — no manual context add is needed when client and server are on the same machine.
Before you start, make sure you have:
- Docker running locally (or skip to Local with pip install)
- Python 3 with
pipavailable for therouter-maestroCLI on the host - An active GitHub Copilot subscription
- Port 8080 free, or adjust
-p 8080:8080in the Docker command below
Custom port: If you use a different host port (e.g.
-p 8123:8080), update the local context so the CLI connects to the correct port:router-maestro context update local --endpoint http://localhost:8123
About the Router-Maestro API key. Router-Maestro has one server key (format
sk-rm-...) that clients must send on inference, administration, and remote-management requests. Public health/docs and the independently configured metrics endpoint are exceptions. It is not an OpenAI / Anthropic / Gemini / GitHub token — it only authenticates clients to your Router-Maestro server. The server auto-generates and persists this key on first start (in~/.config/router-maestro/contexts.jsonor its Docker-mounted equivalent), so you usually never type it by hand: the CLI reads it from the active context and theconfig claude-code/codex/geminiwizards write it into each tool's settings for you. The two times you do touch it explicitly are (1)router-maestro server show-keyto copy it into a rawcurlor environment variable likeROUTER_MAESTRO_API_KEY, and (2)router-maestro context add ... --api-key sk-rm-...when pointing a client machine at a remote server (see Deployment). If an authenticated request returns401, it almost always means the key it sent doesn't match what the server expects — re-runserver show-keyand compare.
https://github.com/user-attachments/assets/35f7c0f5-967a-4f93-aec8-c34b460a0032
One Copilot Catalog, Every Client
Router-Maestro loads the live GitHub Copilot model catalog instead of limiting each client to its native model family. The provider handler selects the compatible upstream transport, using an identity fast path for matching protocols and lazy semantic translation only when protocols differ.
| Client | Ingress protocol | GitHub Copilot models |
|---|---|---|
| Claude Code | Anthropic Messages | Full live catalog, including Responses-only GPT models |
| OpenAI Codex | OpenAI Responses | Full live catalog, including Claude, Gemini, Grok, and MAI models |
| Gemini CLI | Gemini generateContent | Full live catalog across Messages, Chat, and Responses transports |
The CLI wizard and router-maestro web both show the selected server's live
models and advertised context-window choices. Unsupported feature combinations
fail explicitly before provider I/O rather than silently dropping fields.
1. Start the Server (Docker)
docker run -d --name router-maestro \
-p 8080:8080 \
-v ~/.local/share/router-maestro:/home/maestro/.local/share/router-maestro \
-v ~/.config/router-maestro:/home/maestro/.config/router-maestro \
likanwen/router-maestro:latest
Both volumes are required:
.local/share/router-maestropersists GitHub Copilot OAuth tokens..config/router-maestropersists the auto-generated server API key (incontexts.json). Because this directory is shared with the host, the host CLI sees the samelocalcontext as the container — no extra setup needed.
If you want a fixed key for automation, add -e ROUTER_MAESTRO_API_KEY="sk-rm-...". The API key is the Router-Maestro server key, not an OpenAI, Anthropic, Gemini, or GitHub token. Every client, generated tool config, or raw API call must use the same key.
Confirm the server is up:
curl http://localhost:8080/health
# Expected: {"status":"healthy"}
Prefer running without Docker? See Local with pip install.
2. Authenticate with GitHub Copilot
Install the CLI on the host and run auth login against the local server. The OAuth device flow is hosted by the server; the CLI just renders the URL/code and polls for completion, so there is no need to docker exec into the container.
pip install router-maestro
router-maestro auth login github-copilot
# Follow the prompts in this terminal:
# 1. Visit https://github.com/login/device
# 2. Enter the displayed code
# 3. Authorize "GitHub Copilot Chat"
If you ever need the server API key (for example to paste into a raw curl):
router-maestro server show-key
3. Configure Your CLI Tool
The config commands read the endpoint and API key from the active context (local by default) and write them into the target tool's settings.
router-maestro config claude-code # Claude Code (Anthropic-compatible)
router-maestro config codex # OpenAI Codex (CLI / extension / app)
router-maestro config gemini # Gemini CLI
Or open the local configuration portal:
router-maestro web
The portal binds to 127.0.0.1:8765 and opens the system browser. It measures
each selected context through the public /health endpoint, then loads that
context's authenticated /api/admin/models catalog. Context API keys are never
shown; Copy Key retrieves one only for the clipboard action. Project-level
targets are discovered from Claude Code, Codex, and Gemini trusted-folder
stores, plus paths explicitly added to Router-Maestro. Explicit additions do
not modify any client's own trust policy. Codex project files can only set the
model, so the portal requires their selected context to match the
Router-Maestro provider already configured at user level.
Interactive terminals use a searchable model dropdown whose labels and table
show the server's context_window_options (for example, 272K / 1M). Claude
Code then asks for one of the contexts it can encode; older servers without the
new field retain the legacy context choices. If the live catalog contains no
Claude-family model, both user- and project-level wizards map Fable, Opus,
Sonnet, Haiku, and subagent roles to the available models; project mappings
override the user defaults only inside that project. The legacy
ANTHROPIC_SMALL_FAST_MODEL setting is removed. After all client-specific
choices, the final prompt asks whether generated model IDs should retain the
provider/ prefix.
For Copilot models, Router-Maestro derives these choices from CAPI's default and
long_context billing tiers. Raw usable prompt limits such as 922000 are
displayed as 1M, matching VS Code's model picker.
For Codex, also export the same key on the client because the generated config references ROUTER_MAESTRO_API_KEY:
export ROUTER_MAESTRO_API_KEY="sk-rm-..." # add to your shell profile
Done! Run claude, codex, or gemini and your requests route through Router-Maestro.
To smoke-test the full path without launching a client:
curl http://localhost:8080/api/openai/v1/models \
-H "Authorization: Bearer $(router-maestro server show-key)"
# Expected: JSON list of available models
Deploying to another machine or a VPS? See Deployment for the remote-Docker and Compose + Traefik + HTTPS setups.
Core Concepts
Model Identification
Models are identified using the format {provider}/{model-id}:
| Example | Description |
|---|---|
github-copilot/gpt-4o |
GPT-4o via GitHub Copilot |
github-copilot/claude-sonnet-4 |
Claude Sonnet 4 via GitHub Copilot |
openai/gpt-4-turbo |
GPT-4 Turbo via OpenAI |
anthropic/claude-3-5-sonnet |
Claude 3.5 Sonnet via Anthropic |
Model-list endpoints and successful response model fields use this qualified
form. The response value identifies the candidate that actually executed the
request, so it can change after a permitted fallback without becoming
ambiguous. Qualification uses catalog provenance rather than guessing from the
text of an upstream ID. A raw upstream ID may itself contain / and remains the
complete suffix: provider openrouter plus raw ID openrouter/auto is exposed
as openrouter/openrouter/auto. Only a catalog value explicitly marked as an
already-public ID is decoded once.
Bare model IDs and fuzzy aliases remain valid input conveniences, but they do
not select or lock a provider. This includes an exact raw alias that contains a
slash. Use the complete public provider/model-id returned by a model-list
endpoint when provider identity matters. Other slash-containing inputs are
treated as provider-scoped; an unknown provider prefix returns 404 instead of
falling back to a cross-provider fuzzy match.
Fuzzy matching: You don't need to type exact model IDs. Router-Maestro will fuzzy-match common variations:
| You type | Resolves to |
|---|---|
Opus 4.6 |
claude-opus-4-6-20250617 |
opus-4-6 |
claude-opus-4-6-20250617 |
claude-sonnet-4.5 |
claude-sonnet-4-5-20250929 |
anthropic/sonnet-4-5 |
Sonnet 4.5 via Anthropic only |
The highest-confidence fuzzy match wins. A date/version suffix is used only to break ties within the same normalized model family. Low-confidence or effectively tied cross-family matches are rejected as ambiguous instead of being resolved by an unrelated model's newer date.
Auto-Routing
Use the special model name router-maestro for automatic provider selection:
{"model": "router-maestro", "messages": [...]}
The router will try models in priority order and fall back to the next on failure.
Priority & Fallback
Priority determines which model is tried first when using auto-routing.
# Set priorities
router-maestro model priority add github-copilot/claude-sonnet-4 --position 1
router-maestro model priority add github-copilot/gpt-4o --position 2
# View priorities
router-maestro model priority list
Fallback triggers only after a retryable execution failure, such as a transport error, rate limit, retryable upstream status, or malformed upstream response before a streaming response is committed:
| Strategy | Behavior |
|---|---|
priority |
Try next model in priorities list |
same-model |
Try same model on different provider |
none |
Fail immediately |
Configure in ~/.config/router-maestro/priorities.json:
{
"priorities": ["github-copilot/claude-sonnet-4", "github-copilot/gpt-4o"],
"fallback": {"strategy": "priority", "maxRetries": 2}
}
An explicit provider/model-id remains the primary candidate. If it is absent
from the priority list, the configured priorities are still eligible after a
retryable execution failure, with the primary removed from duplicates. Static
capability mismatches and invalid or unsupported request options return the
entry protocol's native 400 response and do not switch models. Once a stream
has emitted its first provider chunk, a later failure is surfaced in that same
stream and Router-Maestro never replays the request on another candidate.
Streaming has a strict terminal contract. A stream succeeds only after an
explicit successful provider terminal; clean EOF without one is an
unexpected_eof, not success. Failures discovered before an SSE response is
returned use the entry protocol's non-2xx JSON error. After the HTTP response
has started, its status is already committed (normally 200), so an error,
incomplete result, or unexpected EOF is encoded as exactly one protocol-native
in-stream terminal instead. The standard Anthropic route may send a ping
before opening a slow upstream stream; a failure after that ping is therefore
post-commit even though no model content has arrived yet.
OpenAI Responses preserves the upstream response's business status. A native
Responses result with status: "incomplete", "failed", or "cancelled"
remains an HTTP 200 Responses object/event with that status; it is not
manufactured into completed. Transport failures and malformed provider
payloads remain errors. When a failed/cancelled Responses result is bridged to
Chat, Anthropic, or Gemini, whose non-stream response schemas cannot represent
that native status, Router-Maestro returns that entry protocol's error envelope.
Cross-Provider Translation
Router-Maestro separates the client protocol from the provider transport. Its generation dispatcher accepts Anthropic Messages, OpenAI Chat, OpenAI Responses, or Gemini and can select an upstream Messages, Chat, or Responses binding. This gives the current implementation a 4×3 transport matrix; it does not imply that every provider or model advertises all three transports.
For example:
# Use Anthropic API with OpenAI provider
POST /v1/messages {"model": "openai/gpt-4o", ...}
# Use OpenAI API with Anthropic provider
POST /api/openai/v1/chat/completions {"model": "anthropic/claude-3-5-sonnet", ...}
When ingress and upstream use the same protocol, Router-Maestro preserves the wire payload on a copy-on-write identity fast path and does not create semantic IR. A cross-protocol attempt decodes semantic IR lazily, at most once per request, and reuses the immutable result across eligible transport attempts. The model route plan chooses provider/model candidates only. For each model, the provider handler exhausts its compatible transports before model fallback can move to the next candidate; that transport switch does not consume the model fallback limit.
Accepted semantic options are either preserved, translated, or rejected; they
are not silently dropped. Unsupported options use the client's native error
shape (OpenAI, Anthropic, or Gemini) with HTTP 400. Reasoning-tier
substitution is downward-only across the ordered
minimal < low < medium < high < xhigh < max ladder. Unknown tier names are
rejected, and minimal is never inferred from a token budget because no
documented budget equivalent exists.
Omitted temperature remains omitted through Chat, Responses, and Gemini
translation. An explicit value, including 1.0, remains explicit. Copilot's
Chat transport accepts the explicit value, while Copilot's Responses transport
rejects every explicit temperature with an OpenAI-native HTTP 400 before
provider I/O. OpenAI Responses reasoning currently represents only
reasoning.effort; reasoning.summary and other sibling fields are rejected
with their exact parameter path instead of being ignored.
For Copilot, the dispatcher can try Messages, Chat, or Responses according to the ingress protocol and catalog capabilities. Consequently Claude Code can use a Copilot GPT model that is available only through Responses while continuing to call the stable Anthropic Messages endpoint. The former Router-Maestro beta Messages and Responses paths remain aliases of this dispatcher for one compatibility cycle; new CLI configuration writes stable URLs.
For standard Anthropic thinking requests, budget and reasoning validation use the capability and output-token snapshot of the same frozen route candidate that will execute the request. Validation never consults a different provider's same-named catalog entry and then silently changes candidates.
OpenAI Chat and Responses preserve refusals as typed refusal data, including
streaming deltas and assistant history. Anthropic has no equivalent content
block and maps refusal content to text. Gemini has no equivalent refusal part,
so a cross-protocol conversion to Gemini rejects it before provider I/O rather
than silently changing it into ordinary model text.
For OpenAI Chat streaming, stream_options: {"include_usage": true} emits a
final usage-only chunk with choices: [] immediately before [DONE].
Explicit false suppresses downstream usage; omitting stream_options keeps
Router-Maestro's legacy streaming shape. Invalid stream options are rejected
with an OpenAI-native 400 before the provider call.
Contexts
A context is a named connection profile stored on the client machine. It contains the endpoint URL and Router-Maestro server API key for one deployment, so the same CLI can manage local Docker containers, remote VPS deployments, and other Router-Maestro servers.
| Context | Use Case |
|---|---|
local |
Default context for router-maestro server start |
docker |
Connect to a local Docker container |
my-vps |
Connect to a remote VPS deployment |
# Add a context with the server API key from `server show-key`
router-maestro context add my-vps --endpoint https://api.example.com --api-key sk-rm-...
# Switch contexts
router-maestro context set my-vps
# All CLI commands now target the remote server
router-maestro model list
CLI Reference
Server
| Command | Description |
|---|---|
server start --port 8080 |
Start the server |
server status |
Show server status |
server show-key |
Show current context API key |
Authentication
| Command | Description |
|---|---|
auth login [provider] |
Authenticate with a provider |
auth logout <provider> |
Remove authentication |
auth list |
List authenticated providers |
Models
| Command | Description |
|---|---|
model list |
List available models |
model refresh |
Refresh models cache |
model priority list |
Show priorities |
model priority add <model> --position <n> |
Add or move a priority |
model fallback show |
Show fallback config |
Contexts (Remote Management)
| Command | Description |
|---|---|
context current |
Show current context |
context list |
List all contexts |
context set <name> |
Switch context |
context update <name> --endpoint <url> [--api-key <key>] |
Update context endpoint/key |
context add <name> --endpoint <url> --api-key <key> |
Add remote context |
context test |
Test connection |
Other
| Command | Description |
|---|---|
config claude-code |
Generate Claude Code settings |
config codex |
Generate Codex config (CLI/Extension/App) |
config gemini |
Generate Gemini CLI .env |
web |
Open the loopback-only local configuration portal |
API Reference
OpenAI-Compatible
# Chat completions — full curl example
curl http://localhost:8080/api/openai/v1/chat/completions \
-H "Authorization: Bearer sk-rm-..." \
-H "Content-Type: application/json" \
-d '{
"model": "github-copilot/gpt-4o",
"messages": [{"role": "user", "content": "Hello"}],
"stream": false
}'
# Responses
POST /api/openai/v1/responses
# List models
GET /api/openai/v1/models
POST /api/openai/beta/v1/responses remains a temporary compatibility alias
for the stable Responses path.
The list id and every successful Chat/Responses response model are
provider-qualified (provider/model-id). A fallback response reports the
candidate that actually served it.
Anthropic-Compatible
# Messages
POST /v1/messages
POST /api/anthropic/v1/messages
{
"model": "github-copilot/claude-sonnet-4",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello"}]
}
# Count tokens
POST /v1/messages/count_tokens
POST /api/anthropic/beta/v1/messages remains a temporary compatibility alias
for the stable dispatcher. Prefer /v1/messages or
/api/anthropic/v1/messages in new Claude Code configuration.
Admin
POST /api/admin/models/refresh # Refresh model cache
Gemini-Compatible
# Generate content (non-streaming)
POST /api/gemini/v1beta/models/{model}:generateContent
{
"contents": [{"role": "user", "parts": [{"text": "Hello"}]}]
}
# Stream generate content (SSE)
POST /api/gemini/v1beta/models/{model}:streamGenerateContent?alt=sse
{
"contents": [{"role": "user", "parts": [{"text": "Hello"}]}]
}
# Count tokens
POST /api/gemini/v1beta/models/{model}:countTokens
{
"contents": [{"role": "user", "parts": [{"text": "Hello"}]}]
}
Gemini's v1beta segment is the Gemini wire-protocol version, not a
Router-Maestro beta endpoint, so it is not deprecated with the RM beta aliases.
countTokens remains a separate operation using the existing native count or
estimator path; it is not one of the generation transports. Router-Maestro does
not currently register a Gemini-native provider binding, but the Gemini ingress
can target Messages, Chat, or Responses providers.
Configuration
File Locations
Following XDG Base Directory specification:
| Type | Path | Contents |
|---|---|---|
| Config | ~/.config/router-maestro/ |
|
providers.json |
Custom provider definitions | |
priorities.json |
Model priorities and fallback | |
contexts.json |
Deployment contexts | |
projects.json |
Projects explicitly added through the local portal | |
| Data | ~/.local/share/router-maestro/ |
|
auth.json |
Provider OAuth and API-key credentials | |
server.json |
Legacy server state; current server API keys are stored in contexts.json |
|
reasoning-capsule-keys.json |
Auto-generated owner-only reasoning replay key ring |
Reasoning Capsule Keys
Cross-protocol reasoning replay uses authenticated rmr1 capsules so opaque
provider state is never exposed as plaintext to another wire protocol. Without
environment configuration, Router-Maestro atomically creates and reuses
~/.local/share/router-maestro/reasoning-capsule-keys.json with mode 0600.
An invalid, unreadable, or corrupt key source fails server startup rather than
silently replacing the key.
Multi-instance deployments must set the same unpadded base64url-encoded 32-byte
key in ROUTER_MAESTRO_REASONING_CAPSULE_KEY on every instance. During
rotation, place comma-separated old keys in
ROUTER_MAESTRO_REASONING_CAPSULE_PREVIOUS_KEYS; they decrypt existing
capsules but never encrypt new ones. See docs/deployment.md
for the rollout order and fail-closed behavior.
Custom Providers
Add OpenAI-compatible providers in ~/.config/router-maestro/providers.json:
{
"providers": {
"ollama": {
"type": "openai-compatible",
"baseURL": "http://localhost:11434/v1",
"models": {
"llama3": {"name": "Llama 3"},
"mistral": {"name": "Mistral 7B"}
},
"options": {
"allow_unauthenticated": true
}
}
}
}
Custom-provider credentials are resolved in this order:
- A non-empty environment variable. By default its name is the provider ID in
uppercase with punctuation replaced by underscores, followed by
_API_KEY. - An API key saved in Router-Maestro's credential repository with
router-maestro auth login <provider>. - No credential, only when
options.allow_unauthenticatedis explicitlytrue. Anonymous requests do not include anAuthorizationheader.
For example, ollama uses OLLAMA_API_KEY and my-provider uses
MY_PROVIDER_API_KEY:
export OLLAMA_API_KEY="sk-..."
Set options.api_key_env to a valid environment-variable name when a provider
needs a different name. Provider definitions and their authentication
requirements are obtained from the active server, so the same login command
works for local and remote contexts. Only a local context may fall back to the
local providers.json while its server is unavailable.
The supported runtime options are api_key_env and
allow_unauthenticated. Router-Maestro preserves unknown option keys from
older providers.json files when loading and saving, but ignores them at
runtime. If one provider definition is invalid, it is skipped with a sanitized
diagnostic while other valid custom providers remain available.
Hot-Reload
Configuration files are automatically reloaded every 5 minutes:
| File | Auto-Reload |
|---|---|
priorities.json |
✓ (5 min) |
providers.json |
✓ (5 min) |
auth.json |
Requires restart |
Force immediate reload:
router-maestro model refresh
Metrics & Observability
Router-Maestro exposes a top-level Prometheus endpoint at /metrics with
HTTP request counters, request duration histograms, and request IDs on
responses via X-Request-ID. Streaming request durations are recorded after
the response body finishes. Generation attempts also expose bounded
entry_protocol, upstream_transport, conversion_mode, outcome, and
ir_materialized labels. Provider, model, and binding identities are kept out
of Prometheus and are available only in opt-in audit attempt records.
curl http://localhost:8080/metrics
By default /metrics is public. Set ROUTER_MAESTRO_METRICS_TOKEN to require
an independent metrics token:
ROUTER_MAESTRO_METRICS_TOKEN="metrics-secret" router-maestro server start
curl http://localhost:8080/metrics -H "Authorization: Bearer metrics-secret"
See docs/observability.md for scrape examples, metric labels, and troubleshooting guidance.
Deployment
Architecture
graph TD
Internet["🌐 Internet (HTTPS)"]
subgraph VPS
Traefik["Traefik (ports 80/443)\nAutomatic HTTPS · Let's Encrypt\nHTTP → HTTPS redirect"]
RM["Router-Maestro (port 8080)\nOpenAI / Anthropic-compatible API\nMulti-provider routing"]
end
Providers["LLM Providers\nGitHub Copilot · OpenAI · Anthropic"]
Internet -->|443| Traefik
Traefik -->|8080| RM
RM --> Providers
- Traefik — reverse proxy that handles TLS termination and auto-renews HTTPS certificates via Let's Encrypt. Only needed for public-facing deployments.
- Router-Maestro — the API server. Listens on port 8080, requires its API key for inference and administration requests, and routes inference to configured LLM providers. Health/docs are public; metrics has its own optional token.
Server and Client API Keys
Router-Maestro currently has one server API key. The same
ROUTER_MAESTRO_API_KEY protects inference routes and /api/admin/*; every
inference client and remote CLI management command must send that key. A
separate administrator key is not currently supported. Public health/docs and
the independently configured metrics endpoint are the exceptions described in
their respective sections.
You can provide the key explicitly with ROUTER_MAESTRO_API_KEY or router-maestro server start --api-key .... If you do not, the server generates a sk-rm-... key on first start and persists it in the local context inside contexts.json (the Docker image runs the same server start command, so the same behavior applies there). To read it later:
router-maestro server show-key # local install / inside the container
docker exec router-maestro router-maestro server show-key # remote Docker host (run over SSH)
docker compose exec router-maestro router-maestro server show-key # Docker Compose
Authentication (router-maestro auth login github-copilot) and config (router-maestro config claude-code / codex / gemini) always run from the client and use the active context's endpoint + key. They never need docker exec because the server hosts the OAuth device flow and exposes it via the admin HTTP API.
Local with pip install
If you would rather not use Docker, run the server directly on the same machine. The local context is auto-created on first start, so the host CLI works against localhost:8080 with zero context setup.
pip install router-maestro
router-maestro server start --port 8080 # leave running in this terminal
router-maestro auth login github-copilot # in a second terminal
router-maestro config claude-code # or: config codex / config gemini
For a fixed key, set ROUTER_MAESTRO_API_KEY before server start or pass --api-key.
Option A: Remote Docker (No HTTPS)
Use when: running on another machine on your LAN/VPN, or behind an existing reverse proxy (Nginx, Caddy, etc.) that handles TLS.
Prerequisites: Docker installed on the server host; SSH access to that host; the Router-Maestro CLI installed on your client machine (pip install router-maestro).
Step 1 — Start the container on the server host
docker run -d --name router-maestro \
-p 8080:8080 \
-v ~/.local/share/router-maestro:/home/maestro/.local/share/router-maestro \
-v ~/.config/router-maestro:/home/maestro/.config/router-maestro \
likanwen/router-maestro:latest
The server generates and persists an API key automatically. For a fixed key, add -e ROUTER_MAESTRO_API_KEY="sk-rm-...".
Step 2 — Read the server API key from the server host
ssh user@server-host docker exec router-maestro router-maestro server show-key
Copy the printed key for the next step.
Step 3 — Add the server as a context on your client machine
router-maestro context add my-server \
--endpoint http://server-host:8080 \
--api-key "sk-rm-..."
router-maestro context set my-server
router-maestro context test # verify endpoint + key
Step 4 — Authenticate with GitHub Copilot from the client
The auth command targets the active context, so this runs against the remote server over HTTP — no docker exec needed.
router-maestro auth login github-copilot
# 1. Visit the URL shown in this terminal
# 2. Enter the displayed code
# 3. Authorize "GitHub Copilot Chat"
Step 5 — Configure your CLI tool from the client
router-maestro config claude-code # or: config codex / config gemini
Step 6 — Verify
curl http://server-host:8080/health
# Expected: {"status":"healthy"}
curl http://server-host:8080/api/openai/v1/models \
-H "Authorization: Bearer sk-rm-..."
# Expected: JSON list of available models
Option B: Production (Docker Compose + Traefik + HTTPS)
Use when: deploying to a public-facing VPS with a domain name. Provides automatic HTTPS via Let's Encrypt with the Cloudflare DNS challenge.
Prerequisites:
- A VPS with Docker and Docker Compose installed
- A domain name (e.g.,
api.example.com) with DNS pointing to your VPS - A Cloudflare account managing your domain's DNS (for automatic HTTPS)
- The Router-Maestro CLI installed on your client machine (
pip install router-maestro)
Step 1 — Clone the repository on the VPS
git clone https://github.com/MadSkittles/Router-Maestro.git
cd Router-Maestro
Step 2 — Configure environment variables
cp .env.example .env
Edit .env with your values:
| Variable | Description | Example |
|---|---|---|
DOMAIN |
Your domain pointing to this VPS | api.example.com |
CF_DNS_API_TOKEN |
Cloudflare API token with Zone:DNS:Edit permission. Generate here |
abc123... |
ACME_EMAIL |
Email for Let's Encrypt certificate expiry notifications | you@example.com |
ROUTER_MAESTRO_API_KEY |
Optional fixed server API key. Leave blank, and do not set it in the shell running Docker Compose, to let the server generate and persist one. | sk-rm-... |
ROUTER_MAESTRO_LOG_LEVEL |
Log verbosity (DEBUG, INFO, WARNING, ERROR) |
INFO |
TRAEFIK_DASHBOARD_AUTH |
(Optional) Basic auth for Traefik dashboard. Generate with htpasswd -nB admin, then escape $ as $$ |
admin:$$2y$$05$$... |
Step 3 — Start the services
docker compose up -d
This starts both Traefik (reverse proxy) and Router-Maestro. Traefik will automatically obtain an HTTPS certificate for your domain.
Step 4 — Read the server API key from the VPS
docker compose exec router-maestro router-maestro server show-key
If you set ROUTER_MAESTRO_API_KEY in .env or in the shell running Docker Compose, this prints that key. Otherwise it prints the generated key stored in the server's mounted config.
Step 5 — Add the VPS as a context on your client machine
router-maestro context add my-vps \
--endpoint https://api.example.com \
--api-key "sk-rm-..."
router-maestro context set my-vps
router-maestro context test
Step 6 — Authenticate with GitHub Copilot from the client
router-maestro auth login github-copilot
# Targets the VPS through the active context — no docker compose exec needed.
# 1. Visit the URL shown
# 2. Enter the displayed code
# 3. Authorize "GitHub Copilot Chat"
Step 7 — Configure your CLI tool from the client
router-maestro model list # confirm models load from the VPS
router-maestro config claude-code # or: config codex / config gemini
For Codex, also export the same key on the client because the generated config references ROUTER_MAESTRO_API_KEY:
export ROUTER_MAESTRO_API_KEY="sk-rm-..." # add to your shell profile
Step 8 — Verify
curl https://api.example.com/health
# Expected: {"status":"healthy"}
curl https://api.example.com/api/openai/v1/models \
-H "Authorization: Bearer sk-rm-..."
# Expected: JSON list of available models
Remote Management
Contexts let you manage any Router-Maestro server (local or remote) from your local CLI:
# Add a remote server with the server API key from `server show-key`
router-maestro context add my-vps --endpoint https://api.example.com --api-key sk-rm-...
# Switch between servers
router-maestro context set my-vps # target remote VPS
router-maestro context set local # target local server
# Test the connection
router-maestro context test
# All commands now target the active context
router-maestro model list
router-maestro auth login github-copilot
Advanced Configuration
For additional deployment options, see docs/deployment.md:
- Alternative DNS providers (AWS Route53, DigitalOcean, GoDaddy, Namecheap, etc.)
- HTTP challenge setup (when DNS challenge is not available)
- Traefik dashboard configuration and security
- Complete environment variables reference
Stream Guards & Audit Tracing
Router-Maestro includes runtime stream protection and optional per-request tracing.
Stream Guards (enabled by default in priorities.json):
- Leak Guard — detects when Copilot-served Claude models emit internal protocol markup (control envelopes, XML tool calls) as plain text. Control envelopes abort the stream (client retries); invoke leaks are recovered into structured tool_use.
- Runaway Guard — aborts streams with degenerate generation patterns (infinite tiny fragments or excessive byte volume).
Configure in ~/.config/router-maestro/priorities.json:
{
"guards": {
"leak_guard": { "enabled": true },
"runaway_guard": { "enabled": true, "max_bytes": 10000000 }
},
"beta_strip": ["output-128k-*"]
}
The four protocol streaming encoders (OpenAI Chat, OpenAI Responses,
Anthropic, and Gemini) attach these guards to their stream processing. The beta
Anthropic URL reuses the same encoder. beta_strip is also live: matching
tokens are removed from the inbound anthropic-beta header before Copilot
transport, and remaining tokens are forwarded. The broader request context is
separate from this stream pipeline: it owns one immutable config/router
generation, request ID, audit, terminal outcome, and cleanup for streaming and
non-stream inference and token counting routes. Streaming cleanup finishes at
the final ASGI body frame, not when the endpoint returns a stream object.
Audit Tracing (opt-in, for debugging):
# Enable via env var
ROUTER_MAESTRO_TRACE=1 router-maestro server start
# Or in priorities.json
{ "audit": { "enabled": true } }
Each traced request writes a directory under
~/.local/share/router-maestro/traces/{request_id}/. Artifacts are lifecycle
records, not a fixed four-file bundle:
inbound.json— the client requestupstream.json,upstream_2.json, ... — upstream request observations in orderupstream_resp.json,upstream_resp_2.json, ... — upstream response observations, numbered independently in response orderoutbound.json— wire status, timing, and semantic terminal outcome
Some early failures have no upstream artifact, while fallback, authentication
retry, catalog, or token-count traffic may produce multiple attempt records.
The recognized Authorization, X-API-Key, and X-Goog-API-Key headers and
common credential-shaped payload keys are redacted before the trace is written
asynchronously. Treat traces as sensitive because prompts, model output, and
unrecognized application-specific headers can still contain private data.
For Docker, mount a volume to persist traces:
docker run ... -v ./traces:/home/maestro/.local/share/router-maestro/traces ...
License
MIT License - see LICENSE file.
Changelog
See CHANGELOG.md for release history.
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Local Integration Tests
The live-backend integration tests are local-only and are not part of GitHub
Actions. They start a local Router-Maestro server, reuse your existing
Router-Maestro config/auth files, and send requests to the real GitHub Copilot
backend. The suite covers model invocation paths only: OpenAI Chat, OpenAI
Responses, Anthropic Messages/count_tokens, Gemini generateContent/stream/countTokens,
tool calls, streaming, usage accounting, Anthropic thinking budgets and
output_config.effort, OpenAI reasoning_effort, Gemini-family API calls, and the
full Copilot model matrix by default. Admin endpoints are intentionally not
covered by these tests.
Prerequisites:
uv run router-maestro auth login github-copilot
Run them explicitly:
make integration-test
Optional overrides:
RM_INTEGRATION_MODEL=github-copilot/gpt-4o make integration-test
RM_INTEGRATION_TOOL_MODEL=github-copilot/gpt-4o make integration-test
RM_INTEGRATION_RESPONSES_MODEL=github-copilot/gpt-5.4-mini make integration-test
RM_INTEGRATION_MODELS=github-copilot/gpt-4o,github-copilot/claude-sonnet-4.5 make integration-test
RM_INTEGRATION_MAX_MODELS=8 make integration-test
RM_INTEGRATION_MAX_REASONING_MODELS=3 make integration-test
RM_INTEGRATION_MAX_REASONING_MODELS=0 make integration-test # full reasoning sweep
Deployed Claude/Codex Validation
The repository also includes a repeatable live-client runner for an already deployed test context. It fetches the context's current model list, runs Claude Code and Codex one-shot smoke cases, then verifies a fresh two-request recall session for every selected model. Codex receives a temporary version-matched model catalog; persistent client configuration is not changed.
make live-validation RM_LIVE_ARGS='--context remote-vm-hk --client all --phase all'
Use --provider, repeatable --model, --model-pattern, or --max-models to bound a canary.
Sanitized JSON/TSV and per-attempt logs are temporary unless --keep-logs or --output-dir is
specified. The automated recall check complements rather than replaces the interactive file/MCP
rounds in skills/router-maestro-live-validation/.
Request Preparation Benchmark
uv run python scripts/benchmark_request_preparation.py --iterations 1000
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file router_maestro-0.8.1.tar.gz.
File metadata
- Download URL: router_maestro-0.8.1.tar.gz
- Upload date:
- Size: 923.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e0cd3b558228a6287df8f9d62a17da36681d33c401eab816c41a3a8935172a8a
|
|
| MD5 |
59ecb50ed76a008d202ed3d5e9863158
|
|
| BLAKE2b-256 |
b6a37865b42418031ee31512296a0d08827408ffc6c9d6f1f2db6534b4f72b26
|
Provenance
The following attestation bundles were made for router_maestro-0.8.1.tar.gz:
Publisher:
release.yml on MadSkittles/Router-Maestro
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
router_maestro-0.8.1.tar.gz -
Subject digest:
e0cd3b558228a6287df8f9d62a17da36681d33c401eab816c41a3a8935172a8a - Sigstore transparency entry: 2621553953
- Sigstore integration time:
-
Permalink:
MadSkittles/Router-Maestro@c9ae6bea3e45ee132ace67c6c33beb3119edbd21 -
Branch / Tag:
refs/tags/v0.8.1 - Owner: https://github.com/MadSkittles
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c9ae6bea3e45ee132ace67c6c33beb3119edbd21 -
Trigger Event:
push
-
Statement type:
File details
Details for the file router_maestro-0.8.1-py3-none-any.whl.
File metadata
- Download URL: router_maestro-0.8.1-py3-none-any.whl
- Upload date:
- Size: 464.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8e86c67a33c4366f5872ecff57ad00d3f541a07fe9a89400125f3f4fe0632d3c
|
|
| MD5 |
83a3f6595023ba5a90e9efeb8769e41f
|
|
| BLAKE2b-256 |
cbefbbcfbb37cb23bcdd536f542e532e454cebd612cfd1710020ce42923e1a8f
|
Provenance
The following attestation bundles were made for router_maestro-0.8.1-py3-none-any.whl:
Publisher:
release.yml on MadSkittles/Router-Maestro
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
router_maestro-0.8.1-py3-none-any.whl -
Subject digest:
8e86c67a33c4366f5872ecff57ad00d3f541a07fe9a89400125f3f4fe0632d3c - Sigstore transparency entry: 2621553956
- Sigstore integration time:
-
Permalink:
MadSkittles/Router-Maestro@c9ae6bea3e45ee132ace67c6c33beb3119edbd21 -
Branch / Tag:
refs/tags/v0.8.1 - Owner: https://github.com/MadSkittles
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c9ae6bea3e45ee132ace67c6c33beb3119edbd21 -
Trigger Event:
push
-
Statement type: