Skip to main content

modelrelay

One small API to call LLMs anywhere. Write your code once, against modelrelay, and decide by configuration where calls go:

  • any OpenAI-compatible API: OpenAI, OpenRouter, an internal LLM proxy;
  • a job-queue gateway (submit a request, then poll until it's done);
  • with a fixed API key, or client credentials that turn into a token that expires.

Anything specific to one company's gateway lives in a small private adapter package that plugs in. It never goes into your projects or into this library.

from modelrelay import llm

resp = llm.chat("Explain recursion in one sentence", model="gpt-4o")
print(resp.text)

Install

pip install modelrelay   # or the latest from git: pip install git+https://github.com/naruminho/modelrelay
modelrelay init      # once per machine: creates ~/.modelrelay/config.toml (OpenRouter template)
modelrelay show      # prints the config in use, secrets masked

Or skip the file editing: modelrelay serve and open http://127.0.0.1:8765/, the setup screen (see Setup screen). It creates the same file.

pip install can't create files in your home folder, so modelrelay init does it. Other templates: modelrelay init --template openai or --template gateway. init never overwrites an existing file.

Usage

from modelrelay import Relay

llm = Relay()  # reads the config (see below)

# You own the history: send whatever messages you want each time.
history = [
    {"role": "system", "content": "You are concise."},
    {"role": "user", "content": "Hi!"},
]
resp = llm.chat(history, model="gpt-4o", temperature=0.2)
history.append(resp.to_message())

# Images / PDFs go to the last user message
llm.chat("What's in these?", model="gpt-4o", files=["photo.png", "invoice.pdf"])

# Image generation through multimodal models
resp = llm.chat("A cat writing code", model="some-image-model")
resp.images[0].save("cat.png")

Streaming and status events

stream() yields events. The last one is always done, with the full response.

for ev in llm.stream(history, model="gpt-4o"):
    if ev.type == "delta":       # text as it's generated (when the transport can stream)
        print(ev.text, end="")
    elif ev.type in ("queued", "running"):   # job-based transports
        print(f"\rthinking... {ev.elapsed:.0f}s", end="")
    elif ev.type == "done":
        final = ev.response

The same code works everywhere. With a job gateway you get status updates instead of text deltas; llm.supports_text_stream tells you which one to expect.

Tool calling

tools = [{
    "name": "run_command",
    "description": "Runs a shell command",
    "parameters": {"type": "object", "properties": {"cmd": {"type": "string"}}, "required": ["cmd"]},
}]

resp = llm.chat(history, model="gpt-4o", tools=tools)
for call in resp.tool_calls:
    output = my_tools[call.name](**call.arguments)
    history += [resp.to_message(), {"role": "tool", "tool_call_id": call.id, "content": output}]

With tools_mode = "emulated", tools are described in the prompt and the answer is parsed back into resp.tool_calls. Use it for providers without native tool calling. Your code doesn't change.

Errors and tracing

modelrelay never hides a failure. It does not guess missing fields, turn unknown statuses into "running", or hand back half an answer as if it were complete. Every problem is an exception that inherits from ModelRelayError:

Exception When
ProviderError HTTP error, network error, or an error field in the response
UnexpectedResponse the response doesn't have the expected shape (the contract changed?); never retried
StreamInterrupted the stream ended before the provider finished; .partial holds what arrived
JobFailed / JobTimeout the job reported an error, or did not finish in max_wait_seconds
InvalidToolCall the model called a tool with arguments that aren't a JSON object; .raw and .response included
AuthError the token could not be obtained
PayloadTooLarge the request is above max_payload_mb; checked before sending
ConfigError invalid or missing configuration

Every error carries its context. It is printed with the message and available in error.context:

HTTP 504: upstream request timeout [status=504, method=POST, url=https://.../chat/completions,
trace_id=3f9c1a2b7d4e, elapsed=30.02, transport=OpenAICompatible, model=region1;gpt-4o]

error.body has the full, untruncated response. The original exception stays in __cause__.

Trace id. Each call gets a trace_id. It appears in logs, errors, events and on response.trace_id.

Logs. Set MODELRELAY_LOG=debug (or call modelrelay.enable_logging()) to see every HTTP call with its status and timing, token renewals, job status changes and each polling retry. Tokens and message contents are never logged.

What adapters can change

Adapters can change anything specific to a gateway, without touching modelrelay:

  • URLs, request body and response parsing: override the hooks in JobsTransport / OpenAICompatible, or subclass Transport and write complete() and stream() from scratch. See Private adapters and ADAPTER_GUIDE.md. Relay needs nothing else.
  • Headers and fixed body fields: extra_headers / extra_body in the config, or override headers_for() for per-request headers.
  • Status names: map them in parse_poll(), and raise on anything unknown.
  • Authentication: client_credentials options, or a TokenProvider of your own.
  • Anything the core doesn't model: response.raw keeps the original payload, and **params in chat() goes straight to the transport.

The public API (chat, stream, Response, Event, the errors) and the adapter hooks follow semantic versioning. A breaking change to either means a new major version.

Configuration

Configs live in one fixed folder, ~/.modelrelay/ (C:\Users\<you>\.modelrelay\ on Windows). No environment variable is needed. config.toml is the default. One file is usually all you need, since it can send chat() and stream() through different paths (see the gateway example). Extra files are optional profiles, for when you want to switch whole setups:

llm = Relay()                     # ~/.modelrelay/config.toml
llm = Relay(profile="openai")     # ~/.modelrelay/openai.toml  (modelrelay init --profile openai --template openai)
llm = Relay(config_path="somewhere/else.toml")

The lookup order is: config_path, then profile, then $MODELRELAY_CONFIG (optional, handy for CI), then ~/.modelrelay/config.toml. If no file is found, you get a ConfigError telling you to run modelrelay init; there are no hidden defaults. llm.config.source tells you which file was used. To build a config in code instead, use Relay(Config.from_dict({...})).

OpenRouter at home

base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[models]   # names used in code -> names the provider expects
"gpt-4o" = "openai/gpt-4o"
"gemini-2.5-flash" = "google/gemini-2.5-flash"

A gateway with a job API, an OpenAI-compatible proxy and expiring tokens (modelrelay init --template gateway)

base_url = "https://gateway.example.com/v1"
transport = "gateway:GatewayJobs"      # chat(): the job API, adapter in ~/.modelrelay/adapters/gateway.py
stream_transport = "openai_compatible" # stream(): the proxy; switch to "gateway:GatewayJobs" if it goes away
tools_mode = "native"                  # or "emulated"
max_payload_mb = 20
ca_bundle = "C:/certs/company-ca.pem"  # or verify_ssl = false (not recommended)

auth = "client_credentials"
[auth_options]
token_url = "https://identity.example.com/token"
token_field = "data.token"   # dotted path in the response
ttl_minutes = 30             # renewed 2 minutes before it expires
client_id = "..."            # or leave them out and set $MODELRELAY_CLIENT_ID /
client_secret = "..."        # $MODELRELAY_CLIENT_SECRET instead

[models]
"gpt-4o" = "region1;gpt-4o"

[transports."gateway:GatewayJobs"]   # each path can have its own base_url
base_url = "https://gateway.example.com/jobs-api"
poll_interval = 0.5
poll_max_interval = 5
max_wait_seconds = 900

[transports.openai_compatible]
base_url = "https://proxy.example.com/v1"

Both paths share the same token.

Several providers

One file can use more than one provider, e.g. text from DeepSeek and images from Google. Each [providers.<name>] holds a connection (the same keys as the top of the file: base_url, api_key, auth, transport...), and a model entry picks one with provider:

provider = "openrouter"                  # used by entries that don't say

[providers.openrouter]
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[providers.google]
base_url = "https://generativelanguage.googleapis.com/v1beta/openai"
api_key = "AIza..."

[models]
"text"  = "deepseek/deepseek-v4.1-flash"                            # openrouter
"image" = { provider = "google", model = "gemini-3.1-flash-image" }
  • Providers share nothing: a key is only ever sent to its own provider's base_url.
  • [apps.<app>.models] entries can pick a provider too.
  • Files without [providers] work as before: the top of the file is the only connection.

Roles and default params (e.g. reasoning effort)

Apps ask for a role; the config picks the model. Two roles cover most apps:

Role What the model must do Example
text read: conversation, tools, and the images the app sends (slide snapshots, pasted screenshots). If the app sends images, pick a model with image input. deepseek/deepseek-v4.1-flash, anthropic/claude-sonnet-5
image generate images (image output) google/gemini-3.1-flash-image

A role can be a table instead of a string. model is the provider model; every other key is a default request param, sent on every call that uses the role:

[models]
"text"  = { model = "deepseek/deepseek-v4.1-flash", reasoning_effort = "medium" }   # OpenAI-style
"deep"  = { model = "deepseek/deepseek-v4.1-flash", reasoning = { effort = "high" } } # OpenRouter-style
"quick" = { model = "deepseek/deepseek-v4.1-flash", reasoning = { enabled = false } } # no thinking
"image" = "google/gemini-3.1-flash-image"
  • Params sent by the app on a call win over the config (relay.chat(..., model="text", reasoning_effort="low")).
  • The same works in [apps.<app>.models]: a table there replaces the shared entry for that app.
  • modelrelay show --app <app> prints the params: text = deepseek/deepseek-v4.1-flash (reasoning_effort=medium).
  • modelrelay passes the params as they are; how much a model honors them depends on the model and on the provider behind it. Measured on OpenRouter with deepseek/deepseek-v4.1-flash (Sep 2026, 3 runs each, a hard math question): it reasons by default (~600 reasoning tokens, billed as output); effort moved that only a little and noisily (minimal ~560, high ~900); reasoning.max_tokens was not enforced; reasoning = { enabled = false } really turned it off (~2 s instead of ~7 s) but got 1 of 3 right instead of 3 of 3. Measure with your own prompts before relying on it.

Per-app models

Several apps can share one config (and one modelrelay serve) and still use different models. The provider, credentials and transports are configured once; each app only lists the model names it wants resolved differently. Everything it doesn't list comes from [models].

base_url = "https://openrouter.ai/api/v1"     # shared by every app
api_key_env = "OPENROUTER_API_KEY"

[models]                                       # the default for every app
"text"  = "anthropic/claude-sonnet-5"
"image" = "google/gemini-2.5-flash-image"

[apps.wotan.models]                            # wotan: a stronger text model, same image model
"text" = "anthropic/claude-opus-5-5"

[apps.sagadeck.models]                         # sagadeck: same text model, a better image model
"image" = "google/gemini-3-pro-image"

With that file:

App asks for wotan gets sagadeck gets any other app gets
text anthropic/claude-opus-5-5 anthropic/claude-sonnet-5 anthropic/claude-sonnet-5
image google/gemini-2.5-flash-image google/gemini-3-pro-image google/gemini-2.5-flash-image
openai/gpt-4o (not an alias) openai/gpt-4o openai/gpt-4o openai/gpt-4o

How an app says who it is:

llm = Relay(app="wotan")                         # every call from this Relay
llm.chat("hi", model="text")                     # -> anthropic/claude-opus-5-5
llm.chat("hi", model="text", app="sagadeck")     # one call as another app
llm.resolve_model("text")                        # -> "anthropic/claude-opus-5-5" (no request made)

Through modelrelay serve, send the X-Modelrelay-App header (Node example):

await fetch("http://127.0.0.1:8765/v1/chat/completions", {
  method: "POST",
  headers: { "Content-Type": "application/json", "X-Modelrelay-App": "sagadeck" },
  body: JSON.stringify({ model: "image", messages: [{ role: "user", content: "a lighthouse at night" }] }),
});
curl -s http://127.0.0.1:8765/v1/chat/completions -H "X-Modelrelay-App: wotan" \
  -H "Content-Type: application/json" -d '{"model": "text", "messages": [{"role": "user", "content": "hi"}]}'

Check what an app will get before running it:

$ modelrelay show --app wotan
config file: C:\Users\you\.modelrelay\config.toml
models for app 'wotan':
  text = anthropic/claude-opus-5-5   <- [apps.wotan.models]
  image = google/gemini-2.5-flash-image

Rules:

  • An app with no section, or a request without the header / app, uses [models]. Nothing changes for existing apps.
  • [apps.<app>] only accepts models. Anything else (another base_url, other credentials) is a different setup: use a profile for that (modelrelay serve --profile bank).
  • The config belongs to the machine, not to the app's repository: your laptop, the bank's server and CI can map the same app to different models without touching the app.

Deploying: the package ships with no config

pip install modelrelay installs no config file and no model names, on purpose: which provider, which credentials and which models are decisions of each machine. On a new machine:

pip install modelrelay
modelrelay init                        # creates ~/.modelrelay/config.toml (OpenRouter template)
modelrelay init --template gateway     # or: company gateway (job API + proxy + expiring tokens)
modelrelay show                        # check it (secrets masked)
modelrelay show --app sagadeck         # check what one app will get
modelrelay serve                       # optional: http://127.0.0.1:8765/v1 for non-Python apps

Every template has the [models] role aliases (text, image) and a commented [apps.*] example. On a shared server, IT can keep one central file and point $MODELRELAY_CONFIG at it.

Every option:

Key Default Meaning
transport openai_compatible transport for chat() (builtin: openai_compatible, jobs; or a plugin)
stream_transport same as transport transport for stream()
base_url OpenAI API root
auth static static, client_credentials or a plugin
api_key / api_key_env – / OPENAI_API_KEY key for static
auth_options {} token_url, token_field, ttl_minutes, refresh_margin_seconds, request_format (json/form), id_field, secret_field, extra_fields, client_id, client_secret (or client_id_env / client_secret_env to read them from other env vars)
models {} model name map (names used in code -> provider model); a value can be a table { model = "...", <param> = ... } with default request params such as reasoning_effort
apps.<app>.models {} per-app overrides of models (see Per-app models)
tools_mode native native or emulated (can also be set per transport)
verify_ssl / ca_bundle true / – TLS verification
timeout_seconds 120 HTTP timeout
max_payload_mb – refuse requests bigger than this before sending
headers {} extra headers on every request
transports.<name> {} options for one transport: base_url, tools_mode, extra_body, extra_headers, polling settings
providers.<name> {} extra connections, same keys as above (see Several providers)
provider – the provider for model entries that don't name one (default: the top of the file)

Tokens are sent as Authorization: Bearer <token>. When a request gets 401/403, a client_credentials token is renewed once and the request retried.

Private adapters

When a gateway doesn't match the defaults, an adapter translates its contract. An adapter is one Python file on your machine, next to the config:

~/.modelrelay/                     (C:\Users\<you>\.modelrelay\ on Windows)
├── config.toml                    transport = "gateway:GatewayJobs"
└── adapters/
    └── gateway.py                 class GatewayJobs(JobsTransport)

modelrelay init --template gateway creates both, with the adapter as a skeleton full of TODOs. ADAPTER_GUIDE.md walks through filling it in, step by step. It is written so an AI agent can follow it. A filled-in adapter looks like this:

# ~/.modelrelay/adapters/gateway.py
from modelrelay import JobsTransport, JobState, UnexpectedResponse, require
from modelrelay.transports import parse_completion

class GatewayJobs(JobsTransport):
    def submit_url(self, req): return f"{self.base_url}/start"
    def poll_url(self, job_id, req): return f"{self.base_url}/status/{job_id}"
    def build_submit(self, req): return {"flow": "llm", "inputs": {"model": req.model, "messages": req.messages}}
    def parse_submit(self, data): return require(data, "execution.id")   # raises if missing
    def parse_poll(self, data):
        state = require(data, "execution.state")
        if state == "DONE":
            return JobState("done", response=parse_completion(require(data, "output")), raw=data)
        if state == "ERROR":
            return JobState("error", error=require(data, "output.message"), raw=data)
        if state in ("QUEUED", "RUNNING"):
            return JobState(state.lower(), raw=data)
        raise UnexpectedResponse(f"Unknown state {state!r}", body=data)   # never guess

modelrelay does the polling, retries, timeouts, token renewal and error reporting. The adapter only says where to send, what to send and how to read the answers. mock_adapter.py is a complete working example.

In transport and auth, "file:Class" loads Class from ~/.modelrelay/adapters/file.py. Your projects stay clean: they only pip install modelrelay and call llm.chat(...). At home the same projects run with a different config.toml (for example OpenRouter) and no adapter at all.

Keep the adapter out of public repositories. If you want it versioned, put ~/.modelrelay/adapters/ in your company's git, never the config (it holds credentials).

Alternative: an installable package. If you'd rather ship the adapter as a package, register it with entry points and name it in the config:

# pyproject.toml of your private package
[project.entry-points."modelrelay.transports"]
my_gateway_jobs = "my_gateway_adapter.transport:GatewayJobs"

[project.entry-points."modelrelay.auth"]             # only if the token exchange is unusual
my_identity = "my_gateway_adapter.auth:MyIdentity"   # subclass ClientCredentials, override fetch_token()

Then transport = "my_gateway_jobs", and every project installs the package too.

Other languages: modelrelay serve

Programs that aren't Python (a Node app, a browser front end) can use modelrelay through a local OpenAI-compatible endpoint. They send plain /v1/chat/completions, and the config decides where each call really goes: gateway, token, adapters.

modelrelay serve                 # http://127.0.0.1:8765/v1  (--port, --host, --profile)
  • POST /v1/chat/completions: non-streaming and stream: true (SSE). Supports tools, multimodal input and generated images, which come back as message.images, the same format OpenRouter uses. Other fields such as temperature or max_tokens go straight to the provider.
  • X-Modelrelay-App: <app> picks that app's models ([apps.<app>.models] over [models]); without it, [models].
  • GET /v1/models lists the names the caller can use (with the header, that app's). GET /health always answers without auth.
  • It listens on 127.0.0.1 with no auth. --api-key or $MODELRELAY_SERVE_KEY requires Authorization: Bearer <key>.
  • A provider error comes back with its HTTP status for 4xx and as 502 otherwise, with the details in error.

In Python you can start it inside your own process: make_server(port=0) returns the server, and server.url gives its address.

Tip: give models role names in [models] ("text", "image") so apps ask for a role and each machine's config picks the actual model; add [apps.<app>.models] when one app needs something else.

Setup screen

modelrelay serve also serves a setup screen at http://127.0.0.1:8765/: add providers (OpenRouter, OpenAI, Google Gemini, DeepSeek, Anthropic, any OpenAI-compatible URL, or a company gateway with client credentials), paste keys, test the connection (it lists the provider's models), pick the model for each role and per app. Saving writes the config file (the previous one stays as config.toml.bak) and the running server uses it right away.

  • With no config yet, serve starts anyway so the screen can create one.
  • Keys never go back to the browser: a saved key shows as …a3f9.
  • It only answers requests addressed to this machine (Host: 127.0.0.1/localhost), and changes must be JSON from the same origin: other computers on the network and other websites open in your browser can't reach it.
  • Behind a login proxy (only admins should see it), pass the public address: modelrelay serve --public-url https://example.com/ia/. The page uses relative URLs, so any path prefix works.
  • Apps that called the server (X-Modelrelay-App) show up on the screen, ready to configure.
  • The model field is a searchable list of what the provider really has (loaded when the screen opens), showing what each model reads and makes (OpenRouter: "lê imagem", "gera imagem", price). The image role only lists models that generate images. A model the provider doesn't have can't be saved.
  • Testar on each row calls the model for real, with the row's provider and reasoning effort: a short message for text, a small image for the image role (it fails if no image comes back; costs a few cents). Models that passed are marked "testado" and come first in the list (remembered by the browser).

Mock gateway

A fake gateway to develop against without spending tokens: expiring tokens, a request time limit, a job queue with nested payloads, tool calls and image output.

python -m modelrelay.testing.mock_server --port 8000 --token-ttl 120
base_url = "http://127.0.0.1:8000/v1"
transport = "modelrelay.testing.mock_adapter:MockJobsTransport"
auth = "client_credentials"
[auth_options]
token_url = "http://127.0.0.1:8000/auth/token"
token_field = "data.token"

Any client id/secret is accepted; a static key test-key also works.

Development

python -m venv .venv
.venv/Scripts/pip install -e ".[dev]"   # or .venv/bin/pip
.venv/Scripts/pytest

examples/basics.py shows everything a project uses, written like a real project: chat, history, stream, tools, files and images, errors. Run it with python examples/basics.py.

pytest runs against the mock and costs nothing. To check a real provider end to end (chat, history, stream, native/emulated tools, image and PDF input, image generation, errors):

python examples/smoke_test.py                       # OpenRouter, needs $OPENROUTER_API_KEY
python examples/smoke_test.py my-config.toml --text-model X --vision-model Y --image-model Z

License

MIT

Release files for modelrelay 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for modelrelay 0.2.0
File Size Uploaded
modelrelay-0.2.0.tar.gz 75.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for modelrelay 0.2.0
File Interpreter ABI Platform
modelrelay-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 141.4 kB

Release files / modelrelay-0.2.0.tar.gz

Download URL modelrelay-0.2.0.tar.gz
Size 75.0 kB
Tags Source
SHA-256 checksum
How to use checksums
059544a2b115ec2df04b368da531dc328c65232169080fef77bce9e66bc48cfd
BLAKE2b-256 checksum
How to use checksums
970e39faa15a7f1fa3ba6f114c9a01d33353aab531b799f231fdc9e029f6dd7b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / modelrelay-0.2.0-py3-none-any.whl

Download URL modelrelay-0.2.0-py3-none-any.whl
Size 66.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2d25c2804c107f4dfbffc78f2944a4f0aa690599a6177c84a3cf5aa6120c6299
BLAKE2b-256 checksum
How to use checksums
b5906ec6d8e02fec95f53b6bef4133bfd596fc46427db67028859b2ba8e1400f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page