lmrelay - a credentialed relay beside a local Ollama
If you work with Ollama, you run into this: by default it is reachable only from localhost,
and it has no built-in authentication. Connecting to Ollama from another machine usually means
changing its systemd configuration, or putting a reverse proxy in front of it. lmrelay solves
that. It installs with pip and runs as a daemon beside Ollama: it listens on a port of its own
and, when you want it to, requires a credential for access.
English | Español | Português | Français | Deutsch | Italiano | Русский | 中文 | 日本語 | हिन्दी | 한국어
Requirements
- Python 3.11 or higher, and four dependencies: FastAPI, starlette, uvicorn and httpx.
- Linux and macOS run every command, including
serve(detached) andenable— a systemd--userunit on Linux, a launchd agent on macOS, and a refusal where neither is installed. - Windows runs
runonly.servereports that the platform has noos.fork, andenablethat there is no systemd or launchd, rather than half-starting. - A local Ollama on 11434 is the default upstream, but it is not required. A relay with only
hosted providers configured is valid, as long as
default_upstreamnames one of them.
Installation
pip install lmrelay
Or the current main, which may be ahead of the release:
pip install git+https://github.com/wachawo/lmrelay.git
Quick start
lmrelay init # writes ~/.lmrelay/lmrelay.toml
lmrelay run # foreground, port 11435
Ollama keeps 11434 and its installation is left exactly as it is. Clients are repointed at 11435 instead. That is the trade: nothing about an existing Ollama has to change, and the relay is opt-in per client.
Auth is off in a fresh state, so on loopback this is a transparent proxy in front of Ollama. That is deliberate: a relay you have just installed should not lock you out of your own Ollama before you have a token. Point a client at 11435 and it works:
OLLAMA_HOST=127.0.0.1:11435 ollama list
Checking it works
Ask the relay for the model list. Either dialect will do; both reach the same Ollama:
curl http://127.0.0.1:11435/api/tags # Ollama's shape
curl http://127.0.0.1:11435/v1/models # OpenAI's shape
Then put a model to work. qwen3:8b here is whatever ollama list shows on your machine:
curl http://127.0.0.1:11435/api/generate -d '{
"model": "qwen3:8b",
"prompt": "Reply with exactly: it works",
"stream": false,
"think": false
}'
curl http://127.0.0.1:11435/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3:8b",
"messages": [{"role": "user", "content": "say ok"}]
}'
qwen3 reasons before it answers, and only Ollama's dialect has a switch for that: the "think": false above. Through /v1/chat/completions the reasoning arrives inside the content as a <think> block, because lmrelay forwards what the upstream produced and does not edit it.
With auth on, every one of these needs the credential:
curl http://127.0.0.1:11435/api/tags \
-H "Authorization: Bearer $LMRELAY_TOKEN"
Running it for real
lmrelay token gen --label laptop # printed once, never again
lmrelay auth true # now start requiring it
lmrelay enable # start at login, and start now
lmrelay status
lmrelay running (pid 40213), healthy
listening 127.0.0.1:11435
config /home/u/.lmrelay/lmrelay.toml
state /home/u/.lmrelay/state.json
upstreams anthropic, ollama, openai (default: ollama)
auth on, 2 tokens
autostart systemd: enabled, active
enable registers a systemd --user unit on Linux or a launchd agent on macOS, then starts
it. From then on stop, restart and reload go through that manager instead of the
pidfile, so the two cannot disagree about who owns the process. On a POSIX box with neither
manager, lmrelay serve runs the relay detached.
Usage
| Command | Does |
|---|---|
lmrelay init |
write ~/.lmrelay/lmrelay.toml |
lmrelay run |
run in the foreground |
lmrelay serve |
run detached, appending to lmrelay.log |
lmrelay stop |
stop the running relay |
lmrelay restart |
stop it, then start it detached again |
lmrelay reload |
re-read the config without dropping a connection |
lmrelay status |
what is running, where, with which upstreams |
lmrelay enable |
start at login, and start now |
lmrelay disable |
undo enable |
lmrelay auth true|false |
require a caller credential, or do not |
lmrelay token gen [--label L] |
mint a token and print it once |
lmrelay token add TOKEN [--label L] |
register a token you chose yourself |
lmrelay token list [--show] |
list tokens, masked unless --show |
lmrelay token delete ID |
remove one by the id token list prints |
lmrelay provider add NAME TOKEN |
add or rotate an upstream |
lmrelay provider list [--show] |
every upstream, from the file and from state |
lmrelay provider delete NAME |
remove a provider that state owns |
run, serve and restart take --host and --port. provider add takes --base-url,
--dialect and a repeatable --header K=V; with a known name — openai, anthropic,
deepseek, grok, ollama — the base URL, dialect and header shape come from a preset,
so lmrelay provider add openai sk-... is the whole command. --config PATH is accepted by
every command that reads the config or the state — that is, every command except init,
which always writes ~/.lmrelay/lmrelay.toml, and disable, which reads neither.
Choosing an upstream
The first path segment selects the upstream if and only if it exactly matches a key in
[upstream]. Otherwise default_upstream handles the request and the path is untouched.
POST /api/chat -> ollama /api/chat
POST /v1/chat/completions -> ollama /v1/chat/completions
POST /openai/v1/chat/completions -> openai /v1/chat/completions
POST /anthropic/v1/messages -> anthropic /v1/messages
POST /deepseek/v1/chat/completions -> deepseek /v1/chat/completions
POST /grok/v1/chat/completions -> grok /v1/chat/completions
So a client only has to learn the port once, and retargeting one at a different provider is a single line:
from openai import OpenAI
from anthropic import Anthropic
OpenAI(base_url="http://relay:11435/openai/v1", api_key=RELAY_TOKEN)
OpenAI(base_url="http://relay:11435/v1", api_key=RELAY_TOKEN) # Ollama
Anthropic(base_url="http://relay:11435/anthropic", api_key=RELAY_TOKEN)
curl http://127.0.0.1:11435/api/chat \
-H "Authorization: Bearer $LMRELAY_TOKEN" \
-d '{
"model": "llama3",
"messages": [{"role": "user", "content": "hi"}]
}'
GET /healthz answers {"status": "ok"} without touching an upstream and without a
credential. Everything else goes through the relay.
Compatibility
lmrelay forwards the method, path, query string and body bytes unchanged, and it does not translate between API dialects.
| Your client speaks | Path it uses | ollama | openai | deepseek | grok | anthropic |
|---|---|---|---|---|---|---|
| Ollama API | /api/chat, /api/generate, /api/tags |
yes | no | no | no | no |
| OpenAI API | /v1/chat/completions, /v1/models |
yes¹ | yes | yes | yes | no |
| Anthropic API | /v1/messages |
no | no | no | no | yes |
¹ Ollama serves an OpenAI-compatible surface at /v1/* alongside its native /api/*. This
is the practically important cell: an OpenAI-shaped client reaches all of ollama,
openai, deepseek and grok by changing only the path prefix.
The four cases that do not work, and the reason each one cannot be made to work, are in the configuration document.
Where lmrelay can tell that a path certainly does not exist upstream, it says so itself rather than letting the provider's 404 look like your mistake:
lmrelay: upstream 'anthropic' speaks the Anthropic API;
'/v1/chat/completions' is an OpenAI-dialect path. lmrelay
forwards requests unchanged and does not translate between
dialects.
Every error lmrelay generates begins with lmrelay: , so it is never mistaken for something
the provider said.
Configuration and Errors - the config file, caller tokens, providers, autostart, streaming behaviour, and what every error means.
Testing
pip install -e '.[test]'
pytest
Most of the suite drives the app in process against a recording upstream, so it needs no
network and no Ollama. tests/test_streaming.py is the
exception: it runs the relay under uvicorn in front of an upstream that answers a chunk at a
time, because the property it checks — that the caller has the first line before the
upstream has written the last — cannot be seen through an in-process client.
License
MIT License. See LICENSE.
Metadata
Release files for lmrelay 0.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lmrelay-0.0.3.tar.gz | 110.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| lmrelay-0.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 151.5 kB
Release files / lmrelay-0.0.3.tar.gz
| Download URL | lmrelay-0.0.3.tar.gz |
|---|---|
| Size | 110.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c97c5135f69cf10a8e30bf5cb1ea643c617747f879f390d7dd5c31cd1fa113b6
|
|
BLAKE2b-256 checksum How to use checksums |
3992ccdf3972e78a2de1545b7259de775e415b0daf0aa0a2c658b723ca1dbde2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency logRelease files / lmrelay-0.0.3-py3-none-any.whl
| Download URL | lmrelay-0.0.3-py3-none-any.whl |
|---|---|
| Size | 41.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
be7580412fba1a86da80ae762ea17e83d46d2c859fdf626e640447112de4de99
|
|
BLAKE2b-256 checksum How to use checksums |
798fbbec64a73523f68bb8b1270b135cd18ee6cf678bd1baeca908266b93f67c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency log