gpu-broker
TLDR: just install it
pip install gpu-broker # or: uvx gpu-broker setup
gpu-broker setup
It finds your GPU and model servers, writes the config, starts the service and opens the dashboard.
gpu-broker lets a chat model, an image generator and a video generator take turns on one graphics card, switching between them for you as requests arrive.
Who it's for
You, with one GPU and more work than fits on it. Your background agents keep a local LLM busy all day. You chat with it too. Now and then you want an image or a short video. The chat model and the image model don't fit in the card's memory together.
Without gpu-broker you stop the LLM, start ComfyUI, make the image, and restart the LLM by hand; your agents stall until you do. With it, you just send the request.
It also fits:
- A small team sharing one workstation. Everyone gets one address; nobody switches models by hand.
- Scripts and agents that need several kinds of model (chat, image, video, 3D) from one machine.
- A home lab with one good GPU.
What it does
- One address for everything. Chat, image and video requests all go to gpu-broker.
- Takes turns on the card. For an image or video job it:
- lets the chats in progress finish,
- stops the chat model,
- runs the job,
- queues whatever arrives meanwhile.
- Puts the chat model back by itself once the card has been quiet for a couple of minutes.
- Keeps people ahead of batch work. Your own chats skip the queue and stream token by token.
- Shows what is happening. A web dashboard shows what is loaded, what runs, what waits, and every switch.
- Speaks the OpenAI (Chat Completions and Responses) and Anthropic APIs, so Open WebUI, the official SDKs and the Codex CLI connect unchanged.
- Cloud first, local when it fails (optional). With
upstreams:routes,claude-*andgpt-*go to the real provider with the client's own key, and fall back to a local model on an outage, a hung connection or used-up quota; a circuit breaker sends them back when the provider recovers (docs/failover.md).
See it
The chat model is loaded and answering. The live charts show the card's memory, load, power
and temperature.
A video was requested. The chat model was stopped to make room; chats that arrive wait in the
queue until the video is done.
Every model it can run. You can borrow the whole GPU for hands-on work in ComfyUI; it goes
back to the chat model when you are done.
The event log, newest first: the chat model stopped for an image, then came back on its own.
Try it in 30 seconds
No GPU needed. The demo runs the real dashboard and API on a simulated card, with simulated users. You need Python 3.12 or newer.
pip install gpu-broker
gpu-broker demo
Or without installing: uvx gpu-broker demo.
- The dashboard opens in your browser. Over SSH, or with
--no-browser, open the printed link. - Within two minutes, someone asks for a video and the chat model steps aside.
gpu-broker demo --quietleaves out the simulated users, so you can send your own requests. It prints acurlline to start from.- Nothing real runs: no model is downloaded, and the "images" are placeholders.
Quickstart
pip install gpu-broker # add [download] to fetch models: pip install 'gpu-broker[download]'
gpu-broker setup # or first see what it would do: gpu-broker setup --dry-run
setup asks no questions. It:
- finds the GPU (NVIDIA or AMD) and its memory;
- finds your model servers: llama.cpp, vLLM, Ollama and ComfyUI on their usual ports, and the systemd units or Docker containers that run them;
- writes
config.yamlandcatalog.yamlfor what it found (the starter files if it found nothing), and a new API token inbroker.env(mode 600); - installs and starts the
gpu-brokerservice: a user service if your model servers aresystemctl --userunits, else a system service when it has root or passwordless sudo. The broker never runs as root. Otherwise it prints the exactservecommand, and runs it for you when you're at a terminal; - runs
check, waits until the broker answers, and opens the dashboard (not over SSH).
Running it again is safe: it keeps every file it finds, and the token. Details, flags and what each estimate means: docs/setup.md.
Manual setup
For full control, set it up by hand: pick how your model servers run.
| your model servers are | route |
|---|---|
| systemd units on this machine | systemd |
| Docker containers | Docker compose |
| systemd units in Proxmox LXCs | Proxmox |
systemd
Create the model servers' units as you normally would. Then:
sudo gpu-broker init # writes /etc/gpu-broker/{config,catalog}.yaml, creates its folders
sudoedit /etc/gpu-broker/catalog.yaml # your units, endpoints and model sizes
gpu-broker check # validates both files, prints the driver and the unit allowlist
sudo BROKER_TOKEN=$(openssl rand -hex 24) gpu-broker serve
initnever overwrites existing files (--forcedoes).--dirwrites elsewhere.- It listens on
127.0.0.1:8095(server.host,server.port). - The broker runs
systemctl start/stopon the catalog's units.- Run it as root, or set
driver.sudo: truewith a sudoers rule limited to those units.
- Run it as root, or set
- To run it as a service:
examples/systemd/gpu-broker.service.- Keep its
TimeoutStopSecaboveserver.graceful_shutdown_s(default 10 s).
- Keep its
Docker compose
This route needs a clone, for the compose files in examples/docker. It also needs Docker and
the NVIDIA Container Toolkit.
git clone https://github.com/emergenthq-net/gpu-broker && cd gpu-broker/examples/docker
mkdir -p conf data models/llama-3.1-8b && cp config.yaml ../catalog.yaml conf/
sudo chown -R 10001:10001 conf data # the broker runs as uid 10001 and rewrites catalog.yaml
echo "BROKER_TOKEN=$(openssl rand -hex 24)" > .env
echo "DOCKER_GID=$(getent group docker | cut -d: -f3)" >> .env
pip install huggingface_hub # provides the `hf` CLI, for the one-off model fetch
hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
--local-dir models/llama-3.1-8b
docker compose up -d --build
- It runs the broker, llama.cpp's
llama-serverand ComfyUI. - The broker starts and stops the other two through the Docker socket. It never creates or removes containers.
- The dashboard is at http://localhost:8095/dash. It asks once for the token from
.env. - Image and video models need their files in ComfyUI's model folders (the
comfy-modelsvolume).- The comments in
examples/catalog.yamlname the files each entry expects.
- The comments in
Proxmox
The broker runs in its own LXC or VM. The model servers are systemd units in other containers.
- Start from
examples/config.proxmox.yaml. - On the host, install
host/gpu-broker-ctlas the broker key's forced command (see Security). - Install
host/gpu-broker-gpu, its GPU reader, next to it.
Use it
# after setup the token is in broker.env (/etc/gpu-broker, or ~/.config/gpu-broker without root)
export BROKER_TOKEN=$(sudo sed -n 's/^BROKER_TOKEN=//p' /etc/gpu-broker/broker.env)
T="Authorization: Bearer $BROKER_TOKEN"
curl -s localhost:8095/v1/chat/completions -H "$T" -H 'Content-Type: application/json' \
-d '{"model":"llama","messages":[{"role":"user","content":"hi"}]}'
curl -s localhost:8095/v1/jobs -H "$T" -H 'Content-Type: application/json' \
-d '{"model":"wan2.2-5b","prompt":"a fox in snow","wait":true}'
More examples and every route: docs/api.md.
Connect your apps
gpu-broker connect --url http://gpu-host:8095 # or click Connect apps on the dashboard
gpu-broker connectfinds the tools on your machine (your shell, Continue, Cline, Aider, Codex, Open WebUI) and points them at your local model. No settings to edit.--dry-runshows the plan first;gpu-broker clientslists what was found and connected.gpu-broker disconnectputs everything back, from backups taken before the first change.
- Any other app built for the OpenAI (Chat Completions or Responses) or Anthropic API works by changing only the base URL and key.
- Details: docs/drop-in.md.
Use it from Claude, ChatGPT or Codex (MCP)
- With the
mcpextra (pip install 'gpu-broker[mcp]'), the broker is also an MCP server: assistants can generate and edit images, make videos and 3D scenes, and ask the local model. - Remote clients and ChatGPT connectors use
/mcpon the broker; apps that launch a command usegpu-broker mcp.gpu-broker connectregisters it in Claude Code, Claude Desktop and Codex. - Details: docs/mcp.md.
Documentation
| page | covers |
|---|---|
| docs/setup.md | gpu-broker setup: what it detects, what it writes, flags |
| docs/catalog.md | the catalog: models, substitution, input files, bundled ComfyUI templates |
| docs/exec-recipes.md | command-line models (runner: exec): recipes, timeouts, GPU holds |
| docs/api.md | every route, request fields, headers |
| docs/drop-in.md | using it in place of the OpenAI and Anthropic APIs |
| docs/failover.md | cloud first, local on failure: routes, the circuit breaker, what fails over |
| docs/mcp.md | the MCP server: tools, /mcp and gpu-broker mcp, registering it |
| docs/hardware.md | host drivers, supported GPUs, choosing the card |
| ARCHITECTURE.md | module layout and invariants |
| SECURITY.md, docs/threat-model.md | security model and threat model |
Command-line models
- Some models are a program, not a ComfyUI graph: image-to-3D tools, for example.
- gpu-broker runs them from a recipe file the host's administrator writes.
- A request names the recipe but never carries a command.
- Details: docs/exec-recipes.md. Example:
examples/recipes/sharp.recipe.
How it works
sequenceDiagram
autonumber
actor U as You (chat)
actor A as Your agent
participant B as gpu-broker
participant L as LLM server
participant C as ComfyUI
U->>B: POST /v1/chat/completions (stream)
B->>L: resident, slot free: forward directly
L-->>U: tokens, streamed as generated
A->>B: POST /v1/jobs {model: video, prompt}
B->>B: queue the job
B->>B: close the LLM pool, drain in-flight calls
B->>L: stop the unit
B->>C: run the video graph
C-->>B: outputs
B-->>A: job done, with output file URLs
Note over B: queue idle for idle_restore_s
B->>C: POST /free
B->>L: start the unit, wait for /health
U->>B: next chat is served directly again
- One GPU thread, fair order. Models switch only between jobs, never mid-job.
- Waiting jobs go interactive first, then by requester: whoever has used the least expected GPU time goes next.
- Each requester's jobs keep their order, and an idle caller cannot bank credit.
- A background job that has waited
scheduler.max_wait_s(default 1 h) counts as interactive. scheduler.policy: fiforestores strict arrival order.
- Every LLM call holds a pool slot. Before a switch, the pool closes and in-flight calls finish.
- People first. Background calls may fill
slots - reserved_interactiveslots; people may use them all. - Safe switches. ComfyUI frees its weights before an LLM starts; the LLM stops before a ComfyUI job.
- Health is re-checked, not assumed, since someone else may stop things.
- Idle restore. After
idle_restore_swith an empty queue,defaults.residentcomes back. - Log. Every state change is a SQLite row and a JSONL line.
- Replay.
gpu-broker replay /var/log/gpu-broker/events.jsonlruns your own log through the queue rules in simulated time and prints the wait per requester and the residency switches, next to what the broker really did. It contacts nothing; it is how scheduler changes are measured.
Compared with llama-swap
llama-swap is excellent if all you run is OpenAI-compatible LLM servers. gpu-broker covers what it doesn't:
| llama-swap | gpu-broker | |
|---|---|---|
| Swap between LLM servers on request | yes | yes |
| LLMs and ComfyUI share one GPU | – | yes |
| Queue with positions, job states and an event log | – | yes (SQLite + JSONL) |
| Substitution, with the reason reported | – | yes |
| Downloads by Hugging Face repo or GitHub URL | – | yes |
| Interactive GPU sessions for ComfyUI | – | yes |
| Concurrent calls to the resident LLM | proxied | up to slots, some reserved for people |
| Where model servers live | processes it launches | systemd units, Docker containers, Proxmox LXCs |
- Only swapping LLMs: use llama-swap.
- One card serving chat and diffusion: use this.
Security
Details: SECURITY.md. Threat model: docs/threat-model.md.
-
Token. A bearer token, compared in constant time.
- With none set,
servewon't start. - Secrets come from the environment, never the config file.
- With none set,
-
Bind address.
127.0.0.1by default. Put TLS in front if you expose it. -
Allowlist. Drivers touch only units named in the catalog plus
comfy.unit.- Names, references and paths are validated; files stay under fixed roots; nothing runs through a shell.
-
No SSRF by default. The broker calls only URLs from its own config and catalog.
- Fetching a caller's
<slot>_urlis opt-in (inputs.allow_urls).
- Fetching a caller's
-
Exec recipes are host configuration. The command comes from the recipe file, never the request.
-
Proxmox: forced command, not a shell. The host pins the broker's SSH key to one script:
command="/usr/local/sbin/gpu-broker-ctl",restrict ssh-ed25519 AAAA... gpu-broker- It accepts only
unit,gpu,gpustream,download,comfy-linkandexec-put/exec-info/exec-run/exec-clean. - Only for the
<container>:<unit>pairs inALLOW_UNITS(/etc/gpu-broker-ctl.conf) and the recipes inRECIPES. - It re-validates every argument and logs each call. With no config file it allows nothing.
- It accepts only
-
Dashboard. The page carries no data and runs under a strict Content-Security-Policy. The token stays in your browser.
FAQ
| question | answer |
|---|---|
| AMD / ROCm? | Yes. See docs/hardware.md. |
| Intel, Apple? | Not yet. Starting and stopping servers is vendor-neutral, but there is no GPU reader for them. |
| Multiple GPUs? | Not yet. One broker manages one GPU, the one gpu.index selects. |
| Does it run models itself? | No. It controls servers you already run (llama.cpp, vLLM, any OpenAI-compatible server with /health, plus ComfyUI). |
| Why not run everything at once? | VRAM. On 24 GB, an 8B LLM at Q4 takes 7–8 GB and a 5B video model at fp16 over 20 GB. Partial offloading makes both slow. |
Limitations
- One NVIDIA or AMD GPU.
- Per-process VRAM needs the host PID namespace; in a plain container you get totals only.
- AMD support is tested against fixture files in the kernel's documented formats, not yet a live card.
- LLM servers must be OpenAI-compatible and answer
GET /healthwith 200.- tok/s and time-to-first-token need llama.cpp's
timingsblock.
- tok/s and time-to-first-token need llama.cpp's
- Images and video run through ComfyUI only, with a graph builder per model family in
gpu_broker/templates/.- Some need custom node packs the broker does not install: see Bundled templates.
- One model at a time. An LLM and a ComfyUI model are never loaded together, even when both would fit.
- Downloads are not wired into ComfyUI. Put the files in ComfyUI's model folders yourself.
- A downloaded model with no template stays
needs_integration.
- A downloaded model with no template stays
- Docker driver: existing containers only; needs the root-equivalent Docker socket; no exec recipes.
- Exec on Proxmox copies each input file over its own SSH call, as base64.
- Many
framesmean many connections; send very large videos asvideo_url.
- Many
- One shared token, with no per-user accounts or rate limits.
Roadmap
- Soft preemption and bounded same-model batching for the queue.
- GPU readings for more vendors, and more than one GPU per host.
- Wiring downloaded files into ComfyUI from the API.
Contributing
Tests never touch a real GPU, host or network:
git clone https://github.com/emergenthq-net/gpu-broker && cd gpu-broker
python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/pytest && .venv/bin/ruff check . && .venv/bin/mypy
House rules (layers, no magic values, new graphs and drivers): CONTRIBUTING.md.
Changelog
- 0.4.0 (2026-10-03):
gpu-broker setup, one command from install to a running broker;gpu-broker init; docs split intodocs/. - 0.3.2 (2026-10-03): the source distribution is self-testing; unpack it, install it with
[dev], runpytest. - 0.3.1 (2026-10-03): shipped tests use neutral ids and paths; CI scans every release for private names.
Every release: CHANGELOG.md and GitHub Releases.
License
Apache-2.0. See LICENSE.
Metadata
Release files for gpu-broker 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gpu_broker-0.5.0.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gpu_broker-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.4 MB
Release files / gpu_broker-0.5.0.tar.gz
| Download URL | gpu_broker-0.5.0.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cd5725baf8259d30bfc98bd23b76aec5397e5686c0b20dbcbc7f8338d04a1df7
|
|
BLAKE2b-256 checksum How to use checksums |
4cc10af111734462df0fdda24dd3e2e02047e11127f6d747ed965e0b9c795941
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / gpu_broker-0.5.0-py3-none-any.whl
| Download URL | gpu_broker-0.5.0-py3-none-any.whl |
|---|---|
| Size | 346.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c2f8f66a355b8e58db48c62f7aa68e7e879ab378835a3bec6eb2d2f9f328925b
|
|
BLAKE2b-256 checksum How to use checksums |
e91a2d251a575f799c9cec25a3110bf8292b7696e21ac64fc92adf883c29adae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log