Skip to main content

gpu-broker

gpu-broker lets a chat assistant, an image generator and a video generator take turns on one graphics card, so a team can share a single GPU without anyone switching models by hand.

CI License: Apache-2.0 Python 3.12+

What it's for

You have one GPU. Your team chats with a local LLM all day; sometimes someone wants an image or a short video. The chat model and the image model don't both fit in the card's memory. Without gpu-broker, someone has to SSH in, stop the chat model, start ComfyUI, make the image, then remember to put the chat model back. Until they do, nobody can chat.

gpu-broker does that juggling for you:

  • One address for everything. Chat, image and video requests all go to gpu-broker, which knows what is loaded and what each request needs.
  • It takes turns on the card. When an image or video request arrives, it lets the chats in progress finish, stops the chat model, runs the job, and lines up whatever arrives meanwhile.
  • It puts the chat model back by itself. Once the card has been quiet for a couple of minutes, the chat model is loaded again, so the next chat answers straight away.
  • You can see what is happening. A web dashboard shows what is loaded, what is running, who is waiting, and a log of every switch.
  • Use it instead of ChatGPT or Claude (coming). One command points the coding and chat tools you already use at your own model, and one command puts them back.

It is for a home lab with one good GPU, a small team sharing one workstation, and agents or scripts that need several kinds of model (chat, image, video, 3D) from the same machine. It speaks the same chat API as OpenAI, so chat apps such as Open WebUI connect to it unchanged.

See it

The dashboard while the team chats: the chat model is loaded and answering requests The team is chatting. The chat model is loaded, and a typical answer starts within a fraction of a second; the live charts show the card's memory, load, power and temperature.

A video job running after the chat model was stopped, with chat requests waiting in the queue Someone asked for a video. gpu-broker stopped the chat model to make room; the video is rendering, and the chats that arrived meanwhile wait in the queue until it is done.

The model list, with buttons to borrow the GPU for an image or video model Every model it can run. From here you can borrow the whole GPU for hands-on work in ComfyUI; it is handed back to the chat model when you are done.

The event log: the chat model stopped for an image, then restored once the card went quiet The event log, newest first: the chat model was stopped to make an image, then loaded again on its own once nothing else needed the card.

Try it in 30 seconds

No GPU needed. The demo runs the real dashboard and API on a simulated graphics card, with a few simulated people using it. You need Python 3.12 or newer.

pipx install git+https://github.com/emergenthq-net/gpu-broker
gpu-broker demo

Or, without installing anything: uvx --from git+https://github.com/emergenthq-net/gpu-broker gpu-broker demo.

The dashboard opens in your browser (over SSH, or with --no-browser, open the link it prints instead). Within two minutes you will see someone ask for a video and the chat model step aside. gpu-broker demo --quiet leaves out the simulated people, so you can send your own requests (the demo prints a curl line to start from). Nothing real runs: no model is downloaded, and the "images" it makes are placeholders that repeat the prompt.

Use it instead of ChatGPT or Claude (coming)

Run gpu-broker connect, or click Connect apps on the dashboard, and the tools on your machine (your shell, Continue, Cline, Aider, Codex, Open WebUI) use your local model instead of a hosted one. You don't edit any settings: gpu-broker finds the tools and configures them. gpu-broker disconnect puts everything back the way it was.

gpu-broker already answers both the OpenAI and the Claude (Anthropic) API, so apps built for either work against it today; docs/drop-in.md has the details. The connect command that sets them up for you is coming next.

How a request flows

sequenceDiagram
    autonumber
    actor U as Chat user
    actor A as Script / agent
    participant B as gpu-broker
    participant L as LLM server
    participant C as ComfyUI
    U->>B: POST /v1/chat/completions (stream)
    B->>L: resident, slot free: forward directly
    L-->>U: tokens, streamed as generated
    A->>B: POST /v1/jobs {model: video, prompt}
    B->>B: queue the job (FIFO)
    B->>B: close the LLM pool, drain in-flight calls
    B->>L: stop the unit
    B->>C: run the video graph
    C-->>B: outputs
    B-->>A: job done, with output file URLs
    Note over B: queue idle for idle_restore_s
    B->>C: POST /free
    B->>L: start the unit, wait for /health
    U->>B: next chat is served directly again

Features

  • Queue and residency: one GPU worker runs jobs in order and swaps models only between jobs.
  • LLMs and ComfyUI on one card: an image or video job stops the LLM; idle restore brings it back.
  • Fast lane for people: interactive chat on the resident model skips the queue and streams token by token.
  • OpenAI-compatible: /v1/chat/completions and /v1/models, so existing chat UIs work unchanged.
  • Drop-in for ChatGPT and Claude: OpenAI and Anthropic SDKs work by changing only the base URL and key (docs/drop-in.md).
  • Substitution: an unknown, missing or too-large model runs on the best installed match, and the job says why.
  • Downloads: ask for a Hugging Face repo or a GitHub URL and it is fetched in the background.
  • Interactive ComfyUI sessions: borrow the whole GPU from the dashboard; it is returned when you go idle.
  • Three host drivers: systemd units, Docker containers, or systemd units inside Proxmox LXCs.

Compared with llama-swap

llama-swap is excellent if all you run is OpenAI-compatible LLM servers. gpu-broker is for the case llama-swap does not cover:

llama-swap gpu-broker
Swap between LLM servers on request yes yes
LLMs and ComfyUI share one GPU – yes
Queue with positions, job states and an event log – yes (SQLite + JSONL)
Substitution, with the reason reported – yes
Downloads by Hugging Face repo or GitHub URL – yes
Interactive GPU sessions for ComfyUI – yes
Concurrent calls to the resident LLM proxied up to slots, some reserved for people
Where model servers live processes it launches systemd units, Docker containers, Proxmox LXCs

If you only swap LLMs, use llama-swap. If one card has to serve chat and diffusion, use this.

Quickstart

gpu-broker is not on PyPI yet. Every route below starts from a clone:

git clone https://github.com/emergenthq-net/gpu-broker && cd gpu-broker

Docker compose

Needs Docker and the NVIDIA Container Toolkit. The compose file runs the broker, llama.cpp's llama-server and ComfyUI. The broker starts and stops the other two through the Docker socket; it never creates or removes containers.

cd examples/docker
mkdir -p conf data models/llama-3.1-8b && cp config.yaml ../catalog.yaml conf/
sudo chown -R 10001:10001 conf data          # the broker runs as uid 10001 and rewrites catalog.yaml
echo "BROKER_TOKEN=$(openssl rand -hex 24)" > .env
echo "DOCKER_GID=$(getent group docker | cut -d: -f3)" >> .env
pip install huggingface_hub                  # provides the `hf` CLI, for the one-off model fetch
hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
  --local-dir models/llama-3.1-8b
docker compose up -d --build

The dashboard is at http://localhost:8095/dash. It asks for the token from .env once. Image and video models also need their files in ComfyUI's model folders (the comfy-models volume). The comments in examples/catalog.yaml name the files each entry expects.

systemd

The model servers are systemd units on the same machine; create them as you normally would.

python3 -m venv .venv && .venv/bin/pip install '.[download]'
sudo mkdir -p /etc/gpu-broker /var/lib/gpu-broker /var/log/gpu-broker
sudo cp examples/config.yaml examples/catalog.yaml /etc/gpu-broker/
.venv/bin/gpu-broker check        # loads config + catalog, prints the driver and the unit allowlist
sudo BROKER_TOKEN=$(openssl rand -hex 24) .venv/bin/gpu-broker serve

To run it as a service, use examples/systemd/gpu-broker.service. Its TimeoutStopSec sits just above server.graceful_shutdown_s (default 10 s), the longest a stop waits for open connections such as a streaming chat; raise both together.

The broker needs to systemctl start/stop the units named in the catalog. Either run it as root, or set driver.sudo: true with a sudoers rule limited to exactly those units. It listens on server.host:server.port, which defaults to 127.0.0.1:8095.

Proxmox

The broker runs in its own LXC or VM; the model servers are systemd units inside other containers. Start from examples/config.proxmox.yaml and install host/gpu-broker-ctl on the host as the broker key's forced command, as described under Security model, with host/gpu-broker-gpu (its GPU reader) next to it.

Use it

T="Authorization: Bearer $BROKER_TOKEN"
# Chat. Served directly by the resident model when it is up; `"stream": true` streams tokens.
curl -s localhost:8095/v1/chat/completions -H "$T" -H 'Content-Type: application/json' \
  -d '{"model":"llama","messages":[{"role":"user","content":"hi"}]}'
# Any model as a job: the LLM stops, ComfyUI renders, and the LLM returns once the queue is idle.
curl -s localhost:8095/v1/jobs -H "$T" -H 'Content-Type: application/json' \
  -d '{"model":"wan2.2-5b","prompt":"a fox in snow","wait":true}'
# Image-to-video or image edit: send the image as base64 (or a data: URL).
curl -s localhost:8095/v1/jobs -H "$T" -H 'Content-Type: application/json' \
  -d "{\"model\":\"wan2.2-5b\",\"prompt\":\"the fox runs\",\"image\":\"$(base64 < fox.png | tr -d '\n')\"}"
curl -s localhost:8095/v1/status -H "$T"       # residency, queue, downloads, recent jobs

Your first catalog

The catalog lists what may be requested and how each model runs. Here is a minimal one, with one LLM and one ComfyUI model:

defaults:
  resident: my-llm              # held on the card whenever nothing else needs it
  idle_restore_s: 120           # empty queue for this long → make `resident` resident again
  session_idle_s: 900           # interactive ComfyUI sessions (dashboard)
  session_yield_s: 120
  session_max_s: 14400
  vram_total_mib: 24564         # your card
  vram_reserve_mib: 600         # headroom; models above total - reserve are never loaded
  background_requesters: [batch-agent]   # x-requester values that never skip the queue

models:
  my-llm:
    kind: llm
    runner: llm_unit            # an OpenAI-compatible server the driver starts and stops
    unit: llama-server          # systemd unit or container name
    endpoint: http://127.0.0.1:8080
    served_name: llama-3.1-8b-instruct
    vram_mib: 7500
    slots: 4                    # matches llama-server -np 4
    reserved_interactive: 1     # background callers get 3 slots; one stays free for people
    variants: {my-llm-precise: {temperature: 0.1}}   # extra model id with request overrides
    caps: [chat, code]
    quality: 60
    status: ready
    aliases: [llama]

  sdxl:
    kind: image
    runner: comfy               # runs a graph on the shared ComfyUI
    template: sdxl              # a builder in gpu_broker/templates/
    params: {ckpt: sd_xl_base_1.0.safetensors}
    vram_mib: 9000
    caps: [t2i]
    quality: 60
    status: ready

Substitution. If a request names an unknown model, one that isn't installed yet, or one that won't fit, the highest-quality ready model of the same kind whose caps cover the request runs instead. The response gives the substitute and the reason. If the model can be downloaded, the download is queued as well.

Input files. An entry that takes files declares them in inputs:

  • {image: required} for image-to-video or image editing
  • {image: optional} for an optional start frame, plus end_image: optional where the model can pin the last frame too
  • {frames: one_of, video: one_of} for a model that takes several views of a scene or one video of it (one_of: exactly one of those slots must be filled). frames can also say how many views the model uses: frames: {need: one_of, min: 2, max: 32} (other counts are a 400)

Only ComfyUI models with a template (single images) and exec models take inputs. Capabilities a model has only with an input go in image_caps (e.g. caps: [t2v], image_caps: [i2v] for an optional start frame), so a text-only job never demands them of a substitute.

A job sends image, end_image and video as base64 or a data: URL (or, if inputs.allow_urls is on, as <slot>_url), and frames as a list of base64 images. The broker checks everything at submit time (size caps, PNG/JPEG/WebP or MP4/MOV/WebM by magic bytes, the frame count, model fit) and answers 400 rather than queueing a job that cannot run. For a ComfyUI model it uploads the images as broker-<job id>-<slot>.<ext> just before the graph runs; ComfyUI cannot delete inputs over its API, so prune them with the path unit in examples/systemd/, which runs on each upload. <slot>_url reaches only public addresses and output files of the broker's own ComfyUI, unless inputs.url_allow_networks lists more. Substitutes for a job with files are only models that take those files, and a model that requires one is never picked for a job without it.

Command-line models (runner: exec)

Some models are a program, not a ComfyUI graph: image → 3D Gaussian splat tools, for example. The broker runs those as exec jobs: it evicts the resident LLM, frees ComfyUI's weights, hands the job's input files to the program, and returns the files it wrote.

The command is never in the catalog or the request. It lives in a recipe file that the host's administrator writes (/etc/gpu-broker/recipes/<name>.recipe). The catalog only names it:

  apple-sharp:
    kind: 3d
    runner: exec
    exec: {recipe: sharp, timeout_s: 660}   # >= recipe timeout_s + 10 s kill grace + 30 s
    inputs: {image: required}
    caps: [image_to_splat]
    vram_mib: 12000
    status: ready
# /etc/gpu-broker/recipes/sharp.recipe
argv=/opt/ml-sharp/.venv/bin/sharp predict -i {in_dir} -o {out_dir} -c {checkpoint} --no-render
checkpoint=/var/lib/gpu-broker/models/apple-sharp/sharp_2572gikvuh.pt
in_dir=/var/lib/gpu-broker/exec/in/{jid}
out_dir=/var/lib/gpu-broker/exec/out/{jid}
outputs=*.ply
timeout_s=600

outputs is one or more file name globs separated by spaces (outputs=result.mp4 run.log): the job's outputs are each glob's files sorted by name, in the order the globs are listed, each file once, so a recipe decides what outputs[0] is. Files whose names start with . are never outputs.

The input files land in in_dir as <slot>[-NN].<ext> (image.png, frames-00.png, ...), streamed from the staging directory. out_dir must end in /{jid}: only that folder is created, in a parent you create once (on Proxmox it takes the parent's owner, so ComfyUI's user can serve and prune it). The program runs with GPU_BROKER_JOB=<job id> in its environment, gets SIGTERM at the recipe's timeout_s and SIGKILL 10 s later, and afterwards every process still carrying that tag is killed (workers that left its process group included). The broker refuses to start when a catalog exec.timeout_s is shorter than all that plus 30 s, or timeouts.exec_clean_s shorter than the driver's clean plus 10 s (the driver reports its timings; the Proxmox driver adds its SSH connect timeout, and the 10 s covers the rest of reaching the host), checks it again before an exec job evicts anything, and before the job runs it waits (up to timeouts.exec_vram_s) until the card shows the entry's vram_mib free (by a GPU reading taken after the evictions returned). If a recipe may still be running after its job ended, or the broker restarts while one runs, it holds the GPU: no job runs until a clean it retries in the background every intervals.held_retry_s confirms the job gone, or an operator calls POST /v1/admin/gpu-held/clear (effective at once). The broker records each exec job's recipe (in the job store, not the request) at submit, and marks every other job as not exec, so a restart still cleans an exec job after its model has left the catalog. A direct chat never holds the GPU, and a model the catalog runs as exec always does. A job with no record (written by an older broker) whose model is gone, and that does not look like a chat, is held until an operator clears it. Finding the job's processes needs a readable /proc: on the local driver, run the recipe as the broker's own user (one that switches user is not tracked), and with /proc mounted hidepid the scan can only say "unknown" while the program's own process is still visible. The hold is shown in /v1/status and on the dashboard, and survives resume and restarts. Request keys listed in exec.params arrive as params.json; exec.choices ({param: [values]}) restricts a param to the listed values, so an unsupported one is a 400 at submit. The files matching outputs become the job's result. When out_dir is under ComfyUI's output folder (comfy.output_dir), each result also gets a ComfyUI /view URL. The format is documented in gpu_broker/drivers/recipes.py, and examples/recipes/sharp.recipe is a working example. The systemd driver runs recipes on the broker's machine and the Proxmox driver runs them inside a container (target=<ct>); the Docker driver does not run them. Job folders are not removed by the broker: install examples/systemd/gpu-broker-exec-prune@.{path,service} where they are written, which deletes them two hours after the job (how long results stay downloadable; raise its -mmin to keep them longer). Recipe files must use LF line ends.

Bundled templates. A catalog template names one of these graph builders. Some use nodes that stock ComfyUI does not ship; install those node packs on your ComfyUI first. "Stock" means the nodes ship with a current ComfyUI release.

template model family needs
sdxl single-checkpoint SD / SDXL stock ComfyUI
qwen_image Qwen-Image 2.1 ComfyUI-GGUF when unet is a .gguf file
chroma Chroma1-HD stock ComfyUI
flux2_klein FLUX.2 Klein (optional LoRA) ComfyUI-GGUF (UnetLoaderGGUF)
flux2_klein_edit FLUX.2 Klein 9B image edit (image required) as flux2_klein
qwen_edit Qwen-Image 2.1 image edit with prompt enhancer (image required) stock ComfyUI (plus ComfyUI-GGUF for a .gguf unet)
wan14b Wan 2.2 14B video; params.mode: t2v or i2v (image required) stock ComfyUI
wan5b Wan 2.2 5B video, optional start image stock ComfyUI
hunyuan HunyuanVideo 1.5 480p text-to-video stock ComfyUI
hunyuan_i2v HunyuanVideo 1.5 720p image-to-video (image required) stock ComfyUI
minimax MiniMax H3 video with audio; optional image / end_image frames stock ComfyUI
ltx25 LTX 2.5 text-to-video with audio ComfyUI-GGUF-Loader (LTXV25ModelsLoader, LTXV25AVDecode; verified at commit 142c614)

ltx25 takes four files in params: unet, clip, video_vae and audio_vae. Its defaults (97 frames at 768x512, 24 fps, 8 steps) took about 5 minutes on an RTX 4090.

A catalog entry may retune a template's request-level defaults with defaults: (for example defaults: {steps: 20} or defaults: {enhance: false} on a qwen_edit model). A request still overrides them: request > entry defaults > template defaults. Only keys the template reads as request options are accepted, so params remain the only way to choose files; any other key fails at catalog load.

The full annotated example is examples/catalog.yaml and the config is examples/config.yaml. Every config key and its default is in gpu_broker/settings.py; gpu-broker check validates both files.

API

Every route except /health and the dashboard page needs Authorization: Bearer $BROKER_TOKEN.

endpoint
POST /v1/chat/completions OpenAI-compatible. Interactive callers on the resident model are served directly; with stream: true the tokens stream as they are generated. Other calls are queued, and a queued call answered with stream: true arrives as a single SSE chunk. Broker details are in x_broker.
POST /v1/jobs {model, kind?, caps?, prompt?/messages?, image?, end_image?, frames?, video?, ...params, wait?, wait_s?}. Files are base64 or data: URLs (frames is a list; <slot>_url when enabled). Returns requested, resolved, substitution, queue_position and download; a malformed file, or one the model cannot take, is a 400.
GET /v1/jobs/{id} The job's state, the model it used, and its outputs.
GET /v1/models Ready LLMs and their variants (OpenAI format).
GET /v1/catalog, /v1/status, /v1/events?since=N The catalog, current residency and queue, and the event log.
GET /v1/gpu, /v1/metrics, /v1/stats, /v1/ui Dashboard data: GPU reading, live samples with job latency and tok/s, per-model stats, labels.
POST /v1/sessions, /v1/sessions/end Borrow the GPU for interactive ComfyUI, and give it back.
POST /v1/admin/quiesce, /v1/admin/resume Drain in-flight calls before a restart, and undo that.
POST /v1/admin/gpu-held/clear Lift a GPU hold once you have made sure the held exec job's processes are gone.
GET /health, GET /dash Liveness check (no auth) and the dashboard.

Two optional headers on chat requests:

  • x-requester labels the caller.
  • x-priority: interactive|background overrides the catalog's background_requesters.

How it works

  • One GPU thread, FIFO. Residency changes only between jobs, never mid-job.
  • Every LLM call holds a pool slot. Before a switch, the GPU thread closes the pool and waits for in-flight calls to finish. Afterwards it reopens the pool on the new resident model. A direct chat can only ever reach the model the pool names.
  • Priority. Background calls may fill slots - reserved_interactive slots, and people may use them all, so a chat never waits behind batch work.
  • Residency. Before an LLM starts, ComfyUI is told to POST /free. Before a ComfyUI job, the resident LLM is stopped. Health checks are re-run rather than trusted, because other operators may stop things.
  • Idle restore. After idle_restore_s with an empty queue, defaults.resident comes back.
  • Log. Every state change is a SQLite row and a JSONL line.

ARCHITECTURE.md covers the module layout and the invariants in detail.

Host drivers

driver.kind model servers are start/stop via GPU readings
systemd (default) units on this machine systemctl [--user], optionally sudo -n this machine's GPU (see Hardware)
docker existing containers docker start/stop over the socket from the broker container (nvidia-smi, or /sys for AMD)
proxmox systemd units inside LXCs SSH to a forced-command script on the host the host's GPU, read by host/gpu-broker-gpu

Catalog units are driver-neutral: unit: llama-server, or unit: {name: llama-server, target: 101}, where target is the Proxmox container id.

Hardware

GPU read through model servers
NVIDIA nvidia-smi CUDA builds (llama.cpp server-cuda, ComfyUI on CUDA PyTorch)
AMD the amdgpu driver's sysfs files and /proc/*/fdinfo; no ROCm tools needed ROCm or Vulkan builds both fine (compose.rocm.yaml)
Intel, Apple not supported yet

gpu: {vendor: auto, index: 0} (env BROKER_GPU_VENDOR, BROKER_GPU_INDEX) picks the card. auto is decided on first use:

  • nvidia-smi installed: NVIDIA. If it fails or hangs, it is retried once a second, all within timeouts.gpu_query_s of wall-clock time, and then reported as an error (remembered for that long, so callers get it at once rather than waiting through the retries again). The broker never falls back to an AMD card because nvidia-smi misbehaved; set vendor: amd for that.
  • nvidia-smi not installed: the first amdgpu card.
  • Both an NVIDIA and an amdgpu card: NVIDIA, and the log says the AMD card is not read.

serve starts even when no GPU can be read yet: the dashboard and /v1/gpu show why, and the sampler keeps retrying. The choice is made in the background: while it runs, /v1/gpu answers {"state": "probing"} at once instead of waiting on nvidia-smi. gpu-broker check prints the choice and one reading. On Proxmox the host script decides instead: GPU_VENDOR, GPU_INDEX, NV_TIMEOUT_S (each nvidia-smi call) and GPU_BUDGET_S (a whole reading, auto's choice and its retries included; keep it plus the broker's timeouts.ssh_connect_s under timeouts.gpu_query_s, default 8 + 10 < 20) in /etc/gpu-broker-ctl.conf, with host/gpu-broker-gpu installed next to gpu-broker-ctl. It decides auto once and caches the answer in /run/gpu-broker-ctl until the conf file changes.

Only VRAM used and total are required. On AMD the dashboard also shows utilisation (gpu_busy_percent), board power (power1_average, or power1_input on RDNA3), edge temperature and the graphics clock (freq1_input). A value the card does not report, or cannot report while runtime-suspended, is shown as "—"; so is any bracketed nvidia-smi value ([N/A], [Not Supported], ...).

Per-process VRAM on AMD comes from each DRM client's drm-memory-vram, counted once per client, read only from fds that link into /dev/dri/. Reading another user's /proc/<pid>/fd needs root or CAP_SYS_PTRACE. Without it, the processes the broker can read are still shown, and the dashboard says "per-process memory needs root or CAP_SYS_PTRACE" (procs_unreadable in /v1/metrics) instead of showing a list that looks complete.

Security model

The full threat model is in SECURITY.md. In short:

  • Token. A bearer token is compared in constant time. With none set, every call is refused and serve won't start. Secrets come from the environment, never the config file.

  • Bind address. 127.0.0.1 by default. Put TLS in front if you expose the broker.

  • Allowlist and validation. Drivers only touch units named in the catalog plus comfy.unit. Unit names, repository references, slugs and paths are validated, files stay under fixed roots, and nothing runs through a shell.

  • No SSRF by default. The broker only calls http(s) URLs from its own config and catalog. Fetching a caller's image_url is opt-in (inputs.allow_urls); see SECURITY.md.

  • Exec recipes are host configuration. A job names a recipe and carries files; the command, its paths and its timeout come from the recipe file, never from the request.

  • Proxmox: forced command, not a shell. The host pins the broker's SSH key to the script:

    command="/usr/local/sbin/gpu-broker-ctl",restrict ssh-ed25519 AAAA... gpu-broker
    

    The script accepts only unit, gpu, gpustream, download, comfy-link, exec-put and exec-run, for the <container>:<unit> pairs listed in /etc/gpu-broker-ctl.conf (ALLOW_UNITS) and the recipes in its RECIPES directory. It re-validates every argument and logs each call. With no config file it allows no unit and knows no recipe.

  • Dashboard. The dashboard page carries no data and runs under a strict Content-Security-Policy. The token is kept only in the viewer's browser.

FAQ

AMD / ROCm? Yes, see Hardware. Intel, Apple? Not yet: starting and stopping servers is vendor-neutral, but there is no GPU reader for them.

Multiple GPUs? Not yet. One broker manages one GPU, the one gpu.index selects.

Does it run models itself? No. It controls servers you already run (llama.cpp, vLLM, or any OpenAI-compatible server with a /health endpoint, plus ComfyUI) and decides which one holds the card.

Why not just run everything at once? VRAM. On a 24 GB card, an 8B LLM at Q4 with a long context takes about 7–8 GB, and a 5B video model at fp16 wants over 20 GB. They don't fit together, and partial offloading makes both slow.

Limitations

  • One NVIDIA or AMD GPU. GPU readings come from nvidia-smi or the amdgpu sysfs files, for the card gpu.index selects, and per-process VRAM by owner needs the host PID namespace; inside a plain container you get totals only.
  • AMD support is tested against fixture files written to the kernel's documented formats (tests/fixtures/amdgpu/), not yet against a live AMD card.
  • LLM servers must be OpenAI-compatible and answer GET /health with 200. The tok/s and time-to-first-token figures need llama.cpp's timings block; other servers still work but show no throughput.
  • Images and video run through ComfyUI only. Each model family needs a graph builder in gpu_broker/templates/ (Python, not a workflow JSON file).
  • Some templates need custom ComfyUI node packs, which the broker does not install. ltx25 needs ComfyUI-GGUF-Loader (verified at commit 142c614) for LTXV25ModelsLoader and LTXV25AVDecode; GGUF model files need ComfyUI-GGUF. Without them ComfyUI rejects the graph and the job fails. The full list is under Bundled templates.
  • One LLM at a time. An LLM and a ComfyUI model are never co-resident, even when both would fit.
  • Downloads land in the models root but aren't wired into ComfyUI automatically. Put the files in ComfyUI's model folders yourself. A downloaded model with no template stays needs_integration.
  • The Docker driver only starts and stops containers that already exist, and needs the Docker socket, which is root-equivalent on the host. It does not run exec recipes.
  • Exec on Proxmox copies each input file over its own SSH call, so a job with many frames pays one connection per frame. Files go in as JSON-embedded base64, so very large videos are better sent as video_url (with inputs.allow_urls).
  • One shared token, with no per-user accounts or rate limits.

Roadmap

  • A demand- and priority-aware scheduler, replacing strict FIFO for queued work.
  • GPU readings for non-NVIDIA cards, and more than one GPU per host.
  • Wiring downloaded files into ComfyUI from the API.

Contributing

Tests never touch a real GPU, host or network:

python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/pytest && .venv/bin/ruff check . && .venv/bin/mypy

See CONTRIBUTING.md for the house rules (layers, no magic values, new graphs and drivers).

License

Apache-2.0. See LICENSE.

Metadata

Release files for gpu-broker 0.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpu-broker 0.3.2
File Size Uploaded
gpu_broker-0.3.2.tar.gz 303.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpu-broker 0.3.2
File Interpreter ABI Platform
gpu_broker-0.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 487.3 kB

Release files / gpu_broker-0.3.2.tar.gz

Download URL gpu_broker-0.3.2.tar.gz
Size 303.6 kB
Tags Source
SHA-256 checksum
How to use checksums
81a17cdadc44c67a5c18971dac739224b82425aa2a78071a66c9836bf712f2d4
BLAKE2b-256 checksum
How to use checksums
fda463c651ebc5f187e9a04a474b8271c5e563c5fe46aa2f8db7dea1642f0236
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / gpu_broker-0.3.2-py3-none-any.whl

Download URL gpu_broker-0.3.2-py3-none-any.whl
Size 183.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8c5c887186204cf480b37341fe519eba822fc423744b06eae7d8346930e8c9f8
BLAKE2b-256 checksum
How to use checksums
1ecd887db2b16b52e54c87edfcbd156a3dfdff9ef2b360599ae860199104315f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.0

2 release files

0.4.0

2 release files

This release

0.3.2 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page