Skip to main content

inferhost

Self-hosted, multi-modal AI server for your own GPU — one command, nothing to compile.

PyPI Python License Docs

Chat LLMs · Vision · Text-to-speech · Image generation — all behind one OpenAI-compatible endpoint.

inferhost dashboard

inferhost turns any GPU machine into a private, local AI inference server. It wraps llama.cpp and stable-diffusion.cpp behind a single OpenAI-compatible API, pulls the official upstream binaries for your hardware (no compiling, no CUDA toolkit), auto-downloads the right model files when you paste a Hugging Face link, and hot-swaps models in and out of VRAM so one GPU can serve a large language model, a text-to-speech voice, and an image-generation pipeline side by side. Everything is driven from a keyboard dashboard (TUI) or a headless CLI, configured through a single optional .env file.

If you are searching for a self-hosted alternative to cloud AI APIs — a local LLM server with vision (multimodal) support, speech synthesis, and Stable Diffusion / Flux image generation on the same OpenAI-compatible gateway — inferhost is one pip install away.

Quick start

uv tool install inferhost      # or:  pipx install inferhost  /  pip install inferhost
inferhost                      # opens the dashboard — press 'a' to add a model

That is the whole setup. The first launch fetches the runtime binaries automatically. To add a model, press a and paste a Hugging Face repo — inferhost lists the files, recommends the best quantization for your VRAM, downloads what is needed (including companions such as vision projectors, vocoders, VAEs, and text encoders), and serves it. Then call it like OpenAI:

inferhost quick start: chat, speech, and image generation on one endpoint
Chat / LLM Text-to-speech Image generation
paste
Qwen/Qwen2.5-7B-Instruct-GGUF
paste
OuteAI/OuteTTS-0.2-500M-GGUF
paste
OlegSkutte/sdxl-turbo-GGUF
# Chat  ->  /v1/chat/completions
curl http://localhost:9001/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"<name-from-dashboard>","messages":[{"role":"user","content":"Hello!"}]}'

# Speech  ->  /v1/audio/speech   (returns WAV)
curl http://localhost:9001/v1/audio/speech -H 'Content-Type: application/json' \
  -d '{"model":"<name>","input":"Hello from inferhost.","voice":"default"}' --output speech.wav

# Image  ->  /v1/images/generations   (returns base64 PNG)
curl http://localhost:9001/v1/images/generations -H 'Content-Type: application/json' \
  -d '{"model":"<name>","prompt":"a red apple on a table","size":"512x512"}' \
  | jq -r '.data[0].b64_json' | base64 -d > out.png

Everything lives on http://localhost:9001/v1 — point any OpenAI client at it (OpenAI Python SDK, LangChain, Open WebUI, LibreChat, Continue, Cursor, or your own app). The model name is whatever shows in the dashboard.

Why inferhost

  • One endpoint, every modality — chat, vision, speech, and images on the same OpenAI-compatible :9001. No per-model servers to wire up.
  • Nothing to compile — official llama-server / sd-server binaries are pulled from upstream for your hardware: NVIDIA and AMD GPUs (Vulkan), AMD ROCm, Intel SYCL, Apple Silicon (Metal), and plain CPU.
  • Paste a link, it figures out the rest — recommends the best GGUF quantization for your VRAM, handles multi-part (sharded) GGUFs, and for multi-file image models (Flux, Z-Image, Qwen-Image) auto-downloads the matching VAE and text encoders from known-good repos.
  • One GPU, many models — llama-swap lazy-loads and hot-swaps models in and out of VRAM on demand, so a 24 GB card can serve a 27B LLM and Flux image generation without manual juggling.
  • TUI or headless — drive everything from a keyboard dashboard, or run inferhost start/stop/status on a remote server with no terminal attached.
  • Fast by default — speculative decoding (DFlash draft models, MTP/NextN, n-gram), KV-cache quantization, MoE expert CPU offload, and honest context windows, all tuned automatically and overridable per model or from .env.
  • Private by design — models, weights, and prompts never leave your machine. No accounts, no telemetry, no cloud dependency.

Supported models

Modality Models How
Chat / LLM any GGUF language model — Qwen, Llama, Gemma, DeepSeek, Mistral, and more, including 2-bit ternary builds (Ternary Bonsai) and sharded multi-part GGUFs paste repo, pick quant
Vision (multimodal) any GGUF vision model with an mmproj projector — Qwen3-VL, DeepSeek-OCR, Gemma vision variants projector auto-detected and downloaded
Speech (TTS) OuteTTS, Qwen3-TTS paste repo (vocoder auto-detected, or pick the Text-to-speech kind explicitly)
Image — single-file Stable Diffusion 1.5, SDXL (incl. Turbo) paste repo, pick file
Image — Flux.1 schnell / dev auto-fetches VAE + CLIP-L + T5XXL
Image — Flux.2 Klein incl. Bonsai-Image (1-bit) auto-fetches VAE + Qwen3-4B
Image — Z-Image Z-Image-Turbo auto-fetches VAE + Qwen3-4B
Image — Qwen-Image Qwen-Image / Qwen-Image-Edit auto-fetches VAE + Qwen2.5-VL + mmproj

All image families above were verified end-to-end on a Vulkan GPU (SDXL-Turbo ~2 s, Flux-schnell ~4 s, Bonsai ~2 s, Z-Image-Turbo ~11 s, Qwen-Image-Edit via /v1/images/edits).

Performance and tuning

Every knob below is set to a sensible default automatically and can be overridden per model from the dashboard's Configure screen (or globally via .env):

  • Speculative decoding, three lanes — attach a DFlash block-diffusion draft model to a supported target, use MTP / NextN prediction heads (auto-detected from the GGUF metadata), and stack the model-free n-gram lane on top for extra decode speed.
  • KV-cache quantization — q8_0 key/value cache compression by default, per-model override, with automatic fallback when a binary build does not support a cache type.
  • VRAM-aware quant picker — the add screen probes your GPU and marks the best-fitting quantization before you download anything.
  • Model pinning and hot-swap — pin a model to keep it resident in VRAM; unpinned models load on first request and unload after an idle TTL.
  • Mixture-of-Experts offload — push MoE expert layers to CPU (--n-cpu-moe) to fit large sparse models such as Qwen3.6-35B-A3B on a single consumer GPU.
  • Reasoning control — per-model thinking mode (on / off / auto) and reasoning budget for hybrid-reasoning models.
  • Per-model vision toggle — trade image input for the DFlash/MTP speculative lane on vision models, and switch back at any time.
  • Honest context windows — the requested context is clamped to the GGUF's real trained context, read from the file header.
  • CPU threads and memory locking — per-model --threads and --mlock for latency-sensitive deployments.

DFlash speculative decoding

DFlash speeds up a large target model by attaching a small z-lab block-diffusion draft model that proposes several of the target's next tokens per step, which the big model verifies in one pass — the target's quality at a fraction of the wall-clock time. It is a per-model attachment (like a vision projector), served by the same upstream llama-server (build b9831 or newer) — nothing extra to compile.

Press f on a highlighted chat model and, if it has a known pairing, the right community draft downloads and wires itself up. Or use Configure → Suggest / Browse / Clear for a progress bar and manual repo entry — pasting an official z-lab draft repo (raw safetensors, no GGUFs) into Browse auto-redirects to its paired GGUF conversion when one is known.

Target family Draft repo
Qwen3.6-27B / 35B-A3B (MoE) Alittlehammmer/*-DFlash-GGUF-llama.cpp
Gemma-4-31B / 26B-A4B (MoE) Alittlehammmer/*-DFlash-GGUF-llama.cpp
Gemma-4-12B williamliao/gemma-4-12B-it-DFlash-GGUF
Qwen3.5-27B / Qwen3-Coder-30B-A3B (MoE) AtomicChat/*-DFlash-GGUF
Qwen3.5-9B Anbeeld/Qwen3.5-9B-DFlash-GGUF
  • Thinking caveat: DFlash acceptance drops sharply (~5-14%) with reasoning on — run the target with reasoning off for the full speedup. inferhost warns when a draft is attached to a model whose reasoning resolves to on.
  • Vision caveat: draft-based speculation (DFlash and MTP) cannot run on a model with a vision projector (--mmproj) — an upstream llama-server limit. inferhost auto-disables the draft lane for vision models and serves them with the model-free n-gram lane only, so images always work. To get the draft speed instead of images, set Vision / image input to no in the model's Configure screen — the model serves text-only and the DFlash/MTP lane switches back on (flip it back any time; the projector stays downloaded).
  • VRAM: the draft is co-resident with the target (usually well under 2 GiB) and folded into the VRAM/pin-feasibility estimate automatically.
  • MoE targets (...-A3B / ...-A4B) are already cheap per step, so DFlash buys a smaller speedup than on a dense model of similar total size.
  • On a llama-server older than b9831, inferhost serves the model draftless with a notice rather than failing — see Usage.

Documentation

Full guides live in docs (and in the docs/ folder):

  • Installation — install, upgrade, uninstall, requirements
  • Usage — the dashboard, keyboard keys, and chat / TTS / image / Flux / Z-Image / Qwen-Image walkthroughs
  • Configuration — every .env variable, KV-cache quant, custom binaries
  • Troubleshooting — ports, tmux mouse, common errors
Architecture
Your app ──HTTP──▶  LiteLLM gateway        llama-swap (loopback)       llama-server  (chat/vision)
                    :9001 (public)   ──▶    127.0.0.1:9090      ──┬──▶  sd-server     (images)
                                                                  └──▶  inferhost-tts (speech)
  • llama.cpp (llama-server) runs chat/vision inference — official upstream binary, backend auto-detected.
  • llama-swap fronts the model backends and lazy-loads / hot-swaps them on demand (loopback only). Image models (sd-server) ride here too, so they swap VRAM with LLMs.
  • inferhost-tts wraps llama.cpp's llama-tts (OuteTTS) — or the separate qwen3-tts.cpp engine for Qwen3-TTS — behind /v1/audio/speech (started only when a TTS model is registered).
  • LiteLLM is the single always-on public gateway on :9001, routing each request to the right backend.

The extra engines (llama-tts, sd-server) are fetched automatically the first time you add a model that needs them. qwen3-tts.cpp has no prebuilt release, so it is compiled from source on demand instead (needs git/cmake/a C++ compiler on the host).

Development

The repo ships a run.sh wrapper for source-tree work (end users never need it — they only type inferhost):

git clone git@github.com:amirrouh/inferhost.git && cd inferhost
./run.sh install     # venv + editable install
./run.sh start       # launch the TUI (downloads binaries on first run)
./run.sh status      # headless status
./run.sh stop        # stop daemons
./run.sh test        # pytest

Run ./run.sh help for the full list.

License

Apache 2.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

inferhost-0.8.5.tar.gz (830.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

inferhost-0.8.5-py3-none-any.whl (124.3 kB view details)

Uploaded Python 3

File details

Details for the file inferhost-0.8.5.tar.gz.

File metadata

  • Download URL: inferhost-0.8.5.tar.gz
  • Upload date:
  • Size: 830.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for inferhost-0.8.5.tar.gz
Algorithm Hash digest
SHA256 c3820ced1a9b60eab7daae0f4441fbaa278b7e22c53971ed22258062b183d20e
MD5 15b2cff608e53944ff150cb61b21591d
BLAKE2b-256 b8a929311d52a78a955e13ddf37090f895d94edc2c51a3d53cefcf8e523e4538

See more details on using hashes here.

Provenance

The following attestation bundles were made for inferhost-0.8.5.tar.gz:

Publisher: publish.yml on amirrouh/inferhost

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file inferhost-0.8.5-py3-none-any.whl.

File metadata

  • Download URL: inferhost-0.8.5-py3-none-any.whl
  • Upload date:
  • Size: 124.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for inferhost-0.8.5-py3-none-any.whl
Algorithm Hash digest
SHA256 2f2c63504ae49706056fa680e707ea034c5174297557097423285230763eaaf8
MD5 016e5fbc446b651d928ec884b1c95fe6
BLAKE2b-256 be7f9befedb4cc0e8a6ae8284729c994ee1d2c3eacef84302b48eb47ffcb3548

See more details on using hashes here.

Provenance

The following attestation bundles were made for inferhost-0.8.5-py3-none-any.whl:

Publisher: publish.yml on amirrouh/inferhost

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page