Skip to main content

Yunshu

A fast local LLM / VLM inference engine for Apple Silicon.

One process, OpenAI- and Anthropic-compatible, running on-device via MLX. Built for low latency on a single machine: fast first token, fast lossless decode, and prefix reuse that skips work you already paid for. The first fully tuned model is Qwen3.8-27B.

Python 3.13+ License: Apache 2.0 CI

English · 简体中文 · 繁體中文


What it does differently

  • Lossless speculative decode. A DFlash2 drafter when one is installed for the model (picked automatically), else MTP (the checkpoint's own head). Every decode and verify matmul of a drafting request goes through one batch-invariant kernel, so greedy output with speculation on is token-identical to speculation off — the guarantee Splash calls lossless.
  • Per-row KV for concurrent requests. Rows of the shared batch keep their own KV length, so a short request never reads a long one's padding (Qwen3.5 family; MMLU-Pro at 8 in flight 88 → 139 tok/s, 131K-context decode 24 → 47 tok/s).
  • Lossy only when you ask. Every default is lossless. Memory savers that change outputs (int8 KV, KV quantization, 4-bit cached prefixes, int8 SSD cache) are settings you turn on.
  • Prefix cache for hybrid models. Qwen3.5-family models mix attention with recurrent GatedDeltaNet layers, which ordinary KV caches cannot slice. Yunshu keeps exact checkpoints, keyed by image pixels as well as text, 8 GiB in RAM by default plus an optional SSD tier. Repeated or edited long prompts skip prefill.
  • Verify kernels checked against output. GatedDeltaNet, attention and 5-bit matmul verify kernels, partly vendored from oMLX, each adopted only after a same-checkpoint A/B.
  • The full API surface on the fast path. Tool calls, JSON-schema constraints, stop sequences, logprobs, a streaming reasoning/content split, reasoning_effort passed to chat templates that support it (Qwen3.8), and cancellation on client disconnect.

Quickstart

Needs a Mac with Apple Silicon (macOS 14+) and uv.

# Install. The vision extra covers the Qwen3.5 / 3.6 / 3.8 family and every VLM.
uv tool install "yunshu[vision]"

yunshu doctor                                   # checks this Mac and prints fixes
yunshu pull mlx-community/Qwen3.5-9B-MLX-4bit   # downloads to ~/.yunshu/models/
yunshu serve -m mlx-community/Qwen3.5-9B-MLX-4bit

Other ways to install: pipx install "yunshu[vision]", Homebrew (brew install yuhuanstudio/tap/yunshu), or the latest main (uv tool install "yunshu[vision] @ git+https://github.com/YuhuanStudio/Yunshu").

yunshu serve -m org/name uses a model already in the models directory or the Hugging Face cache and downloads only when neither has it. Models live in ~/.yunshu/models; to keep them elsewhere, run yunshu config set models_dir /path/to/models (saved in ~/.yunshu/config.toml).

Qwen3.8-27B (the tuned model)

Measured memory footprint is about 21 GiB at a 1K prompt and 29 GiB at 32K (weights, drafter and KV), so use a Mac with 32 GB or more; 131K context needs more.

yunshu pull Jundot/Qwen3.8-27B-oQ4e-mtp            # the model: 4-bit, ships an MTP head
yunshu pull incoai/Qwen3.8-27B-DFlash2             # the drafter: faster than MTP
yunshu doctor -m Jundot/Qwen3.8-27B-oQ4e-mtp       # "speculative" row shows the path in use
yunshu serve -m Jundot/Qwen3.8-27B-oQ4e-mtp

With the drafter in the models directory or the Hugging Face cache, yunshu serve finds and uses it; no flag is needed. The startup log says Speculative decoding: dflash. To choose yourself: YUNSHU_VLM_DRAFT=mtp forces the checkpoint's MTP head, YUNSHU_VLM_DRAFT=off disables drafting, YUNSHU_VLM_DRAFT=/path/to/drafter picks a specific drafter. Every path is lossless: with speculation on, greedy output equals speculation off.

The server listens on http://127.0.0.1:8000. Any OpenAI client works unchanged:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")  # any key works

# Single-model mode: the model name is a placeholder, the server serves what you loaded.
r = client.chat.completions.create(
    model="local",
    messages=[{"role": "user", "content": "Explain MLX in one sentence."}],
    extra_body={"reasoning_effort": "medium"},
)
print(r.choices[0].message.content)

To run it in the background at login: yunshu service install -m <model> (service guide). Every command has --help. yunshu model list shows local models, including the Hugging Face cache.

From source (development): clone the repository, run uv sync --extra vision (or --all-extras), then uv run yunshu serve -m <model>. uv.lock pins the exact versions (MLX 0.32, mlx-vlm 0.7.3+).

No telemetry. Yunshu sends nothing anywhere. The only outgoing connections are the model downloads you ask for and MCP servers you configure.

Docs:

Performance

Measured on an M5 Max (128 GB), Qwen3.8-27B, 2026-09-28/29. Same Jundot oQ4e-mtp checkpoint unless noted. How each table was measured (scripts and methods) is in docs/BENCHMARKS.md; the raw run data stays with the maintainers.

Engine Capability checks Chat TTFT (warm) 8K prompt: cold / repeat / edited tail Decode tok/s
Yunshu 0.1.1 (default: MTP block 6, batch-invariant, ragged KV) 34/34 0.192 s 8.42 / 0.115 / 0.259 s 73
mlx-vlm 0.7.3 server (APC) 27/28 0.212 s 8.60 / 0.108 / 0.265 s 32
oMLX.app 0.7 (MTP + cache) 31/31 0.312 s 8.60 / 0.361 / 0.376 s 85
Splash 1.1 (own quantized model + DFlash2) 31/31 0.206 s 7.88 / 0.131 / 7.88 s 119
TensorFold 0.3.6.1 (MTP, parallel 8) 23/34 — — 28

Yunshu's matrix has more checks than the older runs (logprobs, the streaming reasoning split); TensorFold fails the image, tool, JSON-schema and logprobs checks.

Lossless single-request decode by output type (same checkpoint, in-process, greedy, 384 tokens; tok/s):

Context Code Prose JSON-like Spec on == off
1K 82.1 57.8 69.0 yes (tested)¹
32K 75.4 51.2 64.2 yes (tested)¹
131K 59.7 43.8 46.0 yes (tested)¹

¹ Speculative and plain greedy output matched token for token on every task at every context above. Matmuls are row-invariant, and decode and verify attention run one per-row kernel whose result for a token does not depend on how many tokens are verified with it.

MMLU-Pro, 300 questions, 8 in flight, max 16384 tokens, reasoning_effort=medium (accuracy and a long-run soak; same settings for every engine):

Engine Correct Wall time Aggregate tok/s Peak footprint
Yunshu (ragged KV) 249 / 300 28.3 min 139 33.7 GiB
Yunshu 0.1.0-era shared batch (padded KV) 250 / 300 46.5 min 88 45 GiB
TensorFold 0.3.6.1 (MTP, parallel 8) 250 / 300 24.9 min 159 35.1 GiB
Splash 1.1 252 / 300 17.2 min 223 67 GiB
oMLX.app 229 / 300 (27 rejected by its prefill memory guard) 29.2 min 120 75 GiB

With YUNSHU_KV_PRECISION=int8 (lossy, opt-in) Yunshu scored 251 / 300 at a 28.1 GiB peak (measured on an earlier build of the ragged cache, 34.6 min).

Default speculative path on Qwen3.8-27B: DFlash2 drafter (cost-aware chain depth, 8-bit drafter weights), server, greedy, 128 generated tokens, one request, unique prompts, tok/s. Two corpora: novel prose (novel_en, hard to draft) and Python code (code_python, easy):

Context novel_en code_python
1K 57.1 82.0
8K 48.2 89.1
32K 46.1 70.5
131K 32.9 79.9

TTFT at 131K is about 207 s (cold prefill, ~640 tok/s, at the hardware ceiling); the capability matrix is 34/34. Against the checkpoint's MTP head on the same server, novel_en at 1K is 57.1 versus 47.5 tok/s. The numbers move a few tok/s from run to run because acceptance depends on the text.

Same-corpus comparison (novel_en, single request, tok/s; the quantization may differ between engines; both measured 2026-09-29):

Context Yunshu (DFlash2) TensorFold 0.3.6.1 (DFlash2)
1K 57.1 76.4
8K 48.2 67.2
32K 46.1 63.4
131K 32.9 38.9
TTFT at 131K 207 s 280 s

TensorFold decodes novel prose faster than Yunshu at every context here (about 1.2x at 131K, 1.3-1.4x at 1K-32K); Yunshu's prefill is faster (207 vs 280 s at 131K). Splash and oMLX were not run on this corpus; their figures in the next table come from an earlier, different prompt set and should not be read against the columns above.

Earlier speed sweep with the MTP draft (unique prompts, no cache hits; 128 generated tokens; tok/s unless noted):

Yunshu oMLX Splash TensorFold (MTP)
TTFT at 8K / 131K tokens 8.6 / 207 s 8.5 / 214 s 7.9 / 207 s 9.8 / 293 s
Decode after 1K / 32K / 131K 59² / 59 / 47 71 / 60 / 38 101 / 48 / 68 26 / 57 / 19
8 concurrent 1K prompts, aggregate 64 53 70 65

² Single-request decode depends on how many drafted tokens are accepted, which varies with the prompt; Yunshu's 1K figure is the mean of 8 runs (single runs ranged 40–70). The other cells are single runs.

Where Yunshu stands:

  • Prefix reuse and warm TTFT are the best measured; cold prefill is at the hardware ceiling (every engine within ~10%).
  • Accuracy matches the others; 0 errors in every long run.
  • Behind Splash on long-context decode (131K: 47 vs 68 tok/s) and on concurrent long outputs (MMLU-Pro: 139 vs 223 tok/s). TensorFold is also ahead there (159) because it drafts for every row; Yunshu drafts only for a request that is alone. Multi-row speculative decoding is in progress (YUNSHU_ROUND_DRIVER, experimental).
  • A long prompt that arrives while others decode stalls their decode during its prefill; every engine measured does this.
  • A 60-minute mixed soak on the 2026-09-28 build (chat, long documents, images, tools, JSON schema, thinking, disconnects) finished 699 requests with 0 server errors and no memory growth (footprint 17–26 GiB).

Supported models

Tier Models Path What you get
1 — tuned and measured Qwen3.5 / 3.6 / 3.8 family (text + images) VLM batch runner prefix cache (RAM + SSD), MTP / DFlash lossless spec decode, all API features above
2 — supported any mlx-lm text model single-request fast path (generate_step) KV prefix cache, tools, JSON schema, logprobs; opt-in n-gram spec, KV quant
2 — supported any other mlx-vlm model (GLM, Qwen-VL, Gemma-4, Qwen3-Omni, Nemotron-Omni, …) the same VLM batch runner continuous batching, prefix cache (unless the model uses a sliding window), images / audio / video, all API features above; no speculative decode

Only tier 1 was re-measured in the 2026-09-28 round; tier-2 VLMs moved onto the runner afterwards and still need their real-model smoke run.

Other modalities

These ship in the same server behind optional extras. None were re-verified in the 2026-09-28 round, which covered LLM/VLM only.

Modality Endpoint Backend Extra
Native speech-to-speech (Qwen3-Omni Thinker→Talker, streaming) POST /v1/omni/speech/stream mlx-vlm omni
Realtime voice WS /v1/realtime omni, or ASR → LLM → TTS audio
Text over WebSocket (multiplexed, cancel by id) WS /v1/stream, WS /v1/responses any served model --
ASR /v1/audio/transcriptions mlx-audio / Whisper audio
TTS /v1/audio/speech mlx-audio audio
Image generation /v1/images/generations diffusion generation
Embeddings / rerank (text + multimodal) /v1/embeddings, /v1/rerank mlx-lm / mlx-embeddings embeddings

For speech-to-speech, serve a Qwen3-Omni model (uv sync --extra omni) and try examples/talk.py (microphone) or examples/quickstart.py (writes a WAV, no audio hardware). Upstream mlx-vlm 0.7.3 was checked to keep multi-turn omni output correct (maintainer check, 2026-09-28); the server's Realtime path was not.

Text-to-video generation is not offered over HTTP (the route was removed); video input to VLMs works.

Transports beyond SSE (docs/guides/TRANSPORTS.md): a multiplexed text WebSocket (WS /v1/stream: many requests on one socket, cancel / stop / max_tokens update by id, heartbeats, backpressure; client in python/yunshu_client, demo in examples/ws_stream.py), OpenAI's Responses WebSocket mode (client.responses.connect()), Realtime in the GA and beta schemas, and yunshu serve --uds PATH for a Unix socket (curl --unix-socket, OpenAI SDK over an httpx uds transport). WebRTC and HTTP/2 are not offered yet.

Also: MCP server/client, an Anthropic-compatible /v1/messages surface and an Ollama-compatible /api layer (route status: docs/guides/API_SURFACE.md).

Architecture

  Client (any OpenAI / Anthropic SDK)
        │   OpenAI / Anthropic / MCP / Realtime-WS / SSE
  ┌─────┴───────────────────────────────────────────────┐
  │  Gateway (FastAPI)     routers + middleware           │
  ├─────────────────────────────────────────────────────┤
  │  Engine                                               │
  │   · VLM batch runner (every mlx-vlm model)            │
  │       continuous batching · prefix cache (RAM + SSD)  │
  │       Qwen3.5 family: MTP / DFlash + batch-invariant  │
  │   · LLM fast path (mlx-lm generate_step)              │
  │       KV prefix cache · constrained decoding          │
  │   · other modalities: omni, ASR/TTS, image,           │
  │     embeddings                                        │
  └─────────────────────────────────────────────────────┘
        one MLX thread · runs on-device via Apple MLX

Serving model

All GPU work runs on one MLX thread. VLM (mlx-vlm) models serve concurrent requests in one continuous batch, each row with its own sampling settings; a request that is alone uses speculative decoding (Qwen3.5 family), and requests that arrive meanwhile join the shared batch without it. Text-only mlx-lm models use the single-request fast path, so their concurrent requests run one at a time. Every response returns as soon as its own generation finishes.

Configuration

Every setting is a YUNSHU_* name listed in the configuration reference (generated from one registry). Set it as an environment variable, in a TOML file (yunshu serve --config yunshu.toml), or with yunshu serve --set KEY=VALUE; yunshu config shows each effective value and where it came from. A bad value stops startup; a misspelled name gets a warning.

Built on

MLX · mlx-lm · mlx-vlm · mlx-audio. Some kernels are vendored from oMLX (Apache-2.0) and TensorFold (MIT); see THIRD_PARTY_NOTICES.md.

License

Apache 2.0 — see LICENSE.

Metadata

Release files for yunshu 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for yunshu 0.1.1
File Size Uploaded
yunshu-0.1.1.tar.gz 2.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for yunshu 0.1.1
File Interpreter ABI Platform
yunshu-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 3.9 MB

Release files / yunshu-0.1.1.tar.gz

Download URL yunshu-0.1.1.tar.gz
Size 2.3 MB
Tags Source
SHA-256 checksum
How to use checksums
f82f82b0608757e8ba60ffaa1ba92e418562249fa8ff8518b316c6491e05e054
BLAKE2b-256 checksum
How to use checksums
f6f4c2bbf374ef0191d260e50ead818e7f1b3dbaec9aa717ab1f25195fa0fdfa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / yunshu-0.1.1-py3-none-any.whl

Download URL yunshu-0.1.1-py3-none-any.whl
Size 1.5 MB
Tags Python 3
SHA-256 checksum
How to use checksums
ab452f86eab63a5f45b677d5ceae5990d061202f207f89b67fde975fcdb848bd
BLAKE2b-256 checksum
How to use checksums
21011dd40618133fd297d155d7c274a5e08382b919966fb64173644630faeaed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page