Blog | Documentation | Quick Start | Cookbook | SGLang | Join Slack
⭐ Star SGLang-Omni to help more builders discover open infrastructure for multimodal and speech serving!
News
- [2026/08] 🎵 Day-0 support for MiniMax Music 3: lyrics + caption → 32 kHz stereo song on
/v1/audio/speech. [Cookbook] - [2026/08] 🚀 SGLang-Omni v0.1.2 is on PyPI. Install with
uv pip install "sglang-omni==0.1.2". [Installation] [Release notes] - [2026/08] 🚀 TTS architecture refactor: shared pipeline state, engine construction, reference encoding, capability metadata, and vocoder scheduling. [Roadmap] [Blog]
- [2026/06] 🔥 MOSS-TTS Local Transformer v1.5 on SGLang-Omni with native-streaming 48 kHz speech. [Blog] [Cookbook]
- [2026/06] 🔥 Higgs Audio v3 TTS for real-time, controllable speech. [Blog] [Cookbook]
About
SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with SGLang for high-performance autoregressive scheduling and model execution where applicable.
- Multi-stage runtime: SGLang-Omni models generation as coordinated stages: preprocessing, encoders, autoregressive engines, talkers, decoders, vocoders, and aggregators.
- Stage-specialized scheduling: Each stage runs behind a scheduler matched to its workload, from SGLang-backed autoregressive scheduling to lightweight preprocessing and streaming vocoder loops.
- Transport-aware execution: A control plane coordinates requests while the relay data plane moves tensor payloads across shared-memory, NCCL, NIXL, and Mooncake backends.
- API surface: OpenAI-compatible endpoints expose multimodal chat, speech generation, batch speech, streaming speech, uploaded voices, and transcription.
What SGLang-Omni Serves
- Omni chat and speech: Qwen3-Omni, Ming-Omni — multimodal in, text/audio out.
- Music generation: MiniMax Music 3 — lyrics + caption → 32 kHz stereo song.
- Speech generation: Higgs Audio v3, MOSS-TTS, MOSS-TTS Local, Fish Speech S2-Pro, Qwen3-TTS, Voxtral TTS, Ming-Omni-TTS, dots.tts, ZONOS2 —
/v1/audio/speech, batch, streaming, uploaded voices. - Audio transcription and diarization: Qwen3-ASR, Fun-ASR, ARK-ASR, MOSS-Transcribe-Diarize via
/v1/audio/transcriptions. MOSS-TD supports speaker labels and timestamps (response_format=verbose_json). - SGLang-Omni Router: Multi-worker OpenAI-compatible front door — health, readiness, lifecycle, capability discovery. Router guide.
Hardware Support
| Backend | Status | Notes |
|---|---|---|
| NVIDIA CUDA | ✅ Supported | Default target; full model coverage. |
| Intel GPU (XPU) | 🧪 Experimental | Intel Arc GPUs via PyTorch XPU. Qwen3-ASR, Qwen3-TTS, and Qwen3-Omni serve end-to-end (Omni thinker via multi-XPU tensor parallelism). Install per Intel XPU guide; the backend is auto-detected. |
Additional model guides, including experimental and research-oriented paths, are available in the Cookbook.
Quick Start
- Installation
- TTS usage
- Qwen3-Omni usage
- Qwen3-ASR cookbook
- MOSS-Transcribe-Diarize cookbook
- Omni router
- Developer reference
Community & Support
SGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the SGLang Slack or read the developer reference.
Organizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at zhaochenyang@lmsys.org.
Acknowledgments
SGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.
Release files for sglang-omni 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sglang_omni-0.1.2.tar.gz | 1.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sglang_omni-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.5 MB
Release files / sglang_omni-0.1.2.tar.gz
| Download URL | sglang_omni-0.1.2.tar.gz |
|---|---|
| Size | 1.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
73a33de7a7529f5e8fd487163fd622de6c1ff4e9111773ab3159a25279dfbee3
|
|
BLAKE2b-256 checksum How to use checksums |
bbc64246120de86f7b28179a6c9140d3a0b299c174a9fbc0ffb4dc898153dc12
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.11
|
Release files / sglang_omni-0.1.2-py3-none-any.whl
| Download URL | sglang_omni-0.1.2-py3-none-any.whl |
|---|---|
| Size | 1.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2353808d3dd676c0dacc3cb306a6447e3b0a39dd70944b00feeb9a6757c8cf29
|
|
BLAKE2b-256 checksum How to use checksums |
d254e16b78bbe6c7957d5a8f040f4a0d5b92c265b772a641f843d19f23a7c7e8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.11
|