Skip to main content
logo

PyPI GitHub stars license closed issues open issues Ask DeepWiki


Blog | Documentation | Quick Start | Cookbook | SGLang | Join Slack

Star SGLang-Omni to help more builders discover open infrastructure for multimodal and speech serving!

News

  • [2026/08] 🎵 Day-0 support for MiniMax Music 3: lyrics + caption → 32 kHz stereo song on /v1/audio/speech. [Cookbook]
  • [2026/08] 🚀 SGLang-Omni v0.1.3 is on PyPI. Install with uv pip install --prerelease=allow "sglang-omni==0.1.3". [Installation]
  • [2026/08] 🚀 TTS architecture refactor: shared pipeline state, engine construction, reference encoding, capability metadata, and vocoder scheduling. [Roadmap] [Blog]
  • [2026/06] 🔥 MOSS-TTS Local Transformer v1.5 on SGLang-Omni with native-streaming 48 kHz speech. [Blog] [Cookbook]
  • [2026/06] 🔥 Higgs Audio v3 TTS for real-time, controllable speech. [Blog] [Cookbook]

About

SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with SGLang for high-performance autoregressive scheduling and model execution where applicable.

  • Multi-stage runtime: SGLang-Omni models generation as coordinated stages: preprocessing, encoders, autoregressive engines, talkers, decoders, vocoders, and aggregators.
  • Stage-specialized scheduling: Each stage runs behind a scheduler matched to its workload, from SGLang-backed autoregressive scheduling to lightweight preprocessing and streaming vocoder loops.
  • Transport-aware execution: A control plane coordinates requests while the relay data plane moves tensor payloads across shared-memory, NCCL, NIXL, and Mooncake backends.
  • API surface: OpenAI-compatible endpoints expose multimodal chat, speech generation, batch speech, streaming speech, uploaded voices, and transcription.

What SGLang-Omni Serves

Hardware Support

Backend Status Notes
NVIDIA CUDA Supported Default backend with full model coverage.
Intel GPU (XPU) Experimental Intel Arc GPUs via PyTorch XPU. Qwen3-ASR, Qwen3-TTS, and Qwen3-Omni serve end-to-end (Omni thinker via multi-XPU tensor parallelism). Install per Intel XPU guide; the backend is auto-detected.

Additional model guides, including experimental and research-oriented paths, are available in the Cookbook.

Quick Start

Community & Support

SGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the SGLang Slack or read the developer reference.

Organizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at zhaochenyang@lmsys.org.

Acknowledgments

SGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.

Release files for sglang-omni 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sglang-omni 0.1.3
File Size Uploaded
sglang_omni-0.1.3.tar.gz 1.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for sglang-omni 0.1.3
File Interpreter ABI Platform
sglang_omni-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 2.6 MB

Release files / sglang_omni-0.1.3.tar.gz

Download URL sglang_omni-0.1.3.tar.gz
Size 1.2 MB
Tags Source
SHA-256 checksum
How to use checksums
726ca4777c9e2f18f95fbd396ef270f2a378732aed53967badbb3bdd55947aeb
BLAKE2b-256 checksum
How to use checksums
7b7077a1211790b0a8bcc70474408375225c2b46c72d2c3948a08d6a18b599ff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.11

Release files / sglang_omni-0.1.3-py3-none-any.whl

Download URL sglang_omni-0.1.3-py3-none-any.whl
Size 1.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
7c448e3ff2e634f3338633eebb032a777c528569cceeaa81621a411c2c7ceb09
BLAKE2b-256 checksum
How to use checksums
785e3fc83bee568112a4e4927b54e6280e92fd8982faa8b345de4c0b5628946d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.11

Release history Release notifications | RSS feed

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page