Skip to main content
RapidSpeech Logo

English | 简体中文

Open in Colab

RapidSpeech.cpp

On-device speech AI runtime for ASR, TTS, VAD, and voice cloning. Python-simple, C++-native, GGUF-powered.

RapidSpeech.cpp runs speech recognition, text-to-speech, VAD, speaker embedding, and voice cloning on-device. It gives Python developers a simple API while keeping the runtime pure C/C++, backed by ggml and a unified GGUF model format. No cloud API, no speech server, no heavyweight Python model stack.


Python In 60 Seconds

Install

pip install rapidspeech

GPU wheels:

pip install rapidspeech-metal   # macOS / Apple Silicon
pip install rapidspeech-cuda    # Linux / NVIDIA

Text to speech

python python-api-examples/tts/tts-offline.py \
  --model /path/to/omnivoice-f16.gguf \
  --text "Hello, welcome to RapidSpeech." \
  --output output.wav

Speech to text

python python-api-examples/asr/asr-offline.py \
  --model /path/to/funasr-nano-fp16.gguf \
  --audio /path/to/audio.wav

Python API

import rapidspeech

tts = rapidspeech.tts_synthesizer("/path/to/omnivoice-f16.gguf")
tts.set_params(instruct="male, young adult", language="English", seed=42)
pcm = tts.synthesize("Hello from a native speech engine.")
sample_rate = tts.get_sample_rate()
import rapidspeech

asr = rapidspeech.asr_offline("/path/to/funasr-nano-fp16.gguf")
sample_rate = asr.get_model_meta()["audio_sample_rate"]
pcm = ...  # 1-D float32 mono PCM at sample_rate
asr.push_audio(pcm)
asr.process()
print(asr.get_text())

Why RapidSpeech.cpp

  • Built for the edge: run speech models locally on laptops, servers, browsers, and device-class hardware.
  • Python-simple, C++-native: write Python, run a C++/ggml engine underneath.
  • One model format: ASR, TTS, VAD, and speaker models use GGUF.
  • NumPy in, NumPy out: ASR takes float32 PCM; TTS returns float32 PCM.
  • Edge-first backends: CPU, Metal, CUDA, Vulkan, CANN, OpenCL, and WebGPU.

Performance Snapshot

Test environment: Apple M1 Pro, funasr-nano-fp16.gguf, 15s audio.

Configuration RTF Wall Time Notes
CPU -t 4 0.465 12.4s CPU-only inference
GPU -t 4 0.170 5.2s Metal acceleration
GPU -t 4 Q4_K 0.756 - Quantized model: GPU dequant overhead
CPU -t 4 Q4_K 0.530 - Quantized model CPU inference, 596 MB (3.3x compression)

RTF is processing time divided by audio duration. Lower is faster; RTF < 1 is faster than real time.


Supported Today

Task Models Status
ASR SenseVoice-small, FunASR-nano, X-ASR (Zipformer2, streaming) Stable
VAD Silero VAD, FireRedVAD Stable
TTS OmniVoice, OpenVoice2, Kokoro, IndexTTS-2 Active
Speaker CAMPPlus Stable

X-ASR — Chinese/English Zipformer2 transducer (icefall/k2). One GGUF serves both offline full-context decoding and true chunked streaming (per-layer left-context caches, sub-second partials, --chunk-len 16/32/48/96/192 fbank frames). Punctuation and casing, greedy transducer decode, runs on CPU / Metal / CUDA / Vulkan and quantizes to q4_k_m (99.5 MB).

IndexTTS-2 — expressive zero-shot voice-cloning TTS (GPT + S2Mel CFM + BigVGAN-v2 vocoder) with 4-mode emotion control (reference audio / vector / text / Qwen). See docs/index2tts.md.

In Progress

CosyVoice3, Qwen3-ASR, Qwen3-TTS.


Documentation


Native C++ CLI

Download Models

Models are available on:

Build from Source

git clone https://github.com/RapidAI/RapidSpeech.cpp
cd RapidSpeech.cpp
git submodule sync && git submodule update --init --recursive
cmake -B build
cmake --build build --config Release

Self-contained executables (no runtime DLL/.so dependencies) — build the core and ggml statically into each CLI with -DRS_STATIC_EXE=ON:

# Windows / MSVC
cmake -B build -G "Visual Studio 17 2022" -A x64 -DRS_STATIC_EXE=ON
cmake --build build --config Release --parallel
# Linux / macOS
cmake -B build -DRS_STATIC_EXE=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Build artifacts are located in the build/ directory:

  • rs-asr-offline — Offline ASR command-line tool
  • rs-asr-vad-online — VAD-segmented quasi-streaming ASR command-line tool
  • rs-asr-online — True chunked streaming ASR (X-ASR; mic or WAV, live partials)
  • rs-tts-offline — Offline TTS command-line tool
  • rs-server — OpenAI-compatible HTTP API + MCP server (ASR + TTS)
  • rs-quantize — Model quantization tool

Core Commands

Offline ASR

./build/rs-asr-offline \
  -m /path/to/funasr-nano-fp16.gguf \
  -w /path/to/audio.wav \
  -t 4 \
  --gpu true

VAD-segmented ASR

./build/rs-asr-offline \
  -m /path/to/funasr-nano-fp16.gguf \
  -v /path/to/silero_vad_v6.gguf \
  -w /path/to/audio.wav \
  -t 4 \
  --vad-threshold 0.5 \
  --silence-ms 600

Hotword biasing (FunASR-Nano)

# Bias proper nouns / fix homophones (e.g. 郭总 → 虢总); use --hotword-file for a large list
./build/rs-asr-offline \
  -m /path/to/funasr-nano-fp16.gguf \
  -w /path/to/audio.wav \
  --hotwords "虢总,阿里巴巴"

See docs/funasr-nano.md for details.

Streaming ASR (X-ASR)

# WAV, real-time paced with live partials (or --fast to run as fast as possible)
./build/rs-asr-online -m /path/to/xasr-q4_k_m.gguf -w /path/to/audio.wav --chunk-len 32
# Microphone
./build/rs-asr-online -m /path/to/xasr-q4_k_m.gguf --mic --chunk-len 16

See docs/x-asr.md for the model, chunk-size / latency tradeoffs, and GGUF conversion.

Text to speech

./build/rs-tts-offline \
  -m /path/to/omnivoice-f16.gguf \
  -t "Hello, welcome to RapidSpeech!" \
  --instruct "male, young adult, moderate pitch" \
  --lang English \
  --n-steps 32 \
  -o output.wav

Quantization

./build/rs-quantize /path/to/input-f16.gguf /path/to/output-q4_k.gguf q4_k

Server (OpenAI API + MCP)

# Serve ASR + TTS over an OpenAI-compatible HTTP API and MCP
./build/rs-server --asr-model xasr.gguf --tts-model omnivoice.gguf --port 8080

curl http://127.0.0.1:8080/v1/audio/transcriptions -F file=@audio.wav -F model=rapidspeech-asr
curl http://127.0.0.1:8080/v1/audio/speech -H 'content-type: application/json' \
     -d '{"input":"hello","voice":"female"}' --output out.wav

Also runs as an MCP server (stdio for Claude Desktop, or POST /mcp) and exposes WebSocket streaming endpoints (streaming ASR with partial/final, VAD-segmented ASR, segmented + pure streaming TTS), plus a browser test console (--web-dir examples/server → http://host:port/webui.html). See examples/server/README.md.

Python

See Python examples for offline ASR, streaming ASR, offline TTS, streaming TTS, VAD, and voice cloning.


🤝 Contributing

If you are interested in the following areas, we welcome your PRs or participation in discussions:

  • Adapting more models to the framework.
  • Refining and optimizing the project architecture.
  • Improving inference performance.

Acknowledgements

  1. Fun-ASR
  2. llama.cpp
  3. ggml
  4. cppjieba — Chinese word segmentation
  5. WeText — text normalization (ITN/TN)
  6. miniaudio — single-file audio I/O
  7. X-ASR Streaming-focused automatic speech recognition models

Release files for rapidspeech-metal 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for rapidspeech-metal 1.3.0
File
rapidspeech_metal-1.3.0-cp313-cp313-macosx_11_0_arm64.whl CPython 3.13 CPython 3.13 macOS 11.0+ ARM64 Details
rapidspeech_metal-1.3.0-cp312-cp312-macosx_11_0_arm64.whl CPython 3.12 CPython 3.12 macOS 11.0+ ARM64 Details
rapidspeech_metal-1.3.0-cp311-cp311-macosx_11_0_arm64.whl CPython 3.11 CPython 3.11 macOS 11.0+ ARM64 Details
rapidspeech_metal-1.3.0-cp310-cp310-macosx_11_0_arm64.whl CPython 3.10 CPython 3.10 macOS 11.0+ ARM64 Details
rapidspeech_metal-1.3.0-cp39-cp39-macosx_11_0_arm64.whl CPython 3.9 CPython 3.9 macOS 11.0+ ARM64 Details

Total release size: 47.5 MB

Release files / rapidspeech_metal-1.3.0-cp313-cp313-macosx_11_0_arm64.whl

Download URL rapidspeech_metal-1.3.0-cp313-cp313-macosx_11_0_arm64.whl
Size 9.5 MB
Tags CPython 3.13 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
c1ccd9544731157309bb7ed7c0328c6d595968d1b3bfa4ce71fdb021c75e2cbe
BLAKE2b-256 checksum
How to use checksums
4b504a767e9e622955b30a15d03ed54f190aa9b3856d9f47783593157f488274
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release files / rapidspeech_metal-1.3.0-cp312-cp312-macosx_11_0_arm64.whl

Download URL rapidspeech_metal-1.3.0-cp312-cp312-macosx_11_0_arm64.whl
Size 9.5 MB
Tags CPython 3.12 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
37c0c6f033c97676510ab116e819e889f62bfecb0724699ddd131e27e873701d
BLAKE2b-256 checksum
How to use checksums
5a294f61d2278a3632bfdb7e8449e76b521fa5d6878e019ba5c42ed9d1a62364
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release files / rapidspeech_metal-1.3.0-cp311-cp311-macosx_11_0_arm64.whl

Download URL rapidspeech_metal-1.3.0-cp311-cp311-macosx_11_0_arm64.whl
Size 9.5 MB
Tags CPython 3.11 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
226958e83e47b9f56db5f7b876149a68a381e9b2331b2859f792389572034e0c
BLAKE2b-256 checksum
How to use checksums
2c27141798c9d2bdb7eb6951d4d81d5f12067bd16ef1ce1f46e5ba4db17287c4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release files / rapidspeech_metal-1.3.0-cp310-cp310-macosx_11_0_arm64.whl

Download URL rapidspeech_metal-1.3.0-cp310-cp310-macosx_11_0_arm64.whl
Size 9.5 MB
Tags CPython 3.10 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
d4964e489f4b69b92e52010383298e186583e4d63acdf82de430f644d95b4629
BLAKE2b-256 checksum
How to use checksums
1c270777ed8faaea463d5038c180236b2283448c628faf285fc2a22332710f3b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release files / rapidspeech_metal-1.3.0-cp39-cp39-macosx_11_0_arm64.whl

Download URL rapidspeech_metal-1.3.0-cp39-cp39-macosx_11_0_arm64.whl
Size 9.5 MB
Tags CPython 3.9 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
82db67644fccc435f10ae11dd1a0dde57d94bfdbd4b2dca11d9a5d6911093b37
BLAKE2b-256 checksum
How to use checksums
2149bc11fa9518e5886b3d5f069ce66cebf4180820fa5fce2f522e4e0d6c397c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.3.0 This release

5 release files

1.2.0

5 release files

1.1.1

5 release files

1.1.0

5 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page