mlx-speech
Text-to-speech and speech recognition on Apple Silicon
Voice cloning, audio editing, sound effects, and transcription. All running locally in MLX.
Quick start · Models · Model guides · Hugging Face weights · Project website
Website contributions belong in the frontend repository. See website ownership.
mlx-speech is an open-source Python library for text-to-speech (TTS) and automatic speech recognition (ASR) on Apple Silicon Macs. Models share a Python API and command-line interface, with inference implemented in Apple's MLX framework.
Audio processing stays on your Mac. Inference needs neither PyTorch nor a cloud service. Published weights download on first use, and local checkpoint paths support offline loading.
Installation
Requires an Apple Silicon Mac (M1 or later) and Python 3.13+.
pip install mlx-speech
Quick Start
Python:
import mlx_speech
from mlx_speech.audio import write_wav
# Text-to-speech
model = mlx_speech.tts.load("fish-s2-pro")
result = model.generate("Hello from mlx-speech!")
write_wav("output.wav", result.waveform, sample_rate=result.sample_rate)
# Speech-to-text
asr = mlx_speech.asr.load("qwen3-asr-1.7b")
print(asr.generate("audio.wav").text)
CLI:
mlx-speech tts --model fish-s2-pro --text "Hello!" -o output.wav
mlx-speech asr --model qwen3-asr-1.7b --audio speech.wav
Models
Pass the selector to tts.load(), asr.load(), or --model. Names link to
guides; weight links open the Hugging Face repositories.
Text-to-speech, voice cloning, and sound effects
| Model | Use it for | Selector | Weights |
|---|---|---|---|
| Fish S2 Pro | Voice cloning and emotion tags | fish-s2-pro |
int8 |
| VibeVoice Large | Speech synthesis and voice cloning | vibevoice |
int8 |
| LongCat AudioDiT | Diffusion speech synthesis | longcat |
int8 |
| OpenMOSS TTS Local | Speech synthesis and voice cloning | moss-local |
int8 |
| MOSS-TTSD | Multi-speaker dialogue | moss-ttsd |
int8 |
| OpenMOSS Sound Effect | Sound effects from text | moss-sound-effect |
4-bit |
| Step-Audio-EditX | Voice cloning and audio editing | step-audio |
int8 |
| DramaBox | Speech synthesis in 48 kHz stereo | dramabox |
BF16¹ |
| dots.tts SOAR | Voice cloning and waveform streaming | dots-tts-soar |
int8 + base |
| dots.tts MeanFlow | Distilled TTS and waveform streaming | dots-tts-mf |
int8 + base |
| FireRedTTS3 Base | Multilingual voice cloning at 24 kHz | fireredtts3-base |
BF16 |
Speech-to-text
| Model | Use it for | Selector | Weights |
|---|---|---|---|
| Cohere Transcribe | Multilingual transcription | cohere-asr |
int8 |
| Qwen3-ASR-1.7B | English, Chinese, and mixed speech | qwen3-asr-1.7b |
int8 · BF16 |
| Confucius4-R2T2 | Real-time streaming transcription, 30 languages | confucius4-r2t2 |
BF16 |
| NVIDIA Nemotron 3.5 ASR Streaming | Multilingual streaming transcription | nemotron-asr-streaming |
int8 |
| IBM Granite Speech 4.0 1B | Speech recognition with a selective-int8 language model | granite-speech-4.0-1b |
int8 |
Loading local weights, shared repositories, and DramaBox components
Flat model repositories accept an alias or a full repository ID.
tts.load("fish-s2-pro") and
tts.load("appautomaton/fishaudio-s2-pro-8bit-mlx") are equivalent. For a
repository containing multiple artifacts, use an alias or specify
artifact_subdir. Original checkpoint paths work where the model guide
documents their layout.
¹ DramaBox also downloads the
Gemma 3 12B text encoder
automatically. Its optional denoise_ref=True setting uses the MLX
RE-USE / SEMamba enhancer
to clean noisy voice references. Denoising is off by default, and the enhancer
weights carry the NSCLv1 non-commercial license. The
DramaBox guide
covers these components and advanced controls.
More examples
Python: voice cloning, streaming transcription, and model discovery
Voice cloning with emotion tags
import mlx_speech
from mlx_speech.audio import write_wav
model = mlx_speech.tts.load("fish-s2-pro")
result = model.generate(
"[excited] This is amazing!",
reference_audio="reference.wav",
reference_text="Transcript of the reference audio.",
)
write_wav("cloned.wav", result.waveform, sample_rate=result.sample_rate)
Live transcription with Confucius4-R2T2
import mlx_speech
asr = mlx_speech.asr.load("confucius4-r2t2")
session = asr.stream_session(language="English")
for pcm in microphone: # float32, 16 kHz mono
print(session.feed(pcm).committed)
print(session.finalize().text)
Streaming transcription with Nemotron
import mlx_speech
from mlx_speech.audio import load_audio
nemotron = mlx_speech.asr.load("nemotron-asr-streaming")
session = nemotron.stream_session(language="en-US", att_context_size=(56, 3))
waveform, _ = load_audio("audio.wav", sample_rate=16_000, mono=True)
for start in range(0, int(waveform.size), 1_600):
session.feed(waveform[start : start + 1_600])
session.finalize()
print(session.result().text)
Granite transcription and model discovery
import mlx_speech
# Granite defaults to the published selective-int8 artifact
granite = mlx_speech.asr.load("granite-speech-4.0-1b")
print(granite.generate("audio.wav").text)
# Discover models
mlx_speech.tts.list_models()
mlx_speech.tts.list_models(detailed=True) # includes shared-repo artifact paths
mlx_speech.asr.list_models()
CLI: waveform streaming, voice cloning, editing, and sound effects
# Bounded waveform streaming with dots.tts
mlx-speech tts --model dots-tts-soar --text "Hello!" --stream -o streamed.wav
# Voice cloning with emotion tags
mlx-speech tts --model fish-s2-pro \
--text "[whisper] Just between us..." \
--reference-audio ref.wav \
--reference-text "Transcript of reference." \
-o cloned.wav
# Step Audio emotion editing
mlx-speech tts --model step-audio \
--reference-audio input.wav \
--reference-text "Transcript." \
--edit-type emotion --edit-info happy \
-o happy.wav
# Sound effect generation
mlx-speech tts --model moss-sound-effect \
--text "rolling thunder with rainfall" \
--duration-seconds 8 \
-o thunder.wav
# Transcribe audio
mlx-speech asr --model cohere-asr --audio speech.wav
mlx-speech asr --model qwen3-asr-1.7b --audio speech.wav --language Chinese
# File transcription with the streaming-capable Nemotron model
mlx-speech asr --model nemotron-asr-streaming --audio speech.wav --language en-US
mlx-speech asr --model granite-speech-4.0-1b --audio speech.wav
mlx-speech asr --model confucius4-r2t2 --audio speech.wav
# Local checkpoint paths work anywhere an alias does
mlx-speech tts --model models/fish_s2_pro/mlx-int8 --text "Hello!" -o output.wav
mlx-speech asr --model models/ibm/granite_4_0_1b_speech/mlx-int8 --audio speech.wav
# Discover models
mlx-speech tts --list-models
mlx-speech asr --list-models
mlx-speech --help
For sampling controls, diffusion steps, batch generation, and other advanced options, follow the model's guide. Each guide names the supported controls and the script that exposes them.
Conversion
Use the published weights to get started. To convert an original checkpoint, follow its model guide for source files, precision options, and the matching conversion script.
Conversion runs separately from inference. Tools needed to read source checkpoints are not runtime requirements.
Development
git clone https://github.com/appautomaton/mlx-speech.git
cd mlx-speech
uv sync
uv run pytest
uv run ruff check .
The default suite needs no checkpoints. Real-weight tests are separate; see the testing guide.
mlx-speech/
src/mlx_speech/ library code
scripts/ conversion, generation, eval, and audit entry points
models/ local checkpoints (not in git)
tests/ unit, checkpoint, runtime, integration tests
docs/ model-family behavior guides
License
Library code is released under the MIT license. Model weights retain their respective licenses, listed in their model cards.
Built and maintained by App Automaton.
Acknowledgements
Thanks to the linux.do community for its inspiration and support.
Release files for mlx-speech 0.5.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlx_speech-0.5.3.tar.gz | 438.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlx_speech-0.5.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.0 MB
Release files / mlx_speech-0.5.3.tar.gz
| Download URL | mlx_speech-0.5.3.tar.gz |
|---|---|
| Size | 438.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1a7f189320321ebdec8eb371236d76c428f2458b6241ca0af59330ba22084c27
|
|
BLAKE2b-256 checksum How to use checksums |
e44dfc998fbb52b1a8ca8971dbb592ea4afdacc7ed2ae763ae4f39210efdc106
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / mlx_speech-0.5.3-py3-none-any.whl
| Download URL | mlx_speech-0.5.3-py3-none-any.whl |
|---|---|
| Size | 572.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6b2b965e542fcc4c2320c23edacbc853eab9231956e3682693372e303e28a430
|
|
BLAKE2b-256 checksum How to use checksums |
74dbd5f574f5f64529495998850285a75a835f303d76f31fb5c684d3aadb0142
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log