mlx-speech
Local speech synthesis, editing, and transcription on Apple Silicon, running pure MLX. No cloud, no PyTorch at runtime.
mlx-speech is an App Automaton project.
Project page: appautomaton.renocrypt.com/mlx-speech.
The appautomaton org hosts the code on GitHub
and the converted weights on Hugging Face.
Models
Published MLX weights live under the App Automaton Hugging Face org,
appautomaton, and download automatically
when loaded by alias. Flat model repositories load by alias or full repo id —
tts.load("fish-s2-pro") and
tts.load("appautomaton/fishaudio-s2-pro-8bit-mlx") are equivalent. Shared
multi-artifact repositories use an alias or an explicit artifact_subdir so the
runtime never guesses a variant. Original checkpoint directories can also be
loaded by path when the model-family guide documents their layout. Each model
name links to a guide covering behavior, flags, and known limitations.
Text-to-speech
| Selector | Model | Weights |
|---|---|---|
fish-s2-pro |
Fish S2 Pro — dual-AR TTS, voice cloning, emotion tags | int8 |
vibevoice |
VibeVoice Large — hybrid LLM+diffusion TTS, voice cloning | int8 |
longcat |
LongCat AudioDiT — flow-matching diffusion TTS | int8 |
moss-local |
OpenMOSS TTS Local — local-attention multi-VQ TTS | int8 |
moss-ttsd |
MOSS-TTSD — delay-pattern dialogue TTS | int8 |
moss-sound-effect |
OpenMOSS Sound Effect — text-to-sound-effect generation | 4-bit |
step-audio |
Step-Audio-EditX — voice cloning, audio editing | int8 |
dramabox |
DramaBox — Resemble flow-matching diffusion TTS, 48 kHz stereo | bf16¹ |
dots-tts-soar |
dots.tts SOAR — continuous autoregressive flow-matching TTS and voice cloning | int8 + base |
dots-tts-mf |
dots.tts MeanFlow — distilled continuous autoregressive TTS and voice cloning | int8 + base |
Speech-to-text
| Selector | Model | Weights |
|---|---|---|
cohere-asr |
Cohere Transcribe — multilingual ASR | int8 |
qwen3-asr-1.7b |
Qwen3-ASR-1.7B — English, Chinese, and mixed Chinese/English ASR | int8 · bf16 |
nemotron-asr-streaming |
NVIDIA Nemotron 3.5 ASR Streaming — cache-aware multilingual streaming across three stated quality tiers | int8 |
granite-speech-4.0-1b |
IBM Granite Speech 4.0 1B — selective-int8 Granite LM with BF16 acoustic encoder and QFormer | int8 |
¹ tts.load("dramabox") also pulls the Gemma 3 12B backbone
text encoder automatically. Output is 48 kHz stereo. For advanced controls (cfg,
steps, voice reference) use scripts/generate_dramabox.py. Optional
denoise_ref=True cleans a noisy voice reference with the pure-MLX
RE-USE / SEMamba enhancer
(off by default; NSCLv1 non-commercial weights). See
docs/dramabox.md.
Installation
Requires an Apple Silicon Mac (M1 or later) and Python 3.13+.
pip install mlx-speech
Quick Start
Python:
import mlx_speech
from mlx_speech.audio import load_audio, write_wav
# Text-to-speech
model = mlx_speech.tts.load("fish-s2-pro")
result = model.generate("Hello from mlx-speech!")
write_wav("output.wav", result.waveform, sample_rate=result.sample_rate)
# Voice cloning with emotion tags
result = model.generate(
"[excited] This is amazing!",
reference_audio="reference.wav",
reference_text="Transcript of the reference audio.",
)
# Speech-to-text
asr = mlx_speech.asr.load("qwen3-asr-1.7b")
print(asr.generate("audio.wav").text)
# Cache-aware incremental ASR is available on Nemotron
nemotron = mlx_speech.asr.load("nemotron-asr-streaming")
session = nemotron.stream_session(language="en-US", att_context_size=(56, 3))
waveform, _ = load_audio("audio.wav", sample_rate=16_000, mono=True)
for start in range(0, int(waveform.size), 1_600):
session.feed(waveform[start : start + 1_600])
session.finalize()
print(session.result().text)
# Granite defaults to the published selective-int8 artifact
granite = mlx_speech.asr.load("granite-speech-4.0-1b")
print(granite.generate("audio.wav").text)
# Discover models
mlx_speech.tts.list_models()
mlx_speech.tts.list_models(detailed=True) # includes shared-repo artifact paths
mlx_speech.asr.list_models()
CLI:
# Generate speech
mlx-speech tts --model fish-s2-pro --text "Hello!" -o output.wav
# Bounded waveform streaming with dots.tts
mlx-speech tts --model dots-tts-soar --text "Hello!" --stream -o streamed.wav
# Voice cloning with emotion tags
mlx-speech tts --model fish-s2-pro \
--text "[whisper] Just between us..." \
--reference-audio ref.wav \
--reference-text "Transcript of reference." \
-o cloned.wav
# Step Audio emotion editing
mlx-speech tts --model step-audio \
--reference-audio input.wav \
--reference-text "Transcript." \
--edit-type emotion --edit-info happy \
-o happy.wav
# Sound effect generation
mlx-speech tts --model moss-sound-effect \
--text "rolling thunder with rainfall" \
--duration-seconds 8 \
-o thunder.wav
# Transcribe audio
mlx-speech asr --model cohere-asr --audio speech.wav
mlx-speech asr --model qwen3-asr-1.7b --audio speech.wav --language Chinese
# File transcription with the streaming-capable Nemotron model
mlx-speech asr --model nemotron-asr-streaming --audio speech.wav --language en-US
mlx-speech asr --model granite-speech-4.0-1b --audio speech.wav
# Local checkpoint paths work anywhere an alias does
mlx-speech tts --model models/fish_s2_pro/mlx-int8 --text "Hello!" -o output.wav
mlx-speech asr --model models/ibm/granite_4_0_1b_speech/mlx-int8 --audio speech.wav
# Discover models
mlx-speech tts --list-models
mlx-speech asr --list-models
mlx-speech --help
Note: The
mlx-speechCLI covers the common generation, voice cloning, editing, waveform streaming, and transcription paths. For advanced controls (sampling temperature, top-p/k, diffusion steps, batch JSONL, duration tuning, etc.) use the family-specific scripts inscripts/where provided. Each model guide indocs/names its canonical advanced entry point and supported controls.
Conversion
Available model-family conversion entry points include:
python scripts/convert/fish_s2_pro.py
python scripts/convert/longcat_audiodit.py
python scripts/convert/vibevoice.py
python scripts/convert/moss_local.py
python scripts/convert/moss_ttsd.py
python scripts/convert/moss_sound_effect.py
python scripts/convert/step_audio_editx.py
python scripts/convert/cohere_asr.py
python scripts/convert/qwen3_asr.py
python scripts/convert/granite_speech_asr.py
python scripts/convert/dots_tts.py --variant all --precision int8
uv run --with torch python scripts/convert/nemotron_asr.py --quant int8
Conversion is an offline workflow and may require source-format-specific tools;
those tools are not runtime dependencies. Granite conversion reads the original
sharded BF16 safetensors directly and writes a self-contained selective-int8 MLX
artifact without PyTorch or mlx-audio.
Development
git clone https://github.com/appautomaton/mlx-speech.git
cd mlx-speech
uv sync
uv run pytest tests/unit/
uv run ruff check .
mlx-speech/
src/mlx_speech/ library code
scripts/ conversion, generation, eval, and audit entry points
models/ local checkpoints (not in git)
tests/ unit, checkpoint, runtime, integration tests
docs/ model-family behavior guides
License
MIT — see LICENSE
Built and maintained by App Automaton.
Acknowledgements
This project wouldn't exist without the inspiration and generous support of the incredible community at linux.do.
Release files for mlx-speech 0.5.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlx_speech-0.5.2.tar.gz | 403.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlx_speech-0.5.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 931.5 kB
Release files / mlx_speech-0.5.2.tar.gz
| Download URL | mlx_speech-0.5.2.tar.gz |
|---|---|
| Size | 403.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d462d1309d4f998c757fb82dfd4d6dad4d3c16d73f1f4439d446dd1c4956a0da
|
|
BLAKE2b-256 checksum How to use checksums |
bc945ab077f39695a23854c120aad1c65064f90b4638ef22b168b0adec8d9031
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 9, 2026.
Transparency logRelease files / mlx_speech-0.5.2-py3-none-any.whl
| Download URL | mlx_speech-0.5.2-py3-none-any.whl |
|---|---|
| Size | 528.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0d267bc891953fcb67b0d5bea7a85bd4459195be60f5770918fd5d95e7215590
|
|
BLAKE2b-256 checksum How to use checksums |
b5a82f95d29809d0dbee252c55b57cc99849527266e8915bf64a9bb660f137f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 9, 2026.
Transparency log