Skip to main content

mlx-speech

Text-to-speech and speech recognition on Apple Silicon

Voice cloning, audio editing, sound effects, and transcription. All running locally in MLX.

PyPI Python 3.13+ License: MIT CI

Quick start · Models · Model guides · Hugging Face weights · Project website

Website contributions belong in the frontend repository. See website ownership.

mlx-speech is an open-source Python library for text-to-speech (TTS) and automatic speech recognition (ASR) on Apple Silicon Macs. Models share a Python API and command-line interface, with inference implemented in Apple's MLX framework.

Audio processing stays on your Mac. Inference needs neither PyTorch nor a cloud service. Published weights download on first use, and local checkpoint paths support offline loading.

Installation

Requires an Apple Silicon Mac (M1 or later) and Python 3.13+.

pip install mlx-speech

Quick Start

Python:

import mlx_speech
from mlx_speech.audio import write_wav

# Text-to-speech
model = mlx_speech.tts.load("fish-s2-pro")
result = model.generate("Hello from mlx-speech!")
write_wav("output.wav", result.waveform, sample_rate=result.sample_rate)

# Speech-to-text
asr = mlx_speech.asr.load("qwen3-asr-1.7b")
print(asr.generate("audio.wav").text)

CLI:

mlx-speech tts --model fish-s2-pro --text "Hello!" -o output.wav
mlx-speech asr --model qwen3-asr-1.7b --audio speech.wav

Models

Pass the selector to tts.load(), asr.load(), or --model. Names link to guides; weight links open the Hugging Face repositories.

Text-to-speech, voice cloning, and sound effects

Model Use it for Selector Weights
Fish S2 Pro Voice cloning and emotion tags fish-s2-pro int8
VibeVoice Large Speech synthesis and voice cloning vibevoice int8
LongCat AudioDiT Diffusion speech synthesis longcat int8
OpenMOSS TTS Local Speech synthesis and voice cloning moss-local int8
MOSS-TTSD Multi-speaker dialogue moss-ttsd int8
OpenMOSS Sound Effect Sound effects from text moss-sound-effect 4-bit
Step-Audio-EditX Voice cloning and audio editing step-audio int8
DramaBox Speech synthesis in 48 kHz stereo dramabox BF16¹
dots.tts SOAR Voice cloning and waveform streaming dots-tts-soar int8 + base
dots.tts MeanFlow Distilled TTS and waveform streaming dots-tts-mf int8 + base
FireRedTTS3 Base Multilingual voice cloning at 24 kHz fireredtts3-base BF16

Speech-to-text

Model Use it for Selector Weights
Cohere Transcribe Multilingual transcription cohere-asr int8
Qwen3-ASR-1.7B English, Chinese, and mixed speech qwen3-asr-1.7b int8 · BF16
Confucius4-R2T2 Real-time streaming transcription, 30 languages confucius4-r2t2 BF16
NVIDIA Nemotron 3.5 ASR Streaming Multilingual streaming transcription nemotron-asr-streaming int8
IBM Granite Speech 4.0 1B Speech recognition with a selective-int8 language model granite-speech-4.0-1b int8
Loading local weights, shared repositories, and DramaBox components

Flat model repositories accept an alias or a full repository ID. tts.load("fish-s2-pro") and tts.load("appautomaton/fishaudio-s2-pro-8bit-mlx") are equivalent. For a repository containing multiple artifacts, use an alias or specify artifact_subdir. Original checkpoint paths work where the model guide documents their layout.

¹ DramaBox also downloads the Gemma 3 12B text encoder automatically. Its optional denoise_ref=True setting uses the MLX RE-USE / SEMamba enhancer to clean noisy voice references. Denoising is off by default, and the enhancer weights carry the NSCLv1 non-commercial license. The DramaBox guide covers these components and advanced controls.

More examples

Python: voice cloning, streaming transcription, and model discovery

Voice cloning with emotion tags

import mlx_speech
from mlx_speech.audio import write_wav

model = mlx_speech.tts.load("fish-s2-pro")
result = model.generate(
    "[excited] This is amazing!",
    reference_audio="reference.wav",
    reference_text="Transcript of the reference audio.",
)
write_wav("cloned.wav", result.waveform, sample_rate=result.sample_rate)

Live transcription with Confucius4-R2T2

import mlx_speech

asr = mlx_speech.asr.load("confucius4-r2t2")
session = asr.stream_session(language="English")
for pcm in microphone:                 # float32, 16 kHz mono
    print(session.feed(pcm).committed)
print(session.finalize().text)

Streaming transcription with Nemotron

import mlx_speech
from mlx_speech.audio import load_audio

nemotron = mlx_speech.asr.load("nemotron-asr-streaming")
session = nemotron.stream_session(language="en-US", att_context_size=(56, 3))
waveform, _ = load_audio("audio.wav", sample_rate=16_000, mono=True)
for start in range(0, int(waveform.size), 1_600):
    session.feed(waveform[start : start + 1_600])
session.finalize()
print(session.result().text)

Granite transcription and model discovery

import mlx_speech

# Granite defaults to the published selective-int8 artifact
granite = mlx_speech.asr.load("granite-speech-4.0-1b")
print(granite.generate("audio.wav").text)

# Discover models
mlx_speech.tts.list_models()
mlx_speech.tts.list_models(detailed=True)  # includes shared-repo artifact paths
mlx_speech.asr.list_models()
CLI: waveform streaming, voice cloning, editing, and sound effects
# Bounded waveform streaming with dots.tts
mlx-speech tts --model dots-tts-soar --text "Hello!" --stream -o streamed.wav

# Voice cloning with emotion tags
mlx-speech tts --model fish-s2-pro \
  --text "[whisper] Just between us..." \
  --reference-audio ref.wav \
  --reference-text "Transcript of reference." \
  -o cloned.wav

# Step Audio emotion editing
mlx-speech tts --model step-audio \
  --reference-audio input.wav \
  --reference-text "Transcript." \
  --edit-type emotion --edit-info happy \
  -o happy.wav

# Sound effect generation
mlx-speech tts --model moss-sound-effect \
  --text "rolling thunder with rainfall" \
  --duration-seconds 8 \
  -o thunder.wav

# Transcribe audio
mlx-speech asr --model cohere-asr --audio speech.wav
mlx-speech asr --model qwen3-asr-1.7b --audio speech.wav --language Chinese
# File transcription with the streaming-capable Nemotron model
mlx-speech asr --model nemotron-asr-streaming --audio speech.wav --language en-US
mlx-speech asr --model granite-speech-4.0-1b --audio speech.wav
mlx-speech asr --model confucius4-r2t2 --audio speech.wav

# Local checkpoint paths work anywhere an alias does
mlx-speech tts --model models/fish_s2_pro/mlx-int8 --text "Hello!" -o output.wav
mlx-speech asr --model models/ibm/granite_4_0_1b_speech/mlx-int8 --audio speech.wav

# Discover models
mlx-speech tts --list-models
mlx-speech asr --list-models
mlx-speech --help

For sampling controls, diffusion steps, batch generation, and other advanced options, follow the model's guide. Each guide names the supported controls and the script that exposes them.

Conversion

Use the published weights to get started. To convert an original checkpoint, follow its model guide for source files, precision options, and the matching conversion script.

Conversion runs separately from inference. Tools needed to read source checkpoints are not runtime requirements.

Development

git clone https://github.com/appautomaton/mlx-speech.git
cd mlx-speech
uv sync
uv run pytest
uv run ruff check .

The default suite needs no checkpoints. Real-weight tests are separate; see the testing guide.

mlx-speech/
  src/mlx_speech/     library code
  scripts/           conversion, generation, eval, and audit entry points
  models/            local checkpoints (not in git)
  tests/             unit, checkpoint, runtime, integration tests
  docs/              model-family behavior guides

License

Library code is released under the MIT license. Model weights retain their respective licenses, listed in their model cards.

Built and maintained by App Automaton.

Acknowledgements

Thanks to the linux.do community for its inspiration and support.

Release files for mlx-speech 0.5.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mlx-speech 0.5.3
File Size Uploaded
mlx_speech-0.5.3.tar.gz 438.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mlx-speech 0.5.3
File Interpreter ABI Platform
mlx_speech-0.5.3-py3-none-any.whl Python 3 none any Details

Total release size: 1.0 MB

Release files / mlx_speech-0.5.3.tar.gz

Download URL mlx_speech-0.5.3.tar.gz
Size 438.4 kB
Tags Source
SHA-256 checksum
How to use checksums
1a7f189320321ebdec8eb371236d76c428f2458b6241ca0af59330ba22084c27
BLAKE2b-256 checksum
How to use checksums
e44dfc998fbb52b1a8ca8971dbb592ea4afdacc7ed2ae763ae4f39210efdc106
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / mlx_speech-0.5.3-py3-none-any.whl

Download URL mlx_speech-0.5.3-py3-none-any.whl
Size 572.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6b2b965e542fcc4c2320c23edacbc853eab9231956e3682693372e303e28a430
BLAKE2b-256 checksum
How to use checksums
74dbd5f574f5f64529495998850285a75a835f303d76f31fb5c684d3aadb0142
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.3 This release

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page