whisper(ml)x
Fast, accurate speech recognition on Apple Silicon — powered by MLX.
A fork of WhisperX with the inference backend replaced by mlx-whisper, running natively on Apple Silicon via MLX. Word-level timestamps, speaker diarization, and VAD are all retained.
- ⚡️ MLX inference — runs on Apple Silicon GPU via unified memory
- 🎯 Word-level timestamps via wav2vec2 forced alignment
- 👥 Speaker diarization via pyannote-audio
- 🗣️ VAD preprocessing via pyannote or silero
Installation
pip install whispermlx
Or with uv:
uv add whispermlx
Usage
CLI
# Auto-downloads mlx-community/whisper-large-v3-mlx on first run
whispermlx audio.mp3 --model large-v3
# With speaker diarization
whispermlx audio.mp3 --model large-v3 --diarize --hf_token YOUR_TOKEN
# Use any mlx-community model directly
whispermlx audio.mp3 --model mlx-community/whisper-large-v3-turbo
# Carry transcript context across VAD chunks (better punctuation / proper nouns)
whispermlx audio.mp3 --model large-v3 --interleaved_context
Python
import whispermlx
# Short name — auto-maps to mlx-community/whisper-large-v3-mlx
model = whispermlx.load_model("large-v3", device="cpu")
result = model.transcribe("audio.mp3")
print(result["segments"])
# With alignment
model_a, metadata = whispermlx.load_align_model(language_code=result["language"], device="cpu")
result = whispermlx.align(result["segments"], model_a, metadata, "audio.mp3", device="cpu")
# With diarization
from whispermlx.diarize import DiarizationPipeline
diarize_model = DiarizationPipeline(token="YOUR_HF_TOKEN", device="cpu")
diarize_segments = diarize_model("audio.mp3")
result = whispermlx.assign_word_speakers(diarize_segments, result)
Model Names
Short names are automatically mapped to their mlx-community equivalents. Full HF repo IDs also work.
| Short name | HF repo |
|---|---|
tiny, base, small, medium |
mlx-community/whisper-{name}-mlx |
large-v3 |
mlx-community/whisper-large-v3-mlx |
large-v3-turbo / turbo |
mlx-community/whisper-large-v3-turbo |
Comparison with WhisperX
whispermlx tracks upstream WhisperX's pipeline and API, with the inference backend swapped to MLX. Everything downstream of ASR (alignment, diarization, output formats) is unchanged.
Feature parity
| Feature | WhisperX | whispermlx | Notes |
|---|---|---|---|
| VAD preprocessing | ✅ pyannote / silero | ✅ pyannote / silero | Identical |
| Word-level timestamps | ✅ wav2vec2 alignment | ✅ wav2vec2 alignment | Identical; adds Indonesian (id) model |
| Speaker diarization | ✅ pyannote-audio | ✅ pyannote-audio | Identical |
| Output formats | srt, vtt, txt, tsv, json, aud | srt, vtt, txt, tsv, json, aud | Identical |
| CLI flags | Full set | Full set | Identical; adds --log-level |
| Python API | load_model, transcribe, align, assign_word_speakers |
Same signatures | Drop-in compatible |
| Batched inference | ✅ context-aware batching | ❌ per-segment | batch_size accepted but unused; --interleaved_context runs as sequential context carry-over instead of batched streams |
| CUDA / NVIDIA GPU | ✅ | ❌ | Apple Silicon only |
| Apple Silicon GPU | ❌ (CPU only) | ✅ MLX unified memory | Native, no CUDA |
API compatibility
| Parameter | WhisperX | whispermlx |
|---|---|---|
device |
Controls ASR + PyTorch models | Controls VAD/alignment/diarization only; MLX inference auto-uses GPU |
compute_type |
float16 / float32 / int8 | Accepted, ignored |
device_index, threads, download_root, local_files_only, use_auth_token |
Used | Accepted for compatibility, ignored |
asr_options |
Full faster-whisper options | Only initial_prompt is used |
--model |
faster-whisper model names | Short names map to mlx-community repos; full HF repo IDs also work |
Dependencies
| Purpose | WhisperX | whispermlx |
|---|---|---|
| ASR backend | ctranslate2 + faster-whisper | mlx-whisper |
| Alignment | transformers (wav2vec2) | transformers (wav2vec2) |
| VAD + diarization | pyannote-audio | pyannote-audio |
| Torch | CUDA/CPU wheels | CPU wheels only |
| Extra | triton, torchcodec | numba, tqdm |
Speaker Diarization
Requires a Hugging Face access token and acceptance of the pyannote speaker-diarization-community-1 model agreement.
Acknowledgements
Built on top of WhisperX by Max Bain et al., mlx-whisper, pyannote-audio, and OpenAI Whisper.
@article{bain2022whisperx,
title={WhisperX: Time-Accurate Speech Transcription of Long-Form Audio},
author={Bain, Max and Huh, Jaesung and Han, Tengda and Zisserman, Andrew},
journal={INTERSPEECH 2023},
year={2023}
}
Metadata
Release files for whispermlx 3.14.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| whispermlx-3.14.0.tar.gz | 16.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| whispermlx-3.14.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.0 MB
Release files / whispermlx-3.14.0.tar.gz
| Download URL | whispermlx-3.14.0.tar.gz |
|---|---|
| Size | 16.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7b0b6a8ad914a9d27e9b46c71c87f83e271d4a33d870a62a5ab4c4a4738d8c62
|
|
BLAKE2b-256 checksum How to use checksums |
01a34c0e6b304a7ca0b6fec415111ed913be054b5937a473e6019998a4b62d08
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / whispermlx-3.14.0-py3-none-any.whl
| Download URL | whispermlx-3.14.0-py3-none-any.whl |
|---|---|
| Size | 16.5 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c68551d990b04cd32ec615eba59982a0f7e998102111ff45836a4efef02d0e2d
|
|
BLAKE2b-256 checksum How to use checksums |
3287e0ebecd4891be90f5eb0df98826652e537e6607e0fd6fdf56994c554ae90
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|