Skip to main content

whisper(ml)x

Fast, accurate speech recognition on Apple Silicon — powered by MLX.

Documentation

A fork of WhisperX with the inference backend replaced by mlx-whisper, running natively on Apple Silicon via MLX. Word-level timestamps, speaker diarization, and VAD are all retained.

  • ⚡️ MLX inference — runs on Apple Silicon GPU via unified memory
  • 🎯 Word-level timestamps via wav2vec2 forced alignment
  • 👥 Speaker diarization via pyannote-audio
  • 🗣️ VAD preprocessing via pyannote or silero

Installation

pip install whispermlx

Or with uv:

uv add whispermlx

Usage

CLI

# Auto-downloads mlx-community/whisper-large-v3-mlx on first run
whispermlx audio.mp3 --model large-v3

# With speaker diarization
whispermlx audio.mp3 --model large-v3 --diarize --hf_token YOUR_TOKEN

# Use any mlx-community model directly
whispermlx audio.mp3 --model mlx-community/whisper-large-v3-turbo

# Carry transcript context across VAD chunks (better punctuation / proper nouns)
whispermlx audio.mp3 --model large-v3 --interleaved_context

Python

import whispermlx

# Short name — auto-maps to mlx-community/whisper-large-v3-mlx
model = whispermlx.load_model("large-v3", device="cpu")
result = model.transcribe("audio.mp3")
print(result["segments"])

# With alignment
model_a, metadata = whispermlx.load_align_model(language_code=result["language"], device="cpu")
result = whispermlx.align(result["segments"], model_a, metadata, "audio.mp3", device="cpu")

# With diarization
from whispermlx.diarize import DiarizationPipeline
diarize_model = DiarizationPipeline(token="YOUR_HF_TOKEN", device="cpu")
diarize_segments = diarize_model("audio.mp3")
result = whispermlx.assign_word_speakers(diarize_segments, result)

Model Names

Short names are automatically mapped to their mlx-community equivalents. Full HF repo IDs also work.

Short name HF repo
tiny, base, small, medium mlx-community/whisper-{name}-mlx
large-v3 mlx-community/whisper-large-v3-mlx
large-v3-turbo / turbo mlx-community/whisper-large-v3-turbo

Comparison with WhisperX

whispermlx tracks upstream WhisperX's pipeline and API, with the inference backend swapped to MLX. Everything downstream of ASR (alignment, diarization, output formats) is unchanged.

Feature parity

Feature WhisperX whispermlx Notes
VAD preprocessing ✅ pyannote / silero ✅ pyannote / silero Identical
Word-level timestamps ✅ wav2vec2 alignment ✅ wav2vec2 alignment Identical; adds Indonesian (id) model
Speaker diarization ✅ pyannote-audio ✅ pyannote-audio Identical
Output formats srt, vtt, txt, tsv, json, aud srt, vtt, txt, tsv, json, aud Identical
CLI flags Full set Full set Identical; adds --log-level
Python API load_model, transcribe, align, assign_word_speakers Same signatures Drop-in compatible
Batched inference ✅ context-aware batching ❌ per-segment batch_size accepted but unused; --interleaved_context runs as sequential context carry-over instead of batched streams
CUDA / NVIDIA GPU ✅ ❌ Apple Silicon only
Apple Silicon GPU ❌ (CPU only) ✅ MLX unified memory Native, no CUDA

API compatibility

Parameter WhisperX whispermlx
device Controls ASR + PyTorch models Controls VAD/alignment/diarization only; MLX inference auto-uses GPU
compute_type float16 / float32 / int8 Accepted, ignored
device_index, threads, download_root, local_files_only, use_auth_token Used Accepted for compatibility, ignored
asr_options Full faster-whisper options Only initial_prompt is used
--model faster-whisper model names Short names map to mlx-community repos; full HF repo IDs also work

Dependencies

Purpose WhisperX whispermlx
ASR backend ctranslate2 + faster-whisper mlx-whisper
Alignment transformers (wav2vec2) transformers (wav2vec2)
VAD + diarization pyannote-audio pyannote-audio
Torch CUDA/CPU wheels CPU wheels only
Extra triton, torchcodec numba, tqdm

Speaker Diarization

Requires a Hugging Face access token and acceptance of the pyannote speaker-diarization-community-1 model agreement.

Acknowledgements

Built on top of WhisperX by Max Bain et al., mlx-whisper, pyannote-audio, and OpenAI Whisper.

@article{bain2022whisperx,
  title={WhisperX: Time-Accurate Speech Transcription of Long-Form Audio},
  author={Bain, Max and Huh, Jaesung and Han, Tengda and Zisserman, Andrew},
  journal={INTERSPEECH 2023},
  year={2023}
}

Metadata

Release files for whispermlx 3.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for whispermlx 3.14.0
File Size Uploaded
whispermlx-3.14.0.tar.gz 16.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for whispermlx 3.14.0
File Interpreter ABI Platform
whispermlx-3.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 33.0 MB

Release files / whispermlx-3.14.0.tar.gz

Download URL whispermlx-3.14.0.tar.gz
Size 16.5 MB
Tags Source
SHA-256 checksum
How to use checksums
7b0b6a8ad914a9d27e9b46c71c87f83e271d4a33d870a62a5ab4c4a4738d8c62
BLAKE2b-256 checksum
How to use checksums
01a34c0e6b304a7ca0b6fec415111ed913be054b5937a473e6019998a4b62d08
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / whispermlx-3.14.0-py3-none-any.whl

Download URL whispermlx-3.14.0-py3-none-any.whl
Size 16.5 MB
Tags Python 3
SHA-256 checksum
How to use checksums
c68551d990b04cd32ec615eba59982a0f7e998102111ff45836a4efef02d0e2d
BLAKE2b-256 checksum
How to use checksums
3287e0ebecd4891be90f5eb0df98826652e537e6607e0fd6fdf56994c554ae90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

3.14.0 This release

2 release files

3.12.2

2 release files

3.11.0

2 release files

3.10.1

2 release files

3.10.0

2 release files

3.9.3

2 release files

3.9.2

2 release files

3.9.1

2 release files

3.9.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page