Skip to main content

Whisper

Speech recognition with Whisper in MLX. Whisper is a set of open source speech recognition models from OpenAI, ranging from 39 million to 1.5 billion parameters.1

Setup

Install ffmpeg:

# on macOS using Homebrew (https://brew.sh/)
brew install ffmpeg

Install the mlx-whisper package with:

pip install mlx-whisper

Run

CLI

At its simplest:

mlx_whisper audio_file.mp3

This will make a text file audio_file.txt with the results.

Use -f to specify the output format and --model to specify the model. There are many other supported command line options. To see them all, run mlx_whisper -h.

You can also pipe the audio content of other programs via stdin:

some-process | mlx_whisper -

The default output file name will be content.*. You can specify the name with the --output-name flag.

API

Transcribe audio with:

import mlx_whisper

text = mlx_whisper.transcribe(speech_file)["text"]

The default model is "mlx-community/whisper-tiny". Choose the model by setting path_or_hf_repo. For example:

result = mlx_whisper.transcribe(speech_file, path_or_hf_repo="models/large")

This will load the model contained in models/large. The path_or_hf_repo can also point to an MLX-style Whisper model on the Hugging Face Hub. In this case, the model will be automatically downloaded. A collection of pre-converted Whisper models are in the Hugging Face MLX Community.

The transcribe function also supports word-level timestamps. You can generate these with:

output = mlx_whisper.transcribe(speech_file, word_timestamps=True)
print(output["segments"][0]["words"])

To see more transcription options use:

>>> help(mlx_whisper.transcribe)

Converting models

[!TIP] Skip the conversion step by using pre-converted checkpoints from the Hugging Face Hub. There are a few available in the MLX Community organization.

To convert a model, first clone the MLX Examples repo:

git clone https://github.com/ml-explore/mlx-examples.git

Then run convert.py from mlx-examples/whisper. For example, to convert the tiny model use:

python convert.py --torch-name-or-path tiny --mlx-path mlx_models/tiny

Note you can also convert a local PyTorch checkpoint which is in the original OpenAI format.

To generate a 4-bit quantized model, use -q. For a full list of options:

python convert.py --help

By default, the conversion script will make the directory mlx_models and save the converted weights.npz and config.json there.

Each time it is run, convert.py will overwrite any model in the provided path. To save different models, make sure to set --mlx-path to a unique directory for each converted model. For example:

model="tiny"
python convert.py --torch-name-or-path ${model} --mlx-path mlx_models/${model}_fp16
python convert.py --torch-name-or-path ${model} --dtype float32 --mlx-path mlx_models/${model}_fp32
python convert.py --torch-name-or-path ${model} -q --q_bits 4 --mlx-path mlx_models/${model}_quantized_4bits
  1. Refer to the arXiv paper, blog post, and code for more details. ↩

Metadata

Release files for mlx-whisper 0.4.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for mlx-whisper 0.4.3
File Interpreter ABI Platform
mlx_whisper-0.4.3-py3-none-any.whl Python 3 none any Details

Release files / mlx_whisper-0.4.3-py3-none-any.whl

Download URL mlx_whisper-0.4.3-py3-none-any.whl
Size 890.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6b82b6597a994643a3e5496c7bc229a672e5ca308458455bfe276e76ae024489
BLAKE2b-256 checksum
How to use checksums
22b7a35232812a2ccfffcb7614ba96a91338551a660a0e9815cee668bf5743f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.17

Release history Release notifications | RSS feed

This release

0.4.3 This release

1 release file

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page