Moshi - MLX
See the top-level README.md for more information on Moshi.
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec. Mimi operates at a framerate of 12.5 Hz, and compresses 24 kHz audio down to 1.1 kbps, in a fully streaming manner (latency of 80ms, the frame size), yet performs better than existing, non-streaming, codec.
This is the MLX implementation for Moshi. For Mimi, this uses our Rust based implementation through the Python binding provided in rustymimi, available in the rust/ folder of our main repository.
Requirements
You will need at least Python 3.10, we recommend Python 3.12.
pip install moshi_mlx # moshi MLX, from PyPI, best with Python 3.12.
# Or the bleeding edge versions for Moshi and Moshi-MLX.
pip install -e "git+https://git@github.com/kyutai-labs/moshi#egg=moshi_mlx&subdirectory=moshi_mlx"
We have tested the MLX version with MacBook Pro M3.
If you are not using Python 3.12, you might get an error when installing
moshi_mlx or rustymimi (which moshi_mlx depends on). Then,you will need to install the Rust toolchain, or switch to Python 3.12.
Usage
Once you have installed moshi_mlx, you can run
python -m moshi_mlx.local -q 4 # weights quantized to 4 bits
python -m moshi_mlx.local -q 8 # weights quantized to 8 bits
# And using a different pretrained model:
python -m moshi_mlx.local -q 4 --hf-repo kyutai/moshika-mlx-q4
python -m moshi_mlx.local -q 8 --hf-repo kyutai/moshika-mlx-q8
# be careful to always match the `-q` and `--hf-repo` flag.
This uses a command line interface, which is barebone. It does not perform any echo cancellation, nor does it try to compensate for a growing lag by skipping frames.
You can use --hf-repo to select a different pretrained model, by setting the proper Hugging Face repository.
See the model list for a reference of the available models.
Alternatively you can use python -m moshi_mlx.local_web to use
the web UI, the connection is via http, at localhost:8998.
License
The present code is provided under the MIT license.
Citation
If you use either Mimi or Moshi, please cite the following paper,
@techreport{kyutai2024moshi,
author = {Alexandre D\'efossez and Laurent Mazar\'e and Manu Orsini and Am\'elie Royer and
Patrick P\'erez and Herv\'e J\'egou and Edouard Grave and Neil Zeghidour},
title = {Moshi: a speech-text foundation model for real-time dialogue},
institution = {Kyutai},
year={2024},
month={September},
url={http://kyutai.org/Moshi.pdf},
}
Release files for moshi-mlx 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| moshi_mlx-0.3.0.tar.gz | 46.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| moshi_mlx-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 102.2 kB
Release files / moshi_mlx-0.3.0.tar.gz
| Download URL | moshi_mlx-0.3.0.tar.gz |
|---|---|
| Size | 46.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
108f0812dfa248cd543de9a46ec11bb2d603387ec3fc92d171157875c783971e
|
|
BLAKE2b-256 checksum How to use checksums |
3404173cf1fc12f73cfc49321e64b45ea1ea15143b49ef570b7171906a45b5a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.5
|
Release files / moshi_mlx-0.3.0-py3-none-any.whl
| Download URL | moshi_mlx-0.3.0-py3-none-any.whl |
|---|---|
| Size | 55.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2ec210340df8bf34025df6c272ea3ed618f5c22e176c3513f5dba4e529420621
|
|
BLAKE2b-256 checksum How to use checksums |
41753c00e7392a7ea80887440da615c300679636f9e6f89da10e1082f4970a95
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.5
|