Generate Text with LLMs and MLX
The easiest way to get started is to install the mlx-lm package:
pip install mlx-lm
Python API
You can use mlx-lm as a module:
from mlx_lm import load, generate
model, tokenizer = load("mistralai/Mistral-7B-v0.1")
response = generate(model, tokenizer, prompt="hello", verbose=True)
To see a description of all the arguments you can do:
>>> help(generate)
The mlx-lm package also comes with functionality to quantize and optionally
upload models to the Hugging Face Hub.
You can convert models in the Python API with:
from mlx_lm import convert
upload_repo = "mlx-community/My-Mistral-7B-v0.1-4bit"
convert("mistralai/Mistral-7B-v0.1", quantize=True, upload_repo=upload_repo)
This will generate a 4-bit quantized Mistral-7B and upload it to the
repo mlx-community/My-Mistral-7B-v0.1-4bit. It will also save the
converted model in the path mlx_model by default.
To see a description of all the arguments you can do:
>>> help(convert)
Command Line
You can also use mlx-lm from the command line with:
python -m mlx_lm.generate --model mistralai/Mistral-7B-v0.1 --prompt "hello"
This will download a Mistral 7B model from the Hugging Face Hub and generate text using the given prompt.
For a full list of options run:
python -m mlx_lm.generate --help
To quantize a model from the command line run:
python -m mlx_lm.convert --hf-path mistralai/Mistral-7B-v0.1 -q
For more options run:
python -m mlx_lm.convert --help
You can upload new models to Hugging Face by specifying --upload-repo to
convert. For example, to upload a quantized Mistral-7B model to the
MLX Hugging Face community you can do:
python -m mlx_lm.convert \
--hf-path mistralai/Mistral-7B-v0.1 \
-q \
--upload-repo mlx-community/my-4bit-mistral
Supported Models
The example supports Hugging Face format Mistral, Llama, and Phi-2 style models. If the model you want to run is not supported, file an issue or better yet, submit a pull request.
Here are a few examples of Hugging Face models that work with this example:
- mistralai/Mistral-7B-v0.1
- meta-llama/Llama-2-7b-hf
- deepseek-ai/deepseek-coder-6.7b-instruct
- 01-ai/Yi-6B-Chat
- microsoft/phi-2
- mistralai/Mixtral-8x7B-Instruct-v0.1
Most Mistral, Llama, Phi-2 and Mixtral style models should work out of the box.
Release files for mlx-lm 0.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlx-lm-0.0.3.tar.gz | 11.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlx_lm-0.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 26.1 kB
Release files / mlx-lm-0.0.3.tar.gz
| Download URL | mlx-lm-0.0.3.tar.gz |
|---|---|
| Size | 11.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
57e11b6e359ebb496e0fa302e96b63af65855da53ea15ce43d05b08ef7f109d9
|
|
BLAKE2b-256 checksum How to use checksums |
3beda363caf42ad94238f04c0ded359a664a19f3bf5d309cb201d8f16c98c583
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.9.17
|
Release files / mlx_lm-0.0.3-py3-none-any.whl
| Download URL | mlx_lm-0.0.3-py3-none-any.whl |
|---|---|
| Size | 14.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5981276d73b891562ac5dffe8318f8cba4429535f69b56509f41ac9ec294eeb7
|
|
BLAKE2b-256 checksum How to use checksums |
64fe1ca7f2eda4bdb7d0ee5382b0ef06c400505ab56bab20084f7b92b53e94cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.9.17
|