mlx-lm

LLMs on Apple silicon with MLX and the Hugging Face Hub

Project description

Generate Text with LLMs and MLX

The easiest way to get started is to install the mlx-lm package:

With pip:

pip install mlx-lm

With conda:

conda install -c conda-forge mlx-lm

The mlx-lm package also has:

Python API

You can use mlx-lm as a module:

from mlx_lm import load, generate

model, tokenizer = load("mlx-community/Mistral-7B-Instruct-v0.3-4bit")

response = generate(model, tokenizer, prompt="hello", verbose=True)

To see a description of all the arguments you can do:

>>> help(generate)

The mlx-lm package also comes with functionality to quantize and optionally upload models to the Hugging Face Hub.

You can convert models in the Python API with:

from mlx_lm import convert

repo = "mistralai/Mistral-7B-Instruct-v0.3"
upload_repo = "mlx-community/My-Mistral-7B-Instruct-v0.3-4bit"

convert(repo, quantize=True, upload_repo=upload_repo)

This will generate a 4-bit quantized Mistral 7B and upload it to the repo mlx-community/My-Mistral-7B-Instruct-v0.3-4bit. It will also save the converted model in the path mlx_model by default.

To see a description of all the arguments you can do:

>>> help(convert)

Streaming

For streaming generation, use the stream_generate function. This returns a generator object which streams the output text. For example,

from mlx_lm import load, stream_generate

repo = "mlx-community/Mistral-7B-Instruct-v0.3-4bit"
model, tokenizer = load(repo)

prompt = "Write a story about Einstein"

for t in stream_generate(model, tokenizer, prompt, max_tokens=512):
    print(t, end="", flush=True)
print()

Command Line

You can also use mlx-lm from the command line with:

mlx_lm.generate --model mistralai/Mistral-7B-Instruct-v0.3 --prompt "hello"

This will download a Mistral 7B model from the Hugging Face Hub and generate text using the given prompt.

For a full list of options run:

mlx_lm.generate --help

To quantize a model from the command line run:

mlx_lm.convert --hf-path mistralai/Mistral-7B-Instruct-v0.3 -q

For more options run:

mlx_lm.convert --help

You can upload new models to Hugging Face by specifying --upload-repo to convert. For example, to upload a quantized Mistral-7B model to the MLX Hugging Face community you can do:

mlx_lm.convert \
    --hf-path mistralai/Mistral-7B-Instruct-v0.3 \
    -q \
    --upload-repo mlx-community/my-4bit-mistral

Supported Models

The example supports Hugging Face format Mistral, Llama, and Phi-2 style models. If the model you want to run is not supported, file an issue or better yet, submit a pull request.

Here are a few examples of Hugging Face models that work with this example:

Most Mistral, Llama, Phi-2, and Mixtral style models should work out of the box.

For some models (such as Qwen and plamo) the tokenizer requires you to enable the trust_remote_code option. You can do this by passing --trust-remote-code in the command line. If you don't specify the flag explicitly, you will be prompted to trust remote code in the terminal when running the model.

For Qwen models you must also specify the eos_token. You can do this by passing --eos-token "<|endoftext|>" in the command line.

These options can also be set in the Python API. For example:

model, tokenizer = load(
    "qwen/Qwen-7B",
    tokenizer_config={"eos_token": "<|endoftext|>", "trust_remote_code": True},
)

Project details

Release history Release notifications | RSS feed

0.31.1

Mar 11, 2026

0.31.0 yanked

Mar 7, 2026

Reason this release was yanked:

Batched KV cache cross contamination

0.30.7

Feb 12, 2026

0.30.6

Feb 4, 2026

0.30.5

Jan 25, 2026

0.30.4

Jan 19, 2026

0.30.2

Jan 6, 2026

0.30.0

Dec 18, 2025

0.29.1

Dec 16, 2025

0.28.4

Dec 3, 2025

0.28.3

Oct 17, 2025

0.28.2

Oct 2, 2025

0.28.1

Sep 27, 2025

0.28.0

Sep 17, 2025

0.27.1

Sep 4, 2025

0.27.0

Aug 29, 2025

0.26.4

Aug 25, 2025

0.26.3

Aug 6, 2025

0.26.2

Jul 30, 2025

0.26.1

Jul 26, 2025

0.26.0

Jul 8, 2025

0.25.3

Jul 1, 2025

0.25.2

Jun 9, 2025

0.25.1

Jun 7, 2025

0.25.0

Jun 2, 2025

0.24.1

May 14, 2025

0.24.0

Apr 28, 2025

0.23.2

Apr 22, 2025

0.23.1

Apr 20, 2025

0.23.0

Apr 18, 2025

0.22.5

Apr 11, 2025

0.22.4

Apr 6, 2025

0.22.3

Apr 3, 2025

0.22.2

Mar 21, 2025

0.22.1

Mar 18, 2025

0.22.0

Mar 13, 2025

0.21.5

Feb 27, 2025

0.21.4

Feb 8, 2025

0.21.3

Feb 7, 2025

0.21.2

Feb 5, 2025

0.21.1

Jan 16, 2025

0.21.0

Jan 10, 2025

0.20.6

Jan 3, 2025

0.20.5

Dec 23, 2024

0.20.4

Dec 13, 2024

0.20.3

Dec 11, 2024

0.20.2

Dec 8, 2024

0.20.1

Nov 25, 2024

0.19.3

Nov 4, 2024

0.19.2

Oct 23, 2024

0.19.1

Oct 14, 2024

0.19.0

Oct 2, 2024

0.18.2

Sep 19, 2024

0.18.1

Aug 30, 2024

0.17.1

Aug 24, 2024

0.17.0

Aug 17, 2024

0.16.1

Jul 23, 2024

0.16.0

Jul 22, 2024

0.15.3

Jul 17, 2024

0.15.2

Jul 8, 2024

0.15.1

Jul 7, 2024

0.15.0

Jun 27, 2024

This version

0.14.3

Jun 3, 2024

0.14.2

Jun 2, 2024

0.14.1

May 31, 2024

0.14.0

May 24, 2024

0.13.1

May 17, 2024

0.13.0

May 10, 2024

0.12.1

Apr 30, 2024

0.12.0

Apr 26, 2024

0.11.0

Apr 23, 2024

0.10.0

Apr 19, 2024

0.9.0

Apr 11, 2024

0.8.0

Apr 8, 2024

0.7.0

Apr 5, 2024

0.6.0

Apr 2, 2024

0.5.0

Mar 25, 2024

0.4.0

Mar 21, 2024

0.3.0

Mar 13, 2024

0.2.0

Mar 13, 2024

0.1.0

Mar 8, 2024

0.0.14

Mar 4, 2024

0.0.13

Feb 21, 2024

0.0.12

Feb 20, 2024

0.0.11

Feb 18, 2024

0.0.10

Feb 13, 2024

0.0.9

Feb 8, 2024

0.0.8

Feb 6, 2024

0.0.7

Feb 4, 2024

0.0.6

Jan 26, 2024

0.0.5

Jan 24, 2024

0.0.3

Jan 15, 2024

0.0.2

Jan 12, 2024

0.0.1

Jan 12, 2024

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mlx_lm-0.14.3.tar.gz (59.9 kB view details)

Uploaded Jun 3, 2024 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

mlx_lm-0.14.3-py3-none-any.whl (82.2 kB view details)

Uploaded Jun 3, 2024 Python 3

File details

Details for the file mlx_lm-0.14.3.tar.gz.

File metadata

Download URL: mlx_lm-0.14.3.tar.gz
Upload date: Jun 3, 2024
Size: 59.9 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.2 CPython/3.9.17

File hashes

Hashes for mlx_lm-0.14.3.tar.gz
Algorithm	Hash digest
SHA256	`b8d09897b6bd5ebaf60a38a48a072fdc64341e49ae13c25653739d7f5c5cff39`
MD5	`c762f4afb90be8b4eef9e1b49e262806`
BLAKE2b-256	`c20a9d54c1f721e04a59f395cedd905ab638b6a163e1bff5b52b0de7364f96b4`

See more details on using hashes here.

File details

Details for the file mlx_lm-0.14.3-py3-none-any.whl.

File metadata

Download URL: mlx_lm-0.14.3-py3-none-any.whl
Upload date: Jun 3, 2024
Size: 82.2 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.2 CPython/3.9.17

File hashes

Hashes for mlx_lm-0.14.3-py3-none-any.whl
Algorithm	Hash digest
SHA256	`0ee4658619d0b2b55fedf04b21258253b244eb754723933f6386637693af3570`
MD5	`5bf559b339d22fc47e821a057c9f88a6`
BLAKE2b-256	`4cf64884761850f305602fc281cdcfe914db7fade67c3ed7991befeb1e8a3cd0`

See more details on using hashes here.

mlx-lm 0.14.3

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Project description

Generate Text with LLMs and MLX

Python API

Streaming

Command Line

Supported Models

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes