Skip to main content

mlx-optiq

Run any LLM locally on your Mac. Quantize it, serve it, and code with it.

Website: https://mlx-optiq.com  |  Docs: https://mlx-optiq.com/docs/  |  Models: https://mlx-optiq.com/models  |  Blog: https://mlx-optiq.com/blog/  |  HF org: https://huggingface.co/mlx-community

mlx-optiq is the local-LLM stack for Apple Silicon: an optimizing compiler and runtime for MLX that turns a full-precision model into the best version for a given memory and latency budget on your Mac, using per-layer sensitivity measurement instead of uniform 4-bit everywhere. The same signal drives weights, KV cache, LoRA fine-tuning, and runtime adapter swapping.

One pip install gives you three ways to work with a local model: the CLI (quantize, serve, fine-tune), OptiQ Lab (a local web workbench), and OptiQ Code (a terminal coding agent that drives your served model).

pip install mlx-optiq

Python 3.11+. Quantizing and local inference need Apple Silicon; optiq code and optiq lab also run on Linux and Windows against any OpenAI-compatible base_url.

What it does

  • Mixed-precision weight quantization that beats uniform 4-bit at the same size. optiq convert measures each layer's sensitivity and allocates bits per layer. A static method assigns bits by architecture for models too large to measure. Methods.
  • SSD expert streaming runs large MoE quants that don't fit in RAM. A 2-bit Qwen3.5-122B-A10B runs on a 36 GB Mac at ~12 GB resident, experts streamed off disk. How.
  • Mixed-precision KV cache for longer context at lower memory. optiq serve runs a per-layer KV quant pipeline.
  • One server, two protocols. optiq serve speaks both the OpenAI and Anthropic APIs from one process. Point Claude Code or either SDK at the same local URL.
  • Speculative decoding via bundled MTP heads or paired drafters (--mtp, --drafter).
  • Distributed inference across Macs. optiq cluster serve shards a model's layers across two or more Macs over Thunderbolt and exposes one OpenAI endpoint. A 2-bit Qwen3.5-122B-A10B (42.8 GiB) runs fully resident across a 36 GB + 24 GB pair at ~20 tok/s, against 4.9 tok/s streaming experts off SSD on one Mac. How.
  • Sensitivity-aware LoRA (SFT + DPO) and runtime hot-swap adapters.
  • OptiQ Lab (pip install mlx-optiq then optiq lab): a local web UI for chat, quantize, fine-tune, and dataset work. Product.
  • OptiQ Code (optiq code): a terminal coding agent that drives whatever optiq serve is serving, offline, engineered for local models (never an empty patch, edit-resilient, stall-proof). Product.

Quickstart

Every mlx-optiq quant loads with stock mlx-lm:

from mlx_lm import load, generate
model, tok = load("mlx-community/Qwen3.5-9B-OptiQ-4bit")
print(generate(model, tok, prompt="Hello", max_tokens=50))

Installing mlx-optiq unlocks the rest. A few starting points:

# Serve with the OpenAI + Anthropic API and ~1.4x speculative decode
optiq serve --model mlx-community/Qwen3.5-9B-OptiQ-4bit --mtp

# Run a huge MoE that doesn't fit in RAM (experts stream off SSD)
optiq serve --model mlx-community/Qwen3.5-122B-A10B-OptiQ-2bit --stream-experts

# Quantize a fresh model (exact sensitivity, or fast structural rules for big bases)
optiq convert Qwen/Qwen3.5-9B --target-bpw 5.0 --candidate-bits 4,8
optiq convert <large-moe> --method static --candidate-bits 2,4 --target-bpw 2.5

# Fine-tune with sensitivity-aware LoRA
optiq lora train mlx-community/Qwen3.5-9B-OptiQ-4bit --data ./jsonl_dir --rank 8

# Shard a model across two Macs over Thunderbolt (one OpenAI endpoint)
optiq cluster up                                    # on every Mac
optiq cluster serve --model mlx-community/Qwen3.5-122B-A10B-OptiQ-2bit

# Code with your local model, in your terminal (offline)
optiq code                                          # interactive, in a repo
optiq code -p "Fix the failing test in parser.py"   # headless

Full guides for serving, KV-quant, LoRA, MTP and per-family setup are in the docs. The models page lists every quant with its Capability Score, and the blog has the deeper write-ups.

Requirements

  • Apple Silicon (M1 or newer), macOS, Python 3.11+.
  • The published quants load with stock mlx-lm. Converting and some MoE / multimodal runtime features track mlx-lm main; install it from git when a model card asks for it.

License

MIT for the package. Quantized models follow their base model's license.

Release files for mlx-optiq 0.4.25

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mlx-optiq 0.4.25
File Size Uploaded
mlx_optiq-0.4.25.tar.gz 2.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for mlx-optiq 0.4.25
File Interpreter ABI Platform
mlx_optiq-0.4.25-py3-none-any.whl Python 3 none any Details

Total release size: 4.3 MB

Release files / mlx_optiq-0.4.25.tar.gz

Download URL mlx_optiq-0.4.25.tar.gz
Size 2.2 MB
Tags Source
SHA-256 checksum
How to use checksums
c9f69d988b4350b7c9f8f93ac66d9ecc197b12cf176174e9efeb562ff5531384
BLAKE2b-256 checksum
How to use checksums
a25d4ab0f13c9412834e5a007f47eaa80e407d41a470eb7c9b53b7de5b5fb795
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.7

Release files / mlx_optiq-0.4.25-py3-none-any.whl

Download URL mlx_optiq-0.4.25-py3-none-any.whl
Size 2.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
707c54e475719c83ec1a9b50ab1d652bee049c9cbf493fb9d307736d13ff73d2
BLAKE2b-256 checksum
How to use checksums
262dcaf7ea026b408f6db9fccf58c4f6cbf96d7be4a7afcdbc3964221307f4aa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.7

Release history Release notifications | RSS feed

0.5.13

2 release files

0.5.12

2 release files

0.5.11

2 release files

0.5.10

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.34

2 release files

0.4.33

2 release files

0.4.32

2 release files

0.4.31

2 release files

0.4.30

2 release files

0.4.29

2 release files

0.4.28

2 release files

0.4.27

2 release files

0.4.26

2 release files

This release

0.4.25 This release

2 release files

0.4.24

2 release files

0.4.23

2 release files

0.4.22

2 release files

0.4.21

2 release files

0.4.20

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.10

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page