Skip to main content

GGUF to MLX Converter for Apple Silicon

Convert supported GGUF language models into MLX-LM-compatible safetensors on Apple Silicon Macs.

Python 3.10+ License: Apache-2.0 Platform Validation

GGUF → dequantized safetensors → optional 4-bit MLX, with strict architecture validation.

Quick start · Supported models · Benchmarks · FAQ


What is gguf2mlx?

GGUF is great for distribution, but MLX and MLX-LM expect a Hugging Face-style directory with config.json, tokenizer assets, and safetensors weights.

gguf2mlx is a command-line converter for Mac users who have a model in GGUF format but need an MLX-LM model directory. It bridges that gap for supported architectures by:

  • reading GGUF metadata and tensors
  • rebuilding MLX-LM-compatible model artifacts
  • failing closed when a model layout is not actually supported

Use it when a model release provides GGUF files but no MLX checkpoint, when you want to load that model through mlx_lm.load(), or when you need an inspectable Hugging Face-style model directory on an M1, M2, M3, or M4 Mac.

If the original Hugging Face weights are available, converting those directly with MLX-LM is preferable because it avoids inheriting quantization error from an already-quantized GGUF file.

Important: gguf2mlx writes HF-style safetensors for MLX-LM and can optionally run mlx_lm.convert to emit 4-bit MLX output. This is MLX-LM re-quantization after conversion, not bit-for-bit preservation of the original GGUF quantization blocks.


Quick start

Install

# Base converter
pip install "gguf2mlx @ git+https://github.com/barrontang/gguf2mlx.git"

# Converter + MLX runtime for loading converted models
pip install "gguf2mlx[mlx] @ git+https://github.com/barrontang/gguf2mlx.git"

# Or with uv
uv add "gguf2mlx[mlx] @ git+https://github.com/barrontang/gguf2mlx.git"

The mlx extra installs both mlx and mlx-lm.

Install from a local checkout

If you cloned the repo and want the gguf2mlx command in your current environment, install it first:

python -m pip install -e .
# or with MLX / MLX-LM support
python -m pip install -e ".[mlx]"

Convert

# Basic conversion
gguf2mlx --input model.gguf --output ./mlx-model

# GGUF -> 4-bit MLX in one command
gguf2mlx --input model-Q4.gguf --output ./mlx-model-4bit --quantize --q-bits 4 --q-group-size 64

# Bounded-memory direct quantization (llama, gemma, mistral, qwen2, stablelm)
gguf2mlx --input model-Q4.gguf --output ./mlx-model-4bit-direct --quantize --direct-quant --q-bits 4 --q-group-size 64 --q-mode affine

# Float32 output
gguf2mlx --input model.gguf --output ./mlx-model-f32 --dtype float32

# Inspect metadata without writing weights
gguf2mlx --input model.gguf --skip-weights

# New command form (equivalent conversion subcommand)
gguf2mlx convert --input model.gguf --output ./mlx-model

Package an MLX model directory into .mlx

gguf2mlx can now package an MLX model directory into a single .mlx bundle, auto-generate a manifest list, and include SHA-256 integrity metadata.

# Package model directory into model.mlx (+ model.mlx.sha256)
gguf2mlx package --model-dir ./mlx-model --output ./mlx-model.mlx

# Verify embedded manifest and SHA-256 integrity
gguf2mlx verify --bundle ./mlx-model.mlx

Load with MLX-LM

python -c "
from mlx_lm import load, generate
model, tok = load('./mlx-model')
print(generate(model, tok, prompt='Hello from MLX', max_tokens=32))
"

Convert the MLX output to 4-bit

gguf2mlx can optionally run mlx_lm.convert for you. The one-command form is:

gguf2mlx \
  --input model-Q4.gguf \
  --output ./mlx-model-4bit \
  --quantize \
  --q-bits 4 \
  --q-group-size 64

If you want the manual two-step version, it is:

# 1) GGUF -> FP16 MLX-LM-style safetensors
gguf2mlx --input model-Q4.gguf --output ./mlx-model

# 2) FP16 safetensors -> 4-bit MLX
mlx_lm.convert \
  --model ./mlx-model \
  --mlx-path ./mlx-model-4bit \
  -q \
  --q-bits 4 \
  --q-group-size 64

Then load the quantized output normally:

python -c "
from mlx_lm import load
model, tok = load('./mlx-model-4bit')
print('loaded')
"

What works today

Capability Status
Strict architecture validation Supported
Atomic staging and cleanup on failure Supported
Correct torch_dtype propagation Supported
Correct zero-valued special token IDs Supported
GGUF context length -> model_max_length Supported
MLX / MLX-LM optional dependency Supported
One-command 4-bit output via mlx_lm.convert Supported
Package .mlx bundles + SHA-256 manifest integrity Supported
Opt-in mlx_lm.load() integration test Supported
Embedded tokenizer.huggingface.json preservation Supported
GGUF BOS/EOS/UNK and space-prefix metadata Supported
Executable Unigram and WordPiece tokenizer tests Supported
Gemma and Phi-3 architecture fixtures Supported
Bounded-memory direct quantization (--direct-quant) Supported
StableLM conversion adapter Supported

Supported conversion matrix

gguf2mlx distinguishes between:

  1. recognized for inspection via GGUF metadata, and
  2. conversion enabled through an explicit tensor adapter.

The current code recognizes 48 architecture identifiers for inspection and enables conversion for 12 identifiers.

Conversion-enabled architectures

Family GGUF architecture IDs Status
Llama llama, mistral Conversion enabled
Qwen qwen2, qwen2moe, qwen3moe Conversion enabled
DeepSeek deepseek2, deepseek3 Conversion enabled
GLM glm4moe Conversion enabled
Gemma gemma Fixture-validated adapter
Phi phi3 Phi-3 4K fixture-validated adapter
StableLM stablelm Conversion enabled
GLM glm-dsa Experimental conversion only

Gemma 2/3 and Phi-3 LongRoPE are different layouts and are not included in the basic Gemma or Phi-3 4K support claim.

Inspection only (not converted)

These may still be recognized by metadata or --skip-weights, but they are rejected during conversion until a dedicated, tested adapter exists:

arctic, baichuan, bert, bitnet, bloom, chameleon, chatglm, codeshell, command-r, command-r-plus, dbrx, exaone, falcon, gemma2, gemma3, gpt2, gptneox, granite, grok-1, jais, minicpm, minicpm3, mpt, nemotron, olmo, olmo2, openelm, orion, phi, phi2, plamo, refact, smolm, starcoder, t5, and xverse.

That means no more silent "Llama fallback" producing invalid outputs for unrelated architectures.


How conversion works

  1. Read GGUF metadata and detect the architecture
  2. Validate that the architecture has a supported adapter
  3. Build config.json using source metadata and selected dtype
  4. Export tokenizer assets from GGUF metadata
  5. Dequantize GGUF tensors to FP16 or FP32
  6. Remap tensor names into the target Hugging Face / MLX-LM layout
  7. Optionally run mlx_lm.convert --quantize into a second staged directory
  8. Write the selected output directory atomically

If any required step fails, conversion fails and the staged output is cleaned up.


Current limitations

This project is intentionally more honest about scope now:

  • 4-bit MLX output uses MLX-LM re-quantization; source GGUF Q4 blocks are not preserved directly
  • Tokenizer fidelity still depends on available GGUF metadata; embedded Hugging Face tokenizer JSON is preserved when present
  • Architecture coverage is adapter-based, not "all GGUF models"
  • Performance claims depend on model, prompt, hardware, and MLX-LM version
  • Phi-3 LongRoPE is rejected until its factor tensors are represented safely in the generated MLX configuration

If you need guaranteed support for a new family, open an issue with the exact GGUF architecture and source model.


Reproducible benchmarks

The repository does not claim that MLX is universally faster than llama.cpp. Conversion and inference performance depend on model architecture, quantization, prompt length, hardware, thermals, and library versions.

Run the included benchmark harness to record conversion time, peak RSS, input size, output size, platform information, and installed package versions:

uv run benchmarks/benchmark_conversion.py \
  --input ./model-Q4_K_M.gguf \
  --output ./benchmark-model-mlx \
  --result-json ./benchmark-results/model.json

To benchmark the complete GGUF-to-4-bit-MLX path:

uv run --extra mlx benchmarks/benchmark_conversion.py \
  --input ./model-Q4_K_M.gguf \
  --output ./benchmark-model-mlx-4bit \
  --quantize \
  --result-json ./benchmark-results/model-4bit.json

Plain conversion output is dequantized and can be substantially larger than the original GGUF quantized file. Use --quantize --q-bits 4 --q-group-size 64 when you want compact MLX-LM quantized output.

To attach a quality guardrail, add --eval-ppl: after conversion the harness loads each output with mlx_lm and measures perplexity on a fixed corpus, so you can see how much accuracy the GGUF-to-MLX round trip costs relative to a baseline. Perplexity failures never abort the benchmark; they are recorded as success: false in the JSON.

uv run --extra mlx benchmarks/benchmark_conversion.py \
  --input ./model-Q4_K_M.gguf \
  --output ./benchmark-model-mlx-4bit \
  --quantize --compare --eval-ppl \
  --ppl-dataset wikitext --ppl-num-samples 32 --ppl-seq-len 512 \
  --result-json ./benchmark-results/model-4bit.json

Choosing the right tool

Starting point Goal Recommended tool
GGUF model Run the GGUF directly llama.cpp or a GGUF application
Original Hugging Face model Create an MLX model mlx_lm.convert
GGUF-only model release Create an MLX-LM directory gguf2mlx
Existing MLX model directory Create a portable archive gguf2mlx package

Frequently asked questions

How do I convert a GGUF model to MLX on a Mac?

Install gguf2mlx[mlx], then run gguf2mlx convert --input model.gguf --output model-mlx. The output directory can be passed to mlx_lm.load() when the GGUF architecture and variant appear in the supported matrix.

Does GGUF-to-MLX conversion restore FP16 model quality?

No. Dequantization expands the stored values into FP16 or FP32, but it cannot recover information removed when the source GGUF was quantized.

Does the 4-bit output preserve the original GGUF Q4 blocks?

No. The current implementation dequantizes the GGUF and then uses MLX-LM to perform a second quantization. Use --direct-quant for the bounded-memory path that skips the FP16 staging directory; see docs/direct-quant-transcoding.md.

Why is the converted model larger than the GGUF file?

A quantized GGUF stores only a few bits per weight plus block metadata. Plain conversion writes FP16 or FP32 safetensors, so a larger output is expected.

Are Gemma and Phi supported?

The base gemma architecture and standard Phi-3 4K layout have fixture-backed adapters. Gemma 2, Gemma 3, Phi-2, Phi-MoE, and Phi-3 LongRoPE remain unsupported.

Is MLX always faster than llama.cpp?

No universal multiplier is claimed. Use the same model quality, prompt, context, sampling settings, and hardware when comparing runtimes.


Development

git clone https://github.com/barrontang/gguf2mlx.git
cd gguf2mlx
uv sync --all-extras

# Tests
pytest

# Opt-in MLX-LM integration test (Apple Silicon)
GGUF2MLX_RUN_E2E=1 pytest tests/test_e2e.py

# Lint
ruff check src/ tests/ benchmarks/

Hybrid Rust migration (in progress)

The repository now includes an initial Rust core scaffold at:

  • rust/gguf2mlx-rs

To build and enable the optional PyO3 extension locally:

python -m pip install maturin
maturin develop --manifest-path rust/gguf2mlx-rs/Cargo.toml

Current integration behavior:

  • Python CLI/UX remains the primary entrypoint.
  • Python can use an optional gguf2mlx_rust extension for architecture detection.
  • If the Rust extension is unavailable, the existing Python logic is used unchanged.

Recent regression coverage includes:

  • strict rejection of unsupported architectures
  • Gemma tensor mapping and GGUF norm-weight restoration fixtures
  • Phi-3 fused QKV and gated-MLP fixtures
  • preservation of zero-valued token IDs
  • preservation of embedded Hugging Face tokenizer JSON
  • executable tokenizer encode/decode checks
  • correct dtype propagation into config
  • atomic staging cleanup on failed writes
  • MLX-LM quantization error handling
  • opt-in mlx_lm.load() validation of quantized output

Roadmap status

Completed:

  • MLX quantized output through bundled mlx_lm.convert
  • MLX-LM load/integration tests on Apple Silicon
  • strict adapter-based architecture validation
  • fixture-backed Gemma and standard Phi-3 4K adapters
  • GGUF tokenizer flags and embedded tokenizer JSON preservation
  • reproducible conversion benchmark harness
  • bounded-memory direct quantization pipeline (--direct-quant, affine 4-bit, llama/gemma/mistral/qwen2/stablelm, Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q2_K–Q8_K/BF16 source qtypes)
  • StableLM conversion adapter (norm_eps, partial_rotary_factor, qk_layernorm, use_parallel_residual)
  • mlx_lm.load() parity validation test for --direct-quant output (tests/test_e2e.py)
  • opt-in real-GGUF convert + mlx_lm.load() finite-logit validation for gemma and phi3 fixtures (tests/test_e2e.py)
  • --eval-ppl perplexity guardrail in the benchmark harness (reuses mlx_lm.perplexity; reports standard vs --direct-quant PPL delta)

Remaining areas for contributors:

  • publish fixed-corpus perplexity numbers from --eval-ppl (standard vs --direct-quant vs mlx_lm.convert) so users can judge round-trip accuracy loss
  • sensitivity-aware bit retention in --direct-quant: keep the source per-tensor precision from K-quant inputs (Q6_K/Q8_0 layers stay higher-bit) instead of flattening every layer to 4-bit — an information advantage mlx_lm.convert cannot replicate from GGUF
  • extend the opt-in real-GGUF load test to the remaining conversion-enabled architectures (several are declared supported but have no end-to-end coverage; the safe_open bug showed how those paths silently rot)
  • canonical chat-template fallback table for known model families when GGUF-embedded Jinja templates are missing or broken, plus post-conversion template render validation
  • publish peak RSS and output-size comparison results from benchmarks/benchmark_conversion.py --compare
  • evaluate GGUF→MLX conversion for high-demand non-text models (e.g. ASR/embedding) where competition is thin (pending verification that such models ship GGUF)
  • add Gemma 2/3 and Phi LongRoPE adapters without broad family fallbacks
  • broader tokenizer fixture coverage for architecture-specific normalizers, byte fallback variants, and added-token edge cases

Contributing

PRs are welcome, especially for:

  • new architecture adapters backed by tensor manifests and load tests
  • tokenizer fidelity improvements
  • end-to-end mlx_lm.load() parity validation and perplexity benchmarks for --direct-quant
  • expanding --direct-quant to additional architectures and source quantization types
  • Apple Silicon integration coverage

When requesting a new model family, include the exact GGUF architecture, model name, quantization type, tokenizer type, and a public fixture or model URL. If the project saves you conversion work, starring the repository helps other Mac users discover it.


License

Apache-2.0 © Barron Tang


If this repo saved you time, please star it.

Release files for gguf2mlx 2.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gguf2mlx 2.1.1
File Size Uploaded
gguf2mlx-2.1.1.tar.gz 55.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gguf2mlx 2.1.1
File Interpreter ABI Platform
gguf2mlx-2.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 90.1 kB

Release files / gguf2mlx-2.1.1.tar.gz

Download URL gguf2mlx-2.1.1.tar.gz
Size 55.4 kB
Tags Source
SHA-256 checksum
How to use checksums
eb1617ebe083427afc42fd5eaa0c513073da7184acd53c35e78fb9943c9bb6d9
BLAKE2b-256 checksum
How to use checksums
4b18055d50ef5a6090d5ef521bb94999523a17e17dc35958132bb73b6c82dcd5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.9

Release files / gguf2mlx-2.1.1-py3-none-any.whl

Download URL gguf2mlx-2.1.1-py3-none-any.whl
Size 34.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ae1c5950667f00d9433231439b37373a55909353678998800d38651b9f3ffc4
BLAKE2b-256 checksum
How to use checksums
879c8b1619d131cef07bc552d6eadd8dadba00fb9320d5b549d8937f1aecbe33
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.9

Release history Release notifications | RSS feed

This release

2.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page