GGUF to MLX Converter for Apple Silicon
Convert supported GGUF language models into MLX-LM-compatible safetensors on Apple Silicon Macs.
GGUF → dequantized safetensors → optional 4-bit MLX, with strict architecture validation.
What is gguf2mlx?
GGUF is great for distribution, but MLX and MLX-LM expect a Hugging Face-style
directory with config.json, tokenizer assets, and safetensors weights.
gguf2mlx is a command-line converter for Mac users who have a model in GGUF
format but need an MLX-LM model directory. It bridges that gap for supported
architectures by:
- reading GGUF metadata and tensors
- rebuilding MLX-LM-compatible model artifacts
- failing closed when a model layout is not actually supported
Use it when a model release provides GGUF files but no MLX checkpoint, when you
want to load that model through mlx_lm.load(), or when you need an inspectable
Hugging Face-style model directory on an M1, M2, M3, or M4 Mac.
If the original Hugging Face weights are available, converting those directly with MLX-LM is preferable because it avoids inheriting quantization error from an already-quantized GGUF file.
Important:
gguf2mlxwrites HF-style safetensors for MLX-LM and can optionally runmlx_lm.convertto emit 4-bit MLX output. This is MLX-LM re-quantization after conversion, not bit-for-bit preservation of the original GGUF quantization blocks.
Quick start
Install
# Base converter
pip install "gguf2mlx @ git+https://github.com/barrontang/gguf2mlx.git"
# Converter + MLX runtime for loading converted models
pip install "gguf2mlx[mlx] @ git+https://github.com/barrontang/gguf2mlx.git"
# Or with uv
uv add "gguf2mlx[mlx] @ git+https://github.com/barrontang/gguf2mlx.git"
The mlx extra installs both mlx and mlx-lm.
Install from a local checkout
If you cloned the repo and want the gguf2mlx command in your current
environment, install it first:
python -m pip install -e .
# or with MLX / MLX-LM support
python -m pip install -e ".[mlx]"
Convert
# Basic conversion
gguf2mlx --input model.gguf --output ./mlx-model
# GGUF -> 4-bit MLX in one command
gguf2mlx --input model-Q4.gguf --output ./mlx-model-4bit --quantize --q-bits 4 --q-group-size 64
# Bounded-memory direct quantization (llama, gemma, mistral, qwen2, stablelm)
gguf2mlx --input model-Q4.gguf --output ./mlx-model-4bit-direct --quantize --direct-quant --q-bits 4 --q-group-size 64 --q-mode affine
# Float32 output
gguf2mlx --input model.gguf --output ./mlx-model-f32 --dtype float32
# Inspect metadata without writing weights
gguf2mlx --input model.gguf --skip-weights
# New command form (equivalent conversion subcommand)
gguf2mlx convert --input model.gguf --output ./mlx-model
Package an MLX model directory into .mlx
gguf2mlx can now package an MLX model directory into a single .mlx bundle,
auto-generate a manifest list, and include SHA-256 integrity metadata.
# Package model directory into model.mlx (+ model.mlx.sha256)
gguf2mlx package --model-dir ./mlx-model --output ./mlx-model.mlx
# Verify embedded manifest and SHA-256 integrity
gguf2mlx verify --bundle ./mlx-model.mlx
Load with MLX-LM
python -c "
from mlx_lm import load, generate
model, tok = load('./mlx-model')
print(generate(model, tok, prompt='Hello from MLX', max_tokens=32))
"
Convert the MLX output to 4-bit
gguf2mlx can optionally run mlx_lm.convert for you. The one-command form is:
gguf2mlx \
--input model-Q4.gguf \
--output ./mlx-model-4bit \
--quantize \
--q-bits 4 \
--q-group-size 64
If you want the manual two-step version, it is:
# 1) GGUF -> FP16 MLX-LM-style safetensors
gguf2mlx --input model-Q4.gguf --output ./mlx-model
# 2) FP16 safetensors -> 4-bit MLX
mlx_lm.convert \
--model ./mlx-model \
--mlx-path ./mlx-model-4bit \
-q \
--q-bits 4 \
--q-group-size 64
Then load the quantized output normally:
python -c "
from mlx_lm import load
model, tok = load('./mlx-model-4bit')
print('loaded')
"
What works today
| Capability | Status |
|---|---|
| Strict architecture validation | Supported |
| Atomic staging and cleanup on failure | Supported |
Correct torch_dtype propagation |
Supported |
| Correct zero-valued special token IDs | Supported |
GGUF context length -> model_max_length |
Supported |
| MLX / MLX-LM optional dependency | Supported |
One-command 4-bit output via mlx_lm.convert |
Supported |
Package .mlx bundles + SHA-256 manifest integrity |
Supported |
Opt-in mlx_lm.load() integration test |
Supported |
Embedded tokenizer.huggingface.json preservation |
Supported |
| GGUF BOS/EOS/UNK and space-prefix metadata | Supported |
| Executable Unigram and WordPiece tokenizer tests | Supported |
| Gemma and Phi-3 architecture fixtures | Supported |
Bounded-memory direct quantization (--direct-quant) |
Supported |
| StableLM conversion adapter | Supported |
Supported conversion matrix
gguf2mlx distinguishes between:
- recognized for inspection via GGUF metadata, and
- conversion enabled through an explicit tensor adapter.
The current code recognizes 48 architecture identifiers for inspection and enables conversion for 12 identifiers.
Conversion-enabled architectures
| Family | GGUF architecture IDs | Status |
|---|---|---|
| Llama | llama, mistral |
Conversion enabled |
| Qwen | qwen2, qwen2moe, qwen3moe |
Conversion enabled |
| DeepSeek | deepseek2, deepseek3 |
Conversion enabled |
| GLM | glm4moe |
Conversion enabled |
| Gemma | gemma |
Fixture-validated adapter |
| Phi | phi3 |
Phi-3 4K fixture-validated adapter |
| StableLM | stablelm |
Conversion enabled |
| GLM | glm-dsa |
Experimental conversion only |
Gemma 2/3 and Phi-3 LongRoPE are different layouts and are not included in the basic Gemma or Phi-3 4K support claim.
Inspection only (not converted)
These may still be recognized by metadata or --skip-weights, but they are
rejected during conversion until a dedicated, tested adapter exists:
arctic, baichuan, bert, bitnet, bloom, chameleon, chatglm,
codeshell, command-r, command-r-plus, dbrx, exaone, falcon,
gemma2, gemma3, gpt2, gptneox, granite, grok-1, jais, minicpm,
minicpm3, mpt, nemotron, olmo, olmo2, openelm, orion, phi,
phi2, plamo, refact, smolm, starcoder, t5, and xverse.
That means no more silent "Llama fallback" producing invalid outputs for unrelated architectures.
How conversion works
- Read GGUF metadata and detect the architecture
- Validate that the architecture has a supported adapter
- Build
config.jsonusing source metadata and selected dtype - Export tokenizer assets from GGUF metadata
- Dequantize GGUF tensors to FP16 or FP32
- Remap tensor names into the target Hugging Face / MLX-LM layout
- Optionally run
mlx_lm.convert --quantizeinto a second staged directory - Write the selected output directory atomically
If any required step fails, conversion fails and the staged output is cleaned up.
Current limitations
This project is intentionally more honest about scope now:
- 4-bit MLX output uses MLX-LM re-quantization; source GGUF Q4 blocks are not preserved directly
- Tokenizer fidelity still depends on available GGUF metadata; embedded Hugging Face tokenizer JSON is preserved when present
- Architecture coverage is adapter-based, not "all GGUF models"
- Performance claims depend on model, prompt, hardware, and MLX-LM version
- Phi-3 LongRoPE is rejected until its factor tensors are represented safely in the generated MLX configuration
If you need guaranteed support for a new family, open an issue with the exact GGUF architecture and source model.
Reproducible benchmarks
The repository does not claim that MLX is universally faster than llama.cpp. Conversion and inference performance depend on model architecture, quantization, prompt length, hardware, thermals, and library versions.
Run the included benchmark harness to record conversion time, peak RSS, input size, output size, platform information, and installed package versions:
uv run benchmarks/benchmark_conversion.py \
--input ./model-Q4_K_M.gguf \
--output ./benchmark-model-mlx \
--result-json ./benchmark-results/model.json
To benchmark the complete GGUF-to-4-bit-MLX path:
uv run --extra mlx benchmarks/benchmark_conversion.py \
--input ./model-Q4_K_M.gguf \
--output ./benchmark-model-mlx-4bit \
--quantize \
--result-json ./benchmark-results/model-4bit.json
Plain conversion output is dequantized and can be substantially larger than the
original GGUF quantized file. Use --quantize --q-bits 4 --q-group-size 64 when
you want compact MLX-LM quantized output.
To attach a quality guardrail, add --eval-ppl: after conversion the harness
loads each output with mlx_lm and measures perplexity on a fixed corpus, so you
can see how much accuracy the GGUF-to-MLX round trip costs relative to a baseline.
Perplexity failures never abort the benchmark; they are recorded as
success: false in the JSON.
uv run --extra mlx benchmarks/benchmark_conversion.py \
--input ./model-Q4_K_M.gguf \
--output ./benchmark-model-mlx-4bit \
--quantize --compare --eval-ppl \
--ppl-dataset wikitext --ppl-num-samples 32 --ppl-seq-len 512 \
--result-json ./benchmark-results/model-4bit.json
Choosing the right tool
| Starting point | Goal | Recommended tool |
|---|---|---|
| GGUF model | Run the GGUF directly | llama.cpp or a GGUF application |
| Original Hugging Face model | Create an MLX model | mlx_lm.convert |
| GGUF-only model release | Create an MLX-LM directory | gguf2mlx |
| Existing MLX model directory | Create a portable archive | gguf2mlx package |
Frequently asked questions
How do I convert a GGUF model to MLX on a Mac?
Install gguf2mlx[mlx], then run gguf2mlx convert --input model.gguf --output model-mlx. The output directory can be passed to mlx_lm.load() when
the GGUF architecture and variant appear in the supported matrix.
Does GGUF-to-MLX conversion restore FP16 model quality?
No. Dequantization expands the stored values into FP16 or FP32, but it cannot recover information removed when the source GGUF was quantized.
Does the 4-bit output preserve the original GGUF Q4 blocks?
No. The current implementation dequantizes the GGUF and then uses MLX-LM to
perform a second quantization. Use --direct-quant for the bounded-memory path
that skips the FP16 staging directory; see docs/direct-quant-transcoding.md.
Why is the converted model larger than the GGUF file?
A quantized GGUF stores only a few bits per weight plus block metadata. Plain conversion writes FP16 or FP32 safetensors, so a larger output is expected.
Are Gemma and Phi supported?
The base gemma architecture and standard Phi-3 4K layout have fixture-backed
adapters. Gemma 2, Gemma 3, Phi-2, Phi-MoE, and Phi-3 LongRoPE remain unsupported.
Is MLX always faster than llama.cpp?
No universal multiplier is claimed. Use the same model quality, prompt, context, sampling settings, and hardware when comparing runtimes.
Development
git clone https://github.com/barrontang/gguf2mlx.git
cd gguf2mlx
uv sync --all-extras
# Tests
pytest
# Opt-in MLX-LM integration test (Apple Silicon)
GGUF2MLX_RUN_E2E=1 pytest tests/test_e2e.py
# Lint
ruff check src/ tests/ benchmarks/
Hybrid Rust migration (in progress)
The repository now includes an initial Rust core scaffold at:
rust/gguf2mlx-rs
To build and enable the optional PyO3 extension locally:
python -m pip install maturin
maturin develop --manifest-path rust/gguf2mlx-rs/Cargo.toml
Current integration behavior:
- Python CLI/UX remains the primary entrypoint.
- Python can use an optional
gguf2mlx_rustextension for architecture detection. - If the Rust extension is unavailable, the existing Python logic is used unchanged.
Recent regression coverage includes:
- strict rejection of unsupported architectures
- Gemma tensor mapping and GGUF norm-weight restoration fixtures
- Phi-3 fused QKV and gated-MLP fixtures
- preservation of zero-valued token IDs
- preservation of embedded Hugging Face tokenizer JSON
- executable tokenizer encode/decode checks
- correct dtype propagation into config
- atomic staging cleanup on failed writes
- MLX-LM quantization error handling
- opt-in
mlx_lm.load()validation of quantized output
Roadmap status
Completed:
- MLX quantized output through bundled
mlx_lm.convert - MLX-LM load/integration tests on Apple Silicon
- strict adapter-based architecture validation
- fixture-backed Gemma and standard Phi-3 4K adapters
- GGUF tokenizer flags and embedded tokenizer JSON preservation
- reproducible conversion benchmark harness
- bounded-memory direct quantization pipeline (
--direct-quant, affine 4-bit, llama/gemma/mistral/qwen2/stablelm, Q4_0/Q4_1/Q5_0/Q5_1/Q8_0/Q2_K–Q8_K/BF16 source qtypes) - StableLM conversion adapter (norm_eps, partial_rotary_factor, qk_layernorm, use_parallel_residual)
mlx_lm.load()parity validation test for--direct-quantoutput (tests/test_e2e.py)- opt-in real-GGUF convert +
mlx_lm.load()finite-logit validation for gemma and phi3 fixtures (tests/test_e2e.py) --eval-pplperplexity guardrail in the benchmark harness (reusesmlx_lm.perplexity; reports standard vs--direct-quantPPL delta)
Remaining areas for contributors:
- publish fixed-corpus perplexity numbers from
--eval-ppl(standard vs--direct-quantvsmlx_lm.convert) so users can judge round-trip accuracy loss - sensitivity-aware bit retention in
--direct-quant: keep the source per-tensor precision from K-quant inputs (Q6_K/Q8_0 layers stay higher-bit) instead of flattening every layer to 4-bit — an information advantagemlx_lm.convertcannot replicate from GGUF - extend the opt-in real-GGUF load test to the remaining conversion-enabled architectures (several are declared supported but have no end-to-end coverage; the
safe_openbug showed how those paths silently rot) - canonical chat-template fallback table for known model families when GGUF-embedded Jinja templates are missing or broken, plus post-conversion template render validation
- publish peak RSS and output-size comparison results from
benchmarks/benchmark_conversion.py --compare - evaluate GGUF→MLX conversion for high-demand non-text models (e.g. ASR/embedding) where competition is thin (pending verification that such models ship GGUF)
- add Gemma 2/3 and Phi LongRoPE adapters without broad family fallbacks
- broader tokenizer fixture coverage for architecture-specific normalizers, byte fallback variants, and added-token edge cases
Contributing
PRs are welcome, especially for:
- new architecture adapters backed by tensor manifests and load tests
- tokenizer fidelity improvements
- end-to-end
mlx_lm.load()parity validation and perplexity benchmarks for--direct-quant - expanding
--direct-quantto additional architectures and source quantization types - Apple Silicon integration coverage
When requesting a new model family, include the exact GGUF architecture, model name, quantization type, tokenizer type, and a public fixture or model URL. If the project saves you conversion work, starring the repository helps other Mac users discover it.
License
Apache-2.0 © Barron Tang
If this repo saved you time, please star it.
Release files for gguf2mlx 2.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gguf2mlx-2.1.1.tar.gz | 55.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gguf2mlx-2.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 90.1 kB
Release files / gguf2mlx-2.1.1.tar.gz
| Download URL | gguf2mlx-2.1.1.tar.gz |
|---|---|
| Size | 55.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
eb1617ebe083427afc42fd5eaa0c513073da7184acd53c35e78fb9943c9bb6d9
|
|
BLAKE2b-256 checksum How to use checksums |
4b18055d50ef5a6090d5ef521bb94999523a17e17dc35958132bb73b6c82dcd5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.9
|
Release files / gguf2mlx-2.1.1-py3-none-any.whl
| Download URL | gguf2mlx-2.1.1-py3-none-any.whl |
|---|---|
| Size | 34.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7ae1c5950667f00d9433231439b37373a55909353678998800d38651b9f3ffc4
|
|
BLAKE2b-256 checksum How to use checksums |
879c8b1619d131cef07bc552d6eadd8dadba00fb9320d5b549d8937f1aecbe33
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.9
|