Skip to main content
VeloxQuant-MLX

VeloxQuant-MLX

Fast KV Cache Quantization for Apple Silicon
TurboQuant · RVQ · VecInfer · RateQuant · PolarQuant · QJL · SpectralQuant · CommVQ · RaBitQ — in MLX

PyPI Downloads Python Platform License Tests DOI

Landing Changelog Blog Blog v2 Ko-fi Buy Me a Chai


VeloxQuant-MLX compresses the KV cache of any mlx_lm model on Apple Silicon — up to 16× smaller with near-lossless quality, in three lines of code. It ships 41 research-adapted compression methods, from zero-calibration 1-bit quantizers to token-eviction caches to cross-layer merging, plus hand-written Metal kernels that make the hottest path up to 14.7× faster.

Why VeloxQuant-MLX:

  • 41 methods behind one identical 3-line API — swap method="..." and go
  • Metal-accelerated hot paths: 6.9–14.7× faster quantize, 98% less peak memory at the OOM-trigger shape
  • Every "-adapted" method documents its honest deviation from the source paper — no silent approximations
  • Validated end-to-end on 12 production models: Llama, Mistral, Qwen, Phi, Gemma 3/4, Falcon
import mlx_lm
from veloxquant_mlx import KVCacheBuilder, KVCacheConfig

model, tokenizer = mlx_lm.load("mlx-community/Mistral-7B-Instruct-v0.3-4bit")
config = KVCacheConfig(method="turboquant_rvq", bit_width_inlier=1, seed=42)
caches = KVCacheBuilder.for_model(model, config)
model.make_cache = lambda *_a, **_k: caches

response = mlx_lm.generate(model, tokenizer, prompt="Explain relativity simply.", max_tokens=200)

Numbers that matter

Metric Value Notes
Max key cache compression 16× VecInfer-1bit, head_dim=128
Metal kernel speedup 13× quantize_vq at S=2048 (range 6.9–14.7× over S=128–8192)
Peak memory reduction 98% 729 MB → 12 MB, Falcon3-7B shape
RVQ-1bit compression 7.5× Near-zero throughput cost
FP16 throughput retained 100% Qwen2.5-7B at 16× compression
SpectralQuant compression 5.33× per-model measured (Qwen2.5-0.5B / Gemma-4-4B), same bit-width
SpectralQuant cosine sim +3pp over TurboQuant on Qwen2.5-0.5B
RaBitQ full KV compression 1-bit keys + MSE-b4 values, Falcon3-7B
RaBitQ context at 8 GB ~103k tokens (est.) KV-only linear extrapolation from measured memory rows; vs ~17k fp16 — 6× more context
CommVQ key compression 64× RoPE-commutative VQ, D=128, n_cb=4
KIVI-2bit key compression 5.8× per-channel keys / per-token values; measured on Llama-3.2-3B, Qwen2.5-7B, Mistral-7B
KIVI-2bit full-KV compression ~4× incl. fp16 residual window (32 tokens); 100–106% of fp16 throughput
Production models validated 12 Llama, Mistral, Qwen, Phi, Gemma 3/4, Falcon

Table of contents

  1. Installation
  2. Quickstart
  3. Method library — all 41 methods at a glance
  4. Metal kernels
  5. Benchmark results
  6. What's inside
  7. Architecture
  8. CLI
  9. Development
  10. Documentation & blog posts
  11. References
  12. Support

Installation

pip install VeloxQuant-MLX

Requirements: Apple Silicon M1+, Python ≥ 3.11, MLX ≥ 0.18, NumPy ≥ 1.26.

Install from source
git clone https://github.com/rajveer43/VeloxQuant-MLX
cd VeloxQuant-MLX
pip install -e ".[dev]"

Quickstart

RVQ 1-bit — 7.5× compression, no calibration (recommended default)

import mlx_lm
from veloxquant_mlx import KVCacheBuilder, KVCacheConfig

model, tokenizer = mlx_lm.load("mlx-community/Mistral-7B-Instruct-v0.3-4bit")

config = KVCacheConfig(method="turboquant_rvq", bit_width_inlier=1, seed=42)
caches = KVCacheBuilder.for_model(model, config)
model.make_cache = lambda *_a, **_k: caches

response = mlx_lm.generate(model, tokenizer,
    prompt="Explain the theory of relativity in simple terms.",
    max_tokens=200,
)

More examples, walked through step by step:


Method library

All 41 methods share the same 3-line integration (method="<id>" in KVCacheConfig). Each links to its full page — mechanism, config, evidence, and honest limitations — on the documentation site.

Quick decision:

  • No calibration, best default → turboquant_rvq b=1 (7.5×, 0.92 cosine)
  • Max compression, Qwen2.5/Gemma → vecinfer 1-bit (16×, Metal-accelerated)
  • Best quality at moderate compression → spectral b=3 (5.33×, ~5s calibration)
  • Heterogeneous layers (sensitivity ratio >2×) → RateQuant on top of RVQ
  • Max context length, fixed RAM → rabitq keys + MSE-b4 values (6× full KV)
  • RoPE-compatible exact VQ → comm_vq (ICML 2025, 64× key compression)

Quantization — compress every token

Method method= What it does Compression New in
TurboQuant RVQ turboquant_rvq Residual VQ, zero calibration — the default 7.5× @ 1-bit
VecInfer vecinfer Dual-transform product VQ, Metal-accelerated 16× 0.4.0
SpectralQuant spectral Rotate keys into eigenbasis — best quality-per-bit 5.33× 0.6.0
RateQuant (allocator) Per-layer mixed precision via reverse-waterfilling 5.2× @ 1.5 avg bit
RaBitQ rabitq 1-bit keys + MSE-b4 values 6× full KV 0.7.0
QJL qjl 1-bit JL sketch, simplest/fastest to set up ~16×
PolarQuant polar Polar-coordinate quant for geometric key distributions varies
CommVQ comm_vq RoPE-commutative VQ, exact inner product (ICML 2025) 64× keys
KIVI kivi Tuning-free asymmetric 2-bit baseline 5.8× 0.8.0
KIVI-Sink kivi_sink Sink-protected low-bit quantization ~5.8× 0.9.0
SKVQ-adapted skvq Channel reordering + clipped dynamic quant behind a sliding fp16 window + sink filter (COLM 2024) varies 0.30.0
SVDq svdq Sub-2-bit keys (~1.25 bit) via prefill SVD ~10× 0.10.0
Kitty kitty Adaptive channel precision, zero calibration varies 0.11.0
KVQuant-NUQ kvquant Non-uniform datatype + outlier isolation varies 0.14.0
NSNQuant-adapted nsnquant Calibration-free universal-codebook VQ — fixed Gaussian codebook (NeurIPS 2025) 1–2 bit/elem 0.28.0
ZipCache-adapted zipcache Per-token mixed bit-width by key-norm saliency varies 0.18.0
GEAR gear Error-feedback: low-rank + sparse residual correction varies 0.17.0
CacheGen cachegen Entropy-coded cache — storage win on correlated KV varies 0.16.0
AMC-adapted amc Saliency-driven tiered rank + precision — one L1-norm score drives both rank and bit-width per token, never evicts (no verified venue — second exception; hardware/RTL half of paper out of scope) varies 0.38.0
A2ATS-adapted a2ats Windowed RoPE + query-aware retrieval VQ — exact RoPE within a trailing window, shared approximate rotation outside it; query-aware codebook assignment for a retrieval-fraction subset (ACL 2025 Findings) ~4× 0.39.0

Low-rank & cross-layer — compress across dimensions or depth

Method method= What it does Compression New in
PALU palu True low-rank latent storage of both K and V varies 0.15.0
XQuant xquant Cross-layer code reuse — adjacent layers share codes varies 0.12.0
MiniCache minicache Cross-layer SLERP merge — deep layer pairs cost one ~2× on deep layers 0.16.0
xKV-adapted xkv Cross-layer shared-subspace SVD — one basis jointly fit across a layer group varies 0.27.0
AdaKV-proxy adakv Per-head adaptive bit budget, layered on KIVI varies 0.13.0
KVTC-adapted kvtc Local PCA + DP-optimal per-component bit allocation + entropy coding (ICLR 2026) — beats fixed-split mixed-precision at matched byte budget on skewed variance varies 0.35.0

Token eviction & merging — drop or merge low-value tokens

Method method= What it does New in
SnapKV-adapted snapkv Prefill observation-window eviction, once at prefill end 0.19.0
StreamingLLM-adapted streaming_llm Sink + recency window, constant memory 0.20.0
H2O-adapted h2o Cumulative attention-mass heavy-hitter eviction 0.21.0
TOVA-adapted tova Memoryless current-step attention-weight eviction 0.22.0
PyramidKV-adapted pyramidkv H2O eviction with a per-layer pyramid budget 0.23.0
SqueezeAttention-adapted squeeze 2D layer×token data-driven budget eviction 0.24.0
ChunkKV-adapted chunkkv Chunk-level eviction (chunk_size=1 == H2O) 0.25.0
CaM-adapted cam Cache merging — merge evicted tokens, don't drop (cam_merge=drop == H2O) 0.26.0
L2Norm-adapted knorm Intrinsic key-norm eviction — low norm ⇒ important (EMNLP 2024) 0.29.0
Q-Filters-adapted qfilters Query-agnostic projection eviction — frozen per-head key-SVD direction 0.31.0
Keyformer-adapted keyformer Gumbel-regularized heavy-hitter eviction (MLSys 2024); keyformer_tau=0 == H2O 0.32.0
MorphKV-adapted morphkv Recent-window correlation retention (ICML 2025); morphkv_window=1 == TOVA 0.33.0
KVzip-adapted kvzip Context-reconstruction reliance eviction (NeurIPS 2025); kvzip_probe=latest == TOVA 0.34.0
CurDKV-adapted curdkv Value-aware leverage-score eviction via approximated CUR decomposition (NeurIPS 2025) — evicts key-similar but value-irrelevant tokens that key-only eviction (H2O) cannot distinguish 0.36.0
NestedKV-adapted nestedkv Multi-scale ensembled prefill eviction — stable + episodic + current key anomaly, combined by a head-adaptive blend and surprise-gated route (no verified venue — one-time exception) 0.37.0

Every "-adapted" method is an honest adaptation, not a faithful port — the cache wrapper sees per-layer K/V but not the model's true query/attention maps, so attention-based signals use a key-as-query proxy. Each method's docs page states its specific limitations plainly.


Metal kernels — new in 0.5.1

The VecInfer quantize_vq hot path is now a 30-line Metal Shading Language shader, JIT-compiled by mx.fast.metal_kernel on first use. Same Python API — no changes required.

Metal kernel benchmark — quantize latency, speedup, and peak memory
Benchmarked on Apple Silicon GPU. Left: quantize latency. Center: speedup factor. Right: peak memory.

Metric Pure MLX Metal kernel Delta
Quantize latency (S=8192) 228 ms 15.6 ms 14.7× faster
Peak memory (Falcon3-7B shape) 729 MB 12 MB 98% reduction
API change required None use_metal_kernels=None auto-detects

Why the memory win: the [N, n_centroids, sub_dim] diff tensor is never materialised — the argmin accumulator lives entirely in thread-local registers.

Honest caveat: the kernel pays a ~50–200 µs launch overhead per call. On tiny models (SmolLM2-135M, ~60 launches/token) that overhead can exceed the savings. Built for the regime that needs it: 7B+ models at realistic context lengths.

Full kernel source and how it was built: blogs/metal-kernels.md. Usage, fallback behaviour, and debugging: docs — Metal GPU kernels.


Benchmark results

10-model comparative study — VecInfer vs RVQ (v0.5.0)

Cross-model comparison — VecInfer vs RVQ-1bit across 10 models
End-to-end mlx_lm.generate · 200-token prompt · 120-token generation · Apple M-series unified memory

Compression ratio:

Model RVQ-1bit VecInfer-1bit
SmolLM2-135M 7.1× 16×
Llama-3.2-1B 7.1× 16×
Llama-3.2-3B 7.5× 16×
Llama-3.1-8B 7.5× 16×
Mistral-7B 7.5× 16×
Qwen2.5-7B 7.5× 16×
Qwen3-8B 7.5× 16×
Phi-4 7.5× 16×
Falcon3-7B 7.8× 16×
gemma-3-4b 7.8× 16×

Throughput (tok/s):

Model fp16 RVQ-1bit VecInfer-1bit
SmolLM2-135M 250.4 188.5 175.8
Llama-3.2-1B 105.4 104.3 91.2
Llama-3.2-3B 47.6 46.2 40.2
Llama-3.1-8B 20.5 20.6 19.6
Mistral-7B 23.6 22.8 9.8
Qwen2.5-7B 21.0 20.7 21.5 ⬆ exceeds fp16 at 16×
Qwen3-8B 20.3 19.6 2.4
Phi-4 10.4 8.1 4.0
Falcon3-7B 17.3 21.7 17.0
gemma-3-4b 26.0 24.2 22.6

RVQ-1bit is the safe default — within 5% of fp16 on most 7–8B models with zero calibration. VecInfer-1bit wins on memory (always 16×) and throughput on strong-GQA models (Qwen2.5, Gemma).

Historical benchmark snapshots (throughput optimisation journey, RateQuant V2, 8-model RVQ sweep) and full methodology: BENCHMARK_RESULTS.md.


What's inside

Module Purpose
veloxquant_mlx/quantizers/turboquant_rvq Two-pass scalar RVQ — Gaussian + Laplacian codebooks, b=1/2/3+
veloxquant_mlx/cache/vecinfer_cache VecInferKVCache — smooth + Hadamard + product VQ
veloxquant_mlx/cache/turboquant_rvq_cache TurboQuantRVQKVCache — mlx_lm-compatible wrapper
veloxquant_mlx/allocators allocate_bits_ratequant, calibrate_layer_sensitivities, VecInfer calibration
veloxquant_mlx/metal Hand-written Metal MSL kernels, JIT via mx.fast.metal_kernel
veloxquant_mlx/spectral SpectralQuantizer, rotation calibration, water-filling bit allocation

Full module reference and API docs: docs — API reference.


Architecture

VeloxQuant-MLX pipelines each quantizer as rotate → quantize (± residual) → pack, built via a Builder/Factory/Strategy layering so every method shares the same KVCacheConfigKVCacheBuildermlx_lm-compatible cache path. Ten design patterns are used throughout (Abstract Base Classes, Factory, Chain of Responsibility, Builder, Strategy, Registry + Plugin, Composite, Observer, DAO, and custom data structures like RingBuffer/MaxHeap/BitPackBuffer/VoronoiTree).

Full pipeline diagrams (TurboQuantRVQ, VecInfer) and design-pattern breakdown: docs — Core concepts.


CLI

# Precompute rotation matrices, JL matrices, codebooks
python -m veloxquant_mlx precompute \
    --head_dim 128 --bits 1 2 3 4 --jl_dim 128 --seed 42 \
    --output_dir ./artifacts/

# Synthetic benchmark — single config
python -m veloxquant_mlx benchmark \
    --method turboquant_rvq --head_dim 128 --bits 2 --seq_len 1000

# End-to-end model benchmarks
python benchmark_scripts/benchmark_vecinfer.py   # VecInfer 10-model sweep
python benchmark_scripts/run_outlier_ratequant.py # RateQuant mixed-precision

Load precomputed artifacts to skip re-computation at runtime:

from veloxquant_mlx.artifacts import NpyArtifactStore

cache = (KVCacheBuilder()
    .with_method("turboquant_rvq")
    .with_head_dim(128).with_bit_width(inlier=2)
    .with_artifact_store(NpyArtifactStore("./artifacts/"))
    .build())

Development

# Full test suite (includes Metal parity tests)
pytest veloxquant_mlx/tests/ -v

# 2-bit improvement validation — fast synthetic run
python test_2bit_improvements.py

# Generate optimization-journey figure
python scripts/plot_optimization_journey.py

Contributions welcome — please open an issue first for anything beyond a small bugfix. See CONTRIBUTING.md for guidelines and CHANGELOG.md for release history.


Documentation & blog posts

Full docs, including per-method pages, guides, and API reference: https://veloxquant-mlx.netlify.app/

Deep-dive writeups live in blogs/ and are also published on the docs site:

File Description Live
blogs/overview.md High-level overview of VeloxQuant-MLX and its goals
blogs/10-model-study.md End-to-end benchmark study across 10 production models
blogs/hands-on.md Hands-on tutorial: compressing your first model
blogs/kivi.md Deep dive into the KIVI asymmetric quantization baseline
blogs/metal-kernels.md How the Metal compute kernel cuts quantize latency 13×
blogs/results.md Detailed benchmark results and analysis
blogs/tensorops-research.md TensorOps research notes and findings
blogs/turboquant-metal-kernels.md TurboQuant + Metal kernels: combined writeup

References

41 methods, each adapted from a published paper with documented deviations (39 from a verified peer-reviewed venue; 2, NestedKV-adapted and AMC-adapted, from unpublished preprints as one-time, stated exceptions — see CITATIONS.md) — full bibliography (implemented methods, related work, and survey papers): CITATIONS.md.

Headline references: TurboQuant (ICLR 2026), VecInfer (2024), RaBitQ (SIGMOD 2024), CommVQ (ICML 2025), KVzip (NeurIPS 2025), KVTC (ICLR 2026), CurDKV (NeurIPS 2025), NestedKV (preprint, arXiv:2605.26678), AMC (preprint, arXiv:2607.10109), A2ATS (ACL 2025 Findings). Built on Apple MLX.


Support

VeloxQuant-MLX has passed 15,000+ downloads on PyPI. It's free, MIT-licensed, and built nights-and-weekends — if it saves your Mac some memory (or you just want to see the 42nd method land), you can buy me a chai or tip on Ko-fi 💜. Stars, issues, and PRs are equally appreciated.


License

MIT — see LICENSE.


Built for Apple Silicon · Engineered for speed · MIT License
Landing page · Issues · Blog: 10-model study · Blog: Metal kernels v1 · Blog: TurboQuant Metal kernels

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

veloxquant_mlx-0.39.0-py3-none-any.whl (681.1 kB view details)

Uploaded Python 3

File details

Details for the file veloxquant_mlx-0.39.0-py3-none-any.whl.

File metadata

File hashes

Hashes for veloxquant_mlx-0.39.0-py3-none-any.whl
Algorithm Hash digest
SHA256 365766966f7f2b9b68ed2bfca45ef298e42aa3f9ba887d0be47c02ce57b5a030
MD5 24e21e98788b8fd925e38940ebbfa8d4
BLAKE2b-256 ffdc9233876d27ae6de3fd1d811d682af580393df5483b13c39746e4402da00a

See more details on using hashes here.

Release history Release notifications | RSS feed

0.59.0

2 files

0.58.0

2 files

0.57.1

2 files

0.57.0

2 files

0.56.0

2 files

0.55.0

2 files

0.54.1

2 files

0.54.0

2 files

0.53.0

2 files

0.52.1

2 files

0.52.0

2 files

0.51.1

2 files

0.51.0

2 files

0.50.2

2 files

0.50.1

2 files

0.49.4

2 files

0.49.3

2 files

0.49.2

2 files

0.49.1

2 files

0.49.0

2 files

0.48.5

2 files

0.48.4

2 files

0.48.3

2 files

0.48.2

2 files

0.48.1

2 files

0.48.0

2 files

0.47.1

2 files

0.47.0

2 files

0.46.0

2 files

0.45.0

2 files

0.44.4

2 files

0.44.3

2 files

0.44.2

2 files

0.44.1

2 files

0.44.0

2 files

0.42.0

2 files

0.41.0

2 files

0.40.0

2 files

0.39.1

2 files

This release

0.39.0 This release

1 file

0.38.0

2 files

0.37.0

2 files

0.36.0

2 files

0.35.0

2 files

0.34.0

2 files

0.33.0

2 files

0.32.0

2 files

0.31.0

2 files

0.30.1

2 files

0.30.0

2 files

0.29.0

2 files

0.28.0

2 files

0.27.0

2 files

0.26.0

2 files

0.25.0

2 files

0.24.1

2 files

0.24.0

2 files

0.23.1

2 files

0.23.0

2 files

0.22.0

2 files

0.21.0

2 files

0.20.0

2 files

0.19.0

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

1 file

0.3.6

1 file

0.3.5

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page