Last released Jul 4, 2026
TurboQuant data-oblivious rotation quantization for MLX-LM, with custom Metal kernels.
Last released Jun 1, 2026
Chunk-level KV cache reuse for faster HuggingFace inference
Supported by