3 projects
mlx-kquant
GGUF K-quant dequantize / quantized-matmul / gather-qmm / quantize ops for MLX, via custom Metal kernels.
gmlx
A local inference platform for Apple Silicon: run, chat with, serve, and fine-tune the GGUF ecosystem's quantized models natively on MLX, straight off the file.
mlx-kld
KL-divergence scoring of quantized students (MLX safetensors or GGUF) against a full-precision teacher, with a self-managing disk cache of teacher logits