Last released Apr 1, 2026
Near-optimal KV-cache compression for HuggingFace transformers using Lloyd-Max + QJL quantization
Supported by