Last released May 8, 2026
Adaptive Memory Runtime for LLMs — compress KV cache 3-7x with vectorized GPU ops, fused compressed attention, and rigorous benchmark suite
Supported by