Last released Sep 10, 2026
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.
Last released Sep 7, 2026
Python client for the ARBI API