Last released Jul 25, 2026
Reproducible KV-cache tiering benchmark suite and signed results for LLM inference storage (Qwen3-480B class on 8x AMD MI308X): throughput +29-40%, TTFT -26-32% vs local NVMe.
Supported by