Last released Sep 28, 2026
A LLM serving engine extension to reduce TTFT and increase throughput, especially under long-context scenarios.