Last released Jul 19, 2026
High-performance FlashAttention-2 and Flash Decoding (GQA + varlen) kernels implemented in Triton, TileLang, and CUDA