Last released Aug 19, 2026
Local MoE-offload LLM inference runtime with OpenAI- and Anthropic-compatible APIs
Last released Jul 25, 2026
High-performance ML primitives, applications, and informative cost API — Triton + CuteDSL kernels for NVIDIA GPUs.
Last released Jun 3, 2026
Fast batched K-Means clustering with Triton GPU kernels