Last released May 4, 2026
Continuous batching & thermal-aware scheduling for LLM inference on edge devices (llama.cpp)
Supported by