Last released Aug 27, 2026
Hardware-aware auto-tuner that finds the smallest llama.cpp GGUF quantization + runtime config meeting a target throughput and quality-loss budget.