Last released Jul 24, 2026
Benchmark the LLM inference capacity of a server (llama-benchy orchestrator with auto-detection, real-workload sizing, charts and reports).
Supported by