🏎️ Dyno — llama.cpp Auto-Tuner & Benchmark
Dyno is an open-source CLI that auto-tunes and benchmarks llama.cpp / ik_llama.cpp inference on NVIDIA, AMD (ROCm), and Apple Silicon GPUs, producing reproducible, shareable results.
pipx install llama-dyno
dyno detect
dyno tune ~/models/mistral-7b.Q4_K_M.gguf --quick
dyno bench ~/models/mistral-7b.Q4_K_M.gguf --ngl 99 --fa
dyno bench --lmstudio llama-3.2-3b # benchmark an LM Studio model
dyno report ~/models/mistral-7b.Q4_K_M.gguf
dyno submit ~/models/mistral-7b.Q4_K_M.gguf
30-Second Quickstart
# 1. Install
pipx install llama-dyno
# 2. Check your hardware
dyno detect
# 3. Auto-tune a model
dyno tune ~/Downloads/my-model.q4_k_m.gguf
# 4. Run the final benchmark
dyno bench ~/Downloads/my-model.q4_k_m.gguf --ngl 99 --fa
# 5. Generate a report
dyno report ~/Downloads/my-model.q4_k_m.gguf
Prerequisites
- Python 3.11+
- A GPU: NVIDIA (drivers + CUDA), AMD (ROCm, via rocm-smi), or Apple Silicon (M-series, uses Metal + unified memory)
- llama-bench binary from llama.cpp or ik_llama.cpp — built with the matching backend (CUDA or Metal)
Install llama.cpp:
brew install llama.cpp # macOS / Linux
# or build from source:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build --config Release
ik_llama.cpp (optional, for MoE models):
git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama.cpp && cmake -B build && cmake --build build --config Release
# Symlink the ik binaries
ln -sf $(pwd)/build/bin/ik_llama-bench ~/.local/bin/
Commands
dyno detect
Fingerprint hardware: GPU model, VRAM, driver/CUDA version, CPU, RAM, and detect which llama.cpp backend is installed (with commit hash).
dyno tune <model.gguf>
Find the fastest config for your GPU + model.
| Flag | Default | Description |
|---|---|---|
--quick |
(default) | ~10 trials, fast iteration |
--thorough |
~25 trials, best accuracy |
Search strategy:
- Validate model loads correctly
- Coarse sweep — ngl (layer offload), flash attention on/off, KV cache quant (f16 / q8_0 / q4_0)
- Hill-climb — batch size (128–4096), threads (auto / half / full cores)
- ik_llama.cpp extras — -fmoe, -rtr, -amb toggles (thorough only)
OOM configs are discarded gracefully with ngl backoff. Live progress table shows every trial.
ik_llama.cpp Support
Dyno automatically detects ik_llama.cpp
when ik_llama-bench is on your PATH. Extra flags tuned:
| Flag | Description | Best for |
|---|---|---|
-fmoe |
Fast MoE computation | MoE models (Mixtral, DeepSeek) |
-rtr |
Runtime tensor reorder | Memory-bandwidth-bound scenarios |
-amb |
Attention memory bound | Large context / many KV heads |
Dyno detects which flags your build supports and only searches those. MoE models get fmoe enabled by default; Dense models skip MoE flags entirely.
LM Studio Support
Dyno can benchmark models served by LM Studio through its
OpenAI-compatible API (http://localhost:1234/v1).
# Show LM Studio status in hardware detection
dyno detect
# Benchmark a loaded model
dyno bench --lmstudio llama-3.2-3b
# Tune and create a report
dyno tune --lmstudio llama-3.2-3b --json-out report.json
LM Studio does not expose engine-reported tok/s — throughput is measured client-side from the streamed SSE response (wall-clock). Results are approximate and may vary with system load.
| Flag | Default | Description |
|---|---|---|
--lmstudio |
Benchmark an LM Studio model | |
--host |
http://localhost:1234/v1 |
LM Studio API base URL |
dyno bench <model.gguf>
Run a specific config 3× and report median tokens/sec with variance.
| Flag | Default | Description |
|---|---|---|
--ngl |
99 | GPU layers to offload |
--fa/--no-fa |
true | Flash attention |
--ctk |
f16 | K cache quant (f16, q8_0, q4_0) |
--ctv |
f16 | V cache quant (f16, q8_0, q4_0) |
--batch |
512 | Batch size |
--ubatch |
512 | Micro batch size |
--threads |
0 (auto) | Thread count |
--runs |
3 | Number of benchmark runs |
--fmoe |
false | Fast MoE (ik_llama.cpp) |
--rtr |
false | Runtime reorder (ik_llama.cpp) |
--amb |
false | Attention memory bound (ik_llama.cpp) |
dyno report <model.gguf>
Generate a shareable report including:
- Full hardware fingerprint
- Model details (name, quant, SHA-256)
- Winning config
- Median scores with variance
- Reproducible llama-server command
- Shareable markdown snippet
- JSON output
dyno submit <model.gguf>
Submit your results to the community results repo (llama-dyno-results):
- Opens a PR via
ghCLI - Falls back to a GitHub Gist
- Saves locally if neither works
Quality & Reproducibility
- Fixed bench params: Every run uses
pp=512/tg=128(fixed prompt/gen tokens) so results are comparable - 3-run median + variance reported
- No fabricated numbers — real hardware detection, real subprocess results
- OOM handling — graceful backoff of ngl
- Clear install hints — suggests brew or git clone when binaries missing
Results Table (Example)
| GPU | Model | Quant | Backend | TG tok/s |
|---|---|---|---|---|
| RTX 4090 | Mixtral-8x7B | Q4_K_M | ik_llama.cpp | 58.7 |
| RTX 4090 | Llama-3-70B | Q4_K_M | ik_llama.cpp | 42.3 |
| RTX 4090 | Llama-3-70B | Q4_K_M | llama.cpp | 35.1 |
| RTX 3090 | Mistral-7B | Q4_K_M | llama.cpp | 112.8 |
| RTX 4060 | Phi-3-mini | Q4_K_M | llama.cpp | 68.5 |
Shell Completions
Dyno supports shell completions via Typer/Click:
# Install completions for your shell
dyno --install-completion
# Show completion script (to manually install)
dyno --show-completion
Supported shells: bash, zsh, fish, powershell.
Architecture
src/llama_dyno/
├── __init__.py # Package metadata
├── cli.py # Typer CLI (detect, tune, bench, report, submit)
├── types.py # Data types (BenchParams, TrialResult, etc.)
├── detect.py # Hardware fingerprinting (pynvml, nvidia-smi, /proc)
├── bench.py # llama-bench subprocess driver + parser
├── tune.py # Search strategy (coarse sweep + hill climb)
├── ollama.py # Ollama REST API runner
├── lmstudio.py # LM Studio OpenAI-compatible API runner
├── report.py # Report generation (JSON + markdown)
└── submit.py # GitHub PR / Gist submission
Development
git clone https://github.com/lachy/llama-dyno
cd llama-dyno
pip install -e ".[dev]"
pytest
Out of Scope (v1)
- GUI
- Multi-GPU
- Intel backends (planned)
- Server hosting for results
License
Apache 2.0
Metadata
Release files for llama-dyno 1.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_dyno-1.5.0.tar.gz | 49.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_dyno-1.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 90.8 kB
Release files / llama_dyno-1.5.0.tar.gz
| Download URL | llama_dyno-1.5.0.tar.gz |
|---|---|
| Size | 49.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
37eebe67f0ac88c463e695495f60d862fb8e4f8083b004acfc6583bca7f3860a
|
|
BLAKE2b-256 checksum How to use checksums |
8b85bcfa6109e13820d9c4f10bd356185f71d93b396b00e6f205ac10e1985116
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Release files / llama_dyno-1.5.0-py3-none-any.whl
| Download URL | llama_dyno-1.5.0-py3-none-any.whl |
|---|---|
| Size | 41.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e2a90a6d2bf00454a82a5fada192250c02b2334f54148e530d9d3ebf96af8c90
|
|
BLAKE2b-256 checksum How to use checksums |
d7eb39573aff95d9097c9b8c8c5989651c4c4ef84f03d7763dea1bf8f8d63e6d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|