Skip to main content

🏎️ Dyno — llama.cpp Auto-Tuner & Benchmark

Dyno is an open-source CLI that auto-tunes and benchmarks llama.cpp / ik_llama.cpp inference on NVIDIA, AMD (ROCm), and Apple Silicon GPUs, producing reproducible, shareable results.

pipx install llama-dyno

dyno detect
dyno tune ~/models/mistral-7b.Q4_K_M.gguf --quick
dyno bench ~/models/mistral-7b.Q4_K_M.gguf --ngl 99 --fa
dyno bench --lmstudio llama-3.2-3b  # benchmark an LM Studio model
dyno report ~/models/mistral-7b.Q4_K_M.gguf
dyno submit ~/models/mistral-7b.Q4_K_M.gguf

30-Second Quickstart

# 1. Install
pipx install llama-dyno

# 2. Check your hardware
dyno detect

# 3. Auto-tune a model
dyno tune ~/Downloads/my-model.q4_k_m.gguf

# 4. Run the final benchmark
dyno bench ~/Downloads/my-model.q4_k_m.gguf --ngl 99 --fa

# 5. Generate a report
dyno report ~/Downloads/my-model.q4_k_m.gguf

Prerequisites

  • Python 3.11+
  • A GPU: NVIDIA (drivers + CUDA), AMD (ROCm, via rocm-smi), or Apple Silicon (M-series, uses Metal + unified memory)
  • llama-bench binary from llama.cpp or ik_llama.cpp — built with the matching backend (CUDA or Metal)

Install llama.cpp:

brew install llama.cpp                        # macOS / Linux
# or build from source:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build --config Release

ik_llama.cpp (optional, for MoE models):

git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama.cpp && cmake -B build && cmake --build build --config Release
# Symlink the ik binaries
ln -sf $(pwd)/build/bin/ik_llama-bench ~/.local/bin/

Commands

dyno detect

Fingerprint hardware: GPU model, VRAM, driver/CUDA version, CPU, RAM, and detect which llama.cpp backend is installed (with commit hash).

dyno tune <model.gguf>

Find the fastest config for your GPU + model.

Flag Default Description
--quick (default) ~10 trials, fast iteration
--thorough ~25 trials, best accuracy

Search strategy:

  1. Validate model loads correctly
  2. Coarse sweep — ngl (layer offload), flash attention on/off, KV cache quant (f16 / q8_0 / q4_0)
  3. Hill-climb — batch size (128–4096), threads (auto / half / full cores)
  4. ik_llama.cpp extras — -fmoe, -rtr, -amb toggles (thorough only)

OOM configs are discarded gracefully with ngl backoff. Live progress table shows every trial.

ik_llama.cpp Support

Dyno automatically detects ik_llama.cpp when ik_llama-bench is on your PATH. Extra flags tuned:

Flag Description Best for
-fmoe Fast MoE computation MoE models (Mixtral, DeepSeek)
-rtr Runtime tensor reorder Memory-bandwidth-bound scenarios
-amb Attention memory bound Large context / many KV heads

Dyno detects which flags your build supports and only searches those. MoE models get fmoe enabled by default; Dense models skip MoE flags entirely.

LM Studio Support

Dyno can benchmark models served by LM Studio through its OpenAI-compatible API (http://localhost:1234/v1).

# Show LM Studio status in hardware detection
dyno detect

# Benchmark a loaded model
dyno bench --lmstudio llama-3.2-3b

# Tune and create a report
dyno tune --lmstudio llama-3.2-3b --json-out report.json

LM Studio does not expose engine-reported tok/s — throughput is measured client-side from the streamed SSE response (wall-clock). Results are approximate and may vary with system load.

Flag Default Description
--lmstudio Benchmark an LM Studio model
--host http://localhost:1234/v1 LM Studio API base URL

dyno bench <model.gguf>

Run a specific config 3× and report median tokens/sec with variance.

Flag Default Description
--ngl 99 GPU layers to offload
--fa/--no-fa true Flash attention
--ctk f16 K cache quant (f16, q8_0, q4_0)
--ctv f16 V cache quant (f16, q8_0, q4_0)
--batch 512 Batch size
--ubatch 512 Micro batch size
--threads 0 (auto) Thread count
--runs 3 Number of benchmark runs
--fmoe false Fast MoE (ik_llama.cpp)
--rtr false Runtime reorder (ik_llama.cpp)
--amb false Attention memory bound (ik_llama.cpp)

dyno report <model.gguf>

Generate a shareable report including:

  • Full hardware fingerprint
  • Model details (name, quant, SHA-256)
  • Winning config
  • Median scores with variance
  • Reproducible llama-server command
  • Shareable markdown snippet
  • JSON output

dyno submit <model.gguf>

Submit your results to the community results repo (llama-dyno-results):

  1. Opens a PR via gh CLI
  2. Falls back to a GitHub Gist
  3. Saves locally if neither works

Quality & Reproducibility

  • Fixed bench params: Every run uses pp=512 / tg=128 (fixed prompt/gen tokens) so results are comparable
  • 3-run median + variance reported
  • No fabricated numbers — real hardware detection, real subprocess results
  • OOM handling — graceful backoff of ngl
  • Clear install hints — suggests brew or git clone when binaries missing

Results Table (Example)

GPU Model Quant Backend TG tok/s
RTX 4090 Mixtral-8x7B Q4_K_M ik_llama.cpp 58.7
RTX 4090 Llama-3-70B Q4_K_M ik_llama.cpp 42.3
RTX 4090 Llama-3-70B Q4_K_M llama.cpp 35.1
RTX 3090 Mistral-7B Q4_K_M llama.cpp 112.8
RTX 4060 Phi-3-mini Q4_K_M llama.cpp 68.5

Shell Completions

Dyno supports shell completions via Typer/Click:

# Install completions for your shell
dyno --install-completion

# Show completion script (to manually install)
dyno --show-completion

Supported shells: bash, zsh, fish, powershell.

Architecture

src/llama_dyno/
├── __init__.py    # Package metadata
├── cli.py         # Typer CLI (detect, tune, bench, report, submit)
├── types.py       # Data types (BenchParams, TrialResult, etc.)
├── detect.py      # Hardware fingerprinting (pynvml, nvidia-smi, /proc)
├── bench.py       # llama-bench subprocess driver + parser
├── tune.py        # Search strategy (coarse sweep + hill climb)
├── ollama.py      # Ollama REST API runner
├── lmstudio.py    # LM Studio OpenAI-compatible API runner
├── report.py      # Report generation (JSON + markdown)
└── submit.py      # GitHub PR / Gist submission

Development

git clone https://github.com/lachy/llama-dyno
cd llama-dyno
pip install -e ".[dev]"
pytest

Out of Scope (v1)

  • GUI
  • Multi-GPU
  • Intel backends (planned)
  • Server hosting for results

License

Apache 2.0

Metadata

Release files for llama-dyno 1.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-dyno 1.5.0
File Size Uploaded
llama_dyno-1.5.0.tar.gz 49.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-dyno 1.5.0
File Interpreter ABI Platform
llama_dyno-1.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 90.8 kB

Release files / llama_dyno-1.5.0.tar.gz

Download URL llama_dyno-1.5.0.tar.gz
Size 49.4 kB
Tags Source
SHA-256 checksum
How to use checksums
37eebe67f0ac88c463e695495f60d862fb8e4f8083b004acfc6583bca7f3860a
BLAKE2b-256 checksum
How to use checksums
8b85bcfa6109e13820d9c4f10bd356185f71d93b396b00e6f205ac10e1985116
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release files / llama_dyno-1.5.0-py3-none-any.whl

Download URL llama_dyno-1.5.0-py3-none-any.whl
Size 41.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e2a90a6d2bf00454a82a5fada192250c02b2334f54148e530d9d3ebf96af8c90
BLAKE2b-256 checksum
How to use checksums
d7eb39573aff95d9097c9b8c8c5989651c4c4ef84f03d7763dea1bf8f8d63e6d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release history Release notifications | RSS feed

This release

1.5.0 This release

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page