Skip to main content

SLM Turbo logo

License: MIT

Profile and auto-tune local AI for your specific GPU.

Automated inference optimizer for LLMs. Profiles your GPU, classifies the bottleneck with a roofline model, and prescribes targeted fixes — KV quantization, prefix caching, chunked prefill, backend selection. Outputs a version-controlled recipe, not magic flags.

Architecture

architecture

How it works

slm-turbo analyze   →  probes model topology + GPU. classifies bottleneck.
                        prints: "memory-bound. 2.1 GB KV cache at 4096 tokens."

slm-turbo optimize  →  generates recipe.yaml. 4 optimizers evaluated per (model, GPU).
                        prints: "+52% throughput, -34% latency expected."

slm-turbo warmup    →  downloads weights, JIT-compiles Triton kernels, verifies recipe.
                        one-time. every serve after is instant.

slm-turbo serve     →  reads recipe.yaml, injects config into vLLM, launches server.
                        metrics on :8000/metrics.

slm-turbo status    →  live: tokens/sec, p50/p99 latency, GPU%, KV cache usage.

What's under the hood

Roofline profiler — analytical by default (math, no GPU load). Falls back to empirical forward-pass when model fits VRAM — real CUDA overhead, real fragmentation, real bandwidth numbers. No spreadsheet guesses.

roofline

The roofline model classifies each inference phase as memory-bound or compute-bound by comparing arithmetic intensity (FLOPs per byte) against the GPU's ridge point (peak compute ÷ memory bandwidth).

Custom CUDA kernels — fused KV cache dequant + attention, hand-written CUDA (JIT-compiled via torch.utils.cpp_extension). Per-channel asymmetric quantization (keys 4-bit, values 4-bit) for the standalone path; per-token min/step quantization for the native packed vLLM KV cache. Dequant in registers — never writes fp16 back to DRAM. Targets sm_75+ (GTX 1650 and up).

Recipe engine — 4 optimizers evaluated per (model, GPU) pair. Declarative YAML output. Human-readable, editable, diffable in git.

Pluggable backends — one recipe, any serving runtime. vLLM today. SGLang and TensorRT-LLM planned. No forks — we inject at each backend's config boundary.

Targets

GPUs NVIDIA Turing+ (sm_75). GTX 1650 stress-tested.
Models Any HuggingFace transformer. VLMs, 1B–70B+.
Backends vLLM · SGLang (planned) · TensorRT-LLM (planned)

Install

# From PyPI — base install (analyze / optimize / doctor)
pip install slm-turbo

# Add the vLLM serving backend (needed for `slm-turbo serve`)
pip install "slm-turbo[gpu]"

# Or with uv — run directly without installing:
uvx slm-turbo analyze --model meta-llama/Llama-2-7b-hf

# Or with uv — install the CLI globally:
uv tool install slm-turbo
uv tool install "slm-turbo[gpu]"    # with the vLLM backend

Requires Linux + NVIDIA GPU (Turing/sm_75 or newer). analyze needs PyTorch + Transformers; serve additionally needs vLLM (the [gpu] extra).

From source (development)

git clone <repo-url> && cd slm-turbo
pip install -e ".[gpu]"     # or: uv sync --extra gpu
slm-turbo doctor

Usage

slm-turbo doctor              # check env: CUDA, PyTorch, vLLM, GPU, permissions
slm-turbo analyze  --model meta-llama/Llama-2-7b-hf
slm-turbo optimize --model meta-llama/Llama-2-7b-hf
slm-turbo warmup   --model meta-llama/Llama-2-7b-hf
slm-turbo serve    --model meta-llama/Llama-2-7b-hf --config ~/.slm-turbo/recipes/llama-2-7b*.yaml
slm-turbo status

Benchmarks

Measured on a NVIDIA GTX 1650 (sm_75, 4 GB VRAM) with PyTorch 2.11 + CUDA 13.1. Reproduce with:

python kernels/benchmark/benchmark.py --op-only                 # op-level
python kernels/benchmark/benchmark.py --model-only --model TinyLlama/TinyLlama-1.1B-Chat-v1.0
python kernels/benchmark/benchmark.py --model-only --model Qwen/Qwen2-0.5B-Instruct
python kernels/benchmark/benchmark.py --model-only --model Qwen/Qwen2-1.5B-Instruct

Decode attention — 4-bit packed vs fp16 (op-level)

The 4-bit KV cache reads 4× less memory per decode step, which is the memory-bandwidth-bound decode phase's bottleneck. Our hand-written CUDA kernel is 2–4.7× faster than fp16 attention:

Sequence len D=64 kernel D=64 fp16 speedup D=128 kernel D=128 fp16 speedup
512 60.5 µs 286.5 µs 4.74× 171.6 µs 474.8 µs 2.77×
1024 136.1 µs 525.4 µs 3.86× 396.3 µs 908.9 µs 2.29×
2048 227.8 µs 967.7 µs 4.25× 861.6 µs 1794.0 µs 2.08×

Decode speedup

KV-cache memory per token drops 16–32× (fp16 → 4-bit packed): e.g. 8 MB → 256 KB for D=64, KV=4 at 2048 tokens.

Model-level decode throughput

Full-model decode (transformers eager, interleaved timing) — the kernel is ~0.84–0.89× eager here because at short context the non-attention layers (matmuls) dominate; the 4-bit win compounds as context grows.

Model layers head_dim eager fp16 slm-turbo 4-bit
TinyLlama-1.1B-Chat 22 64 49.2 tok/s 43.8 tok/s (0.89×)
Qwen2-0.5B-Instruct 24 64 45.9 tok/s 38.7 tok/s (0.84×)
Qwen2-1.5B-Instruct 28 128 35.4 tok/s 31.3 tok/s (0.89×)

Model throughput

Native packed KV cache in vLLM (--custom backend)

The vLLM adapter uses a natively allocated 4-bit KV cache (TurboQuant-style hook: get_kv_cache_shape + quantize-on-write do_kv_cache_update), registered at import time with zero vLLM source modifications.

Metric Stock vLLM slm-turbo CUSTOM
KV cache capacity (same 0.72 GiB) 34,160 tokens 121,456 tokens (3.56×)
TinyLlama decode 48.3 tok/s 40.3 tok/s

Honest trade-off: at short context the packed backend trails stock vLLM on throughput (~17%) because per-layer Python launch overhead dominates; the memory win (3.56× context) and the 2–4.7× long-context decode speedup are where the design pays off. Prefill is intentionally delegated to stock SDPA — on sm_75 there are no tensor cores to win with.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

slm_turbo-0.1.0.tar.gz (46.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

slm_turbo-0.1.0-py3-none-any.whl (47.5 kB view details)

Uploaded Python 3

File details

Details for the file slm_turbo-0.1.0.tar.gz.

File metadata

  • Download URL: slm_turbo-0.1.0.tar.gz
  • Upload date:
  • Size: 46.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.4 {"installer":{"name":"uv","version":"0.10.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"42","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slm_turbo-0.1.0.tar.gz
Algorithm Hash digest
SHA256 cfaf31694ffc4adfda9c9baf0c94ed858c3264786aed43de33c537a2d056de90
MD5 6b6eaba5e151cdc6760ecc007323d292
BLAKE2b-256 bdc5459dbfb0a169d74c43856621e5c5731f36898639948ff9cfdee88f6b4f74

See more details on using hashes here.

File details

Details for the file slm_turbo-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: slm_turbo-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 47.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.4 {"installer":{"name":"uv","version":"0.10.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"42","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for slm_turbo-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 aefd31c164755d7d24f1d8472ec7516182ac4b9ab9ec0c97928189a572c7e890
MD5 41984e91530dd91dd1ec3064fd6823bd
BLAKE2b-256 aae1a6fa20ff58975d8f272d0b30dc364e36577c05ad7be031360842bff2e695

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page