tqkit
Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across LLM inference engines.
What this is
tqkit is a single CLI and Python package that talks to every inference engine that ships TurboQuant+ KV-cache compression:
- llama.cpp (TheTom/llama.cpp@feature/turboquant-kv-cache)
- vLLM (CUDA) (TheTom/vllm@feature/turboquant-kv-cache)
- vLLM (AMD ROCm) (TheTom/vllm@feature/turboquant-amd-noautotune)
- MLX-Swift (TheTom/mlx@feature/turboquant-plus)
- vllm-swift plugin
You bring the inference engine. tqkit autodetects what's installed, runs the canonical benchmark, and prints a reproducible KV-savings table.
Why this exists
KV cache is the dominant memory cost at long context. TurboQuant+ asymmetric (K=FP8, V=4-bit + metadata) shrinks it ~62% (or ~57% accounting for the 4 boundary layers that stay FP16). The savings replicate across engines and hardware vendors. tqkit is the proof, the tool, and the install path.
For a 14B model at 1M tokens of context:
| layout | KV cache size (all-quantized) | fits on MI300X 192GB after weights? |
|---|---|---|
| FP16 | 192 GB | no |
TQ+ asym (turboquant_k8v4) |
73.5 GB headline / ~83 GB realistic with boundary skip | yes |
TQ+ sym 4-bit (turboquant_4bit_nc) |
50.3 GB | yes (more headroom) |
You can verify the math yourself:
pip install tqkit
tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
tq table --model qwen2.5-14b-instruct-1m
Install
pip install tqkit
Usage
tq backends # autodetect installed engines
tq report --model qwen2.5-14b-instruct-1m --ctx 32K # KV cache size for one config
tq table --model qwen2.5-14b-instruct-1m # full layout × ctx grid
tq integrate <backend> # install + serve recipe
tq bench # canonical benchmark (v0.3.0)
Example output:
$ tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
[KV cache] model: Qwen/Qwen2.5-14B-Instruct-1M
[KV cache] arch: layers=48 kv_heads=8 head_dim=128
[KV cache] layout: tq+asym
[KV cache] per-token: 72.0 KB (vs 192.0 KB FP16)
[KV cache] total @ 1M ctx: 72.0 GB (vs 192.0 GB FP16, 62.5% savings)
Integration recipes
One-page docs for plugging TurboQuant+ into each supported backend live under docs/integrate/.
- llama.cpp — NVIDIA, Apple, AMD, CPU
- vLLM (NVIDIA CUDA) — A100, H100, RTX 4090
- vLLM (AMD ROCm) — MI300X (the only TQ+ port for AMD anywhere)
- MLX-Swift — Apple Silicon Macs + iPhone
- vllm-swift — Apple Silicon OpenAI-API server
Docker (AMD ROCm)
docker pull thetom/vllm-turboquant:rocm-7.2
docker run --rm -it \
--device=/dev/kfd --device=/dev/dri --group-add video --ipc=host \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
-p 8000:8000 thetom/vllm-turboquant:rocm-7.2 \
--model Qwen/Qwen2.5-14B-Instruct-1M --kv-cache-dtype turboquant_k8v4
See docker/README.md for build details.
Status
v0.4.0 — alpha. Shipping today:
- KV math +
tq report+tq table - Pinned
canonical_bench.yml+tq config - Engine bridges (
tq bench) for llama.cpp, vLLM (CUDA + AMD), MLX-Swift, vllm-swift - Integration recipes for all 5 backends (
docs/integrate/) - Docker scaffold for AMD ROCm (
docker/Dockerfile.vllm-amd) - Models supported in
tq report: Qwen2.5 7B/14B/32B, Qwen3-8B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3-Next-80B-A3B, Llama-3.1 8B/70B, Mistral-7B - 39 tests, 92% line coverage, ≥85% gate enforced
See CHANGELOG.md for the full version history.
License
Apache 2.0.
Metadata
Release files for tqkit 0.4.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tqkit-0.4.2.tar.gz | 28.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tqkit-0.4.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 57.4 kB
Release files / tqkit-0.4.2.tar.gz
| Download URL | tqkit-0.4.2.tar.gz |
|---|---|
| Size | 28.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
538e13024e143d17cddfb6c398f3f462e0c7a1b541731b92d0bf33d599bf207f
|
|
BLAKE2b-256 checksum How to use checksums |
fc82c37073686832df6b1a3cf0312fd052d4d65d2ba99fc79b8fdad5556a276c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.4
|
Release files / tqkit-0.4.2-py3-none-any.whl
| Download URL | tqkit-0.4.2-py3-none-any.whl |
|---|---|
| Size | 28.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fc1be50c967d71ccc57ea7fd7db3f0f8c56581aed4d7e6a982e33bcdf99a73b5
|
|
BLAKE2b-256 checksum How to use checksums |
0b69dbb5ee2729b1963035c1f6cc5eaa539a259fa489416697837d599fe90485
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.4
|