Skip to main content

tqkit

Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across LLM inference engines.

What this is

tqkit is a single CLI and Python package that talks to every inference engine that ships TurboQuant+ KV-cache compression:

You bring the inference engine. tqkit autodetects what's installed, runs the canonical benchmark, and prints a reproducible KV-savings table.

Why this exists

KV cache is the dominant memory cost at long context. TurboQuant+ asymmetric (K=FP8, V=4-bit + metadata) shrinks it ~62% (or ~57% accounting for the 4 boundary layers that stay FP16). The savings replicate across engines and hardware vendors. tqkit is the proof, the tool, and the install path.

For a 14B model at 1M tokens of context:

layout KV cache size (all-quantized) fits on MI300X 192GB after weights?
FP16 192 GB no
TQ+ asym (turboquant_k8v4) 73.5 GB headline / ~83 GB realistic with boundary skip yes
TQ+ sym 4-bit (turboquant_4bit_nc) 50.3 GB yes (more headroom)

You can verify the math yourself:

pip install tqkit
tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
tq table --model qwen2.5-14b-instruct-1m

Install

pip install tqkit

Usage

tq backends                                            # autodetect installed engines
tq report --model qwen2.5-14b-instruct-1m --ctx 32K    # KV cache size for one config
tq table --model qwen2.5-14b-instruct-1m               # full layout × ctx grid
tq integrate <backend>                                 # install + serve recipe
tq bench                                               # canonical benchmark (v0.3.0)

Example output:

$ tq report --model qwen2.5-14b-instruct-1m --ctx 1M --layout tq+asym
[KV cache] model: Qwen/Qwen2.5-14B-Instruct-1M
[KV cache] arch: layers=48 kv_heads=8 head_dim=128
[KV cache] layout: tq+asym
[KV cache] per-token: 72.0 KB (vs 192.0 KB FP16)
[KV cache] total @ 1M ctx: 72.0 GB (vs 192.0 GB FP16, 62.5% savings)

Integration recipes

One-page docs for plugging TurboQuant+ into each supported backend live under docs/integrate/.

Docker (AMD ROCm)

docker pull thetom/vllm-turboquant:rocm-7.2
docker run --rm -it \
    --device=/dev/kfd --device=/dev/dri --group-add video --ipc=host \
    -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
    -p 8000:8000 thetom/vllm-turboquant:rocm-7.2 \
        --model Qwen/Qwen2.5-14B-Instruct-1M --kv-cache-dtype turboquant_k8v4

See docker/README.md for build details.

Status

v0.4.0 — alpha. Shipping today:

  • KV math + tq report + tq table
  • Pinned canonical_bench.yml + tq config
  • Engine bridges (tq bench) for llama.cpp, vLLM (CUDA + AMD), MLX-Swift, vllm-swift
  • Integration recipes for all 5 backends (docs/integrate/)
  • Docker scaffold for AMD ROCm (docker/Dockerfile.vllm-amd)
  • Models supported in tq report: Qwen2.5 7B/14B/32B, Qwen3-8B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3-Next-80B-A3B, Llama-3.1 8B/70B, Mistral-7B
  • 39 tests, 92% line coverage, ≥85% gate enforced

See CHANGELOG.md for the full version history.

License

Apache 2.0.

Metadata

Release files for tqkit 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tqkit 0.4.2
File Size Uploaded
tqkit-0.4.2.tar.gz 28.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tqkit 0.4.2
File Interpreter ABI Platform
tqkit-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 57.4 kB

Release files / tqkit-0.4.2.tar.gz

Download URL tqkit-0.4.2.tar.gz
Size 28.8 kB
Tags Source
SHA-256 checksum
How to use checksums
538e13024e143d17cddfb6c398f3f462e0c7a1b541731b92d0bf33d599bf207f
BLAKE2b-256 checksum
How to use checksums
fc82c37073686832df6b1a3cf0312fd052d4d65d2ba99fc79b8fdad5556a276c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.4

Release files / tqkit-0.4.2-py3-none-any.whl

Download URL tqkit-0.4.2-py3-none-any.whl
Size 28.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fc1be50c967d71ccc57ea7fd7db3f0f8c56581aed4d7e6a982e33bcdf99a73b5
BLAKE2b-256 checksum
How to use checksums
0b69dbb5ee2729b1963035c1f6cc5eaa539a259fa489416697837d599fe90485
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.4

Release history Release notifications | RSS feed

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page