TurboQuant
Near-optimal KV-cache compression for HuggingFace transformers. Reduces KV cache memory by ~8x at 2 bits with attention quality within ~2.7x of the Shannon limit.
Installation
pip install turboquant-explained
Install PyTorch separately for your hardware first:
- CPU:
pip install torch- CUDA 12.x:
pip install torch --index-url https://download.pytorch.org/whl/cu121- See pytorch.org for all variants.
Quick Start
from transformers import AutoModelForCausalLM, AutoTokenizer
import turboquant
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
# 2-bit keys + values → ~8x memory reduction
cache = turboquant.patch_model(model, b_key=2, b_value=2)
inputs = tokenizer("Hello, world!", return_tensors="pt").to("cuda")
output = model.generate(**inputs, past_key_values=cache, max_new_tokens=200)
print(tokenizer.decode(output[0]))
patch_model reads head_dim from model.config automatically — no manual configuration needed.
Memory savings at a glance
| Bit-width | Memory vs FP16 | Typical use |
|---|---|---|
| b=2 | ~1/8 | Long contexts, aggressive compression |
| b=3 | ~3/16 | Balanced quality / savings |
| b=4 | ~1/4 | Near-lossless |
For implementation details and theory, see the full README on GitHub.
Metadata
Release files for turboquant-explained 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| turboquant_explained-0.1.1.tar.gz | 35.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| turboquant_explained-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 63.0 MB
Release files / turboquant_explained-0.1.1.tar.gz
| Download URL | turboquant_explained-0.1.1.tar.gz |
|---|---|
| Size | 35.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5d78bbddea67d78c9b4e6717b7c269e4b563816da6a1547279b6b028ea1cd2c3
|
|
BLAKE2b-256 checksum How to use checksums |
8df0237ddda6286d5821f36f23a87558c660b03a26c50bbee703de1760917cb2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.8
|
Release files / turboquant_explained-0.1.1-py3-none-any.whl
| Download URL | turboquant_explained-0.1.1-py3-none-any.whl |
|---|---|
| Size | 27.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
50cd2c73e93094b17fe862152c0c5f000713ca7c04eb9c7fdd7f399271a92426
|
|
BLAKE2b-256 checksum How to use checksums |
53ad9e7461e082bcf2a2d68ab8bda1869d3faa0f2ebbec08e60d1d13ab55ffaf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.8
|