pyturboquant-gpu
GPU-accelerated implementation of TurboQuant, a data-oblivious vector quantization algorithm for compressing high-dimensional vectors with near-optimal distortion. Built on PyTorch for seamless GPU acceleration.
Based on the paper: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (Zandieh et al., ICLR 2026).
Installation
pip install pyturboquant-gpu
Requires PyTorch ≥ 2.0. For development:
git clone https://github.com/pyturboquant/pyturboquant-gpu.git
cd pyturboquant-gpu
pip install -e ".[dev]"
Quick Start
MSE-Optimal Quantization
import torch
from pyturboquant_gpu import quantize_mse, dequantize_mse
# Random vectors (e.g., KV cache embeddings) — works on CPU or GPU
vectors = torch.randn(100, 128) # or vectors.cuda()
# Quantize at 3 bits per coordinate
quantized = quantize_mse(vectors, bits=3, seed=42)
# Reconstruct
reconstructed = dequantize_mse(quantized)
mse = torch.mean(torch.sum((vectors - reconstructed) ** 2, dim=1))
print(f"MSE: {mse:.4f}")
Inner-Product-Optimal Quantization
Provides unbiased inner product estimates — essential for attention and nearest-neighbor search:
from pyturboquant_gpu import quantize_prod, dequantize_prod
# 4 bits total (3 bits MSE + 1 bit QJL correction)
quantized = quantize_prod(vectors, bits=4, seed=42)
reconstructed = dequantize_prod(quantized)
# Inner products are unbiased: E[⟨y, x̃⟩] = ⟨y, x⟩
query = torch.randn(128)
true_ip = vectors @ query
approx_ip = reconstructed @ query
print(f"Mean IP error: {torch.mean(torch.abs(true_ip - approx_ip)):.4f}")
GPU Usage
# Move to GPU — all operations automatically use CUDA
vectors_gpu = vectors.cuda()
quantized = quantize_mse(vectors_gpu, bits=3, seed=42)
reconstructed = dequantize_mse(quantized) # result is on GPU
How It Works
TurboQuant is a data-oblivious algorithm — no training or calibration needed:
- Random Rotation: Multiply by a random orthogonal matrix → coordinates follow a known Beta distribution
- Lloyd-Max Scalar Quantization: Each coordinate independently quantized with a precomputed optimal codebook
- QJL Residual Correction (Prod mode): 1-bit Quantized Johnson-Lindenstrauss sketch removes inner-product bias
Theoretical Distortion Bounds (unit vectors)
| Bits | MSE Distortion | Inner Product Distortion |
|---|---|---|
| 1 | ≈ 0.36 | ≈ 1.57/d |
| 2 | ≈ 0.117 | ≈ 0.56/d |
| 3 | ≈ 0.03 | ≈ 0.18/d |
| 4 | ≈ 0.009 | ≈ 0.047/d |
API Reference
quantize_mse(vectors, bits, dim=None, seed=None, device=None)
- vectors: Tensor of shape
(..., d) - bits: int in
[1, 8] - seed: int or None
- device: torch.device or None
- Returns:
QuantizedMSE
dequantize_mse(quantized) → Tensor
quantize_prod(vectors, bits, dim=None, seed=None, device=None)
- bits: int in
[2, 8] - Returns:
QuantizedProd
dequantize_prod(quantized) → Tensor
CPU Version
For a NumPy/SciPy-based CPU-only package:
pip install pyturboquant-cpu
Citation
@article{zandieh2025turboquant,
title={TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate},
author={Zandieh, Amir and Daliri, Majid and Hadian, Majid and Mirrokni, Vahab},
journal={arXiv preprint arXiv:2504.19874},
year={2025}
}
License
Apache 2.0
Release files for pyturboquant-gpu 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pyturboquant_gpu-0.1.0.tar.gz | 18.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pyturboquant_gpu-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 34.3 kB
Release files / pyturboquant_gpu-0.1.0.tar.gz
| Download URL | pyturboquant_gpu-0.1.0.tar.gz |
|---|---|
| Size | 18.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
dbfb2860dfc226d11ebe074fc0b35d398979d92b5eea07ba7baac80fa0394ba0
|
|
BLAKE2b-256 checksum How to use checksums |
dc878df932c3dcf2a405f82328c842225e8ab73605464419f0d083c0aa22d102
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / pyturboquant_gpu-0.1.0-py3-none-any.whl
| Download URL | pyturboquant_gpu-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c2f74263a2cd8c117a28d13293207357b80768c33e3e0aff798f6f44dbb73fb6
|
|
BLAKE2b-256 checksum How to use checksums |
b52f54e4ee04fffde3a2830cdf198d2a47225f7da3838d680c1ac6f42d6e564a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log