Skip to main content

pyturboquant-gpu

GPU-accelerated implementation of TurboQuant, a data-oblivious vector quantization algorithm for compressing high-dimensional vectors with near-optimal distortion. Built on PyTorch for seamless GPU acceleration.

Based on the paper: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (Zandieh et al., ICLR 2026).

Installation

pip install pyturboquant-gpu

Requires PyTorch ≥ 2.0. For development:

git clone https://github.com/pyturboquant/pyturboquant-gpu.git
cd pyturboquant-gpu
pip install -e ".[dev]"

Quick Start

MSE-Optimal Quantization

import torch
from pyturboquant_gpu import quantize_mse, dequantize_mse

# Random vectors (e.g., KV cache embeddings) — works on CPU or GPU
vectors = torch.randn(100, 128)  # or vectors.cuda()

# Quantize at 3 bits per coordinate
quantized = quantize_mse(vectors, bits=3, seed=42)

# Reconstruct
reconstructed = dequantize_mse(quantized)

mse = torch.mean(torch.sum((vectors - reconstructed) ** 2, dim=1))
print(f"MSE: {mse:.4f}")

Inner-Product-Optimal Quantization

Provides unbiased inner product estimates — essential for attention and nearest-neighbor search:

from pyturboquant_gpu import quantize_prod, dequantize_prod

# 4 bits total (3 bits MSE + 1 bit QJL correction)
quantized = quantize_prod(vectors, bits=4, seed=42)
reconstructed = dequantize_prod(quantized)

# Inner products are unbiased: E[⟨y, x̃⟩] = ⟨y, x⟩
query = torch.randn(128)
true_ip = vectors @ query
approx_ip = reconstructed @ query
print(f"Mean IP error: {torch.mean(torch.abs(true_ip - approx_ip)):.4f}")

GPU Usage

# Move to GPU — all operations automatically use CUDA
vectors_gpu = vectors.cuda()
quantized = quantize_mse(vectors_gpu, bits=3, seed=42)
reconstructed = dequantize_mse(quantized)  # result is on GPU

How It Works

TurboQuant is a data-oblivious algorithm — no training or calibration needed:

  1. Random Rotation: Multiply by a random orthogonal matrix → coordinates follow a known Beta distribution
  2. Lloyd-Max Scalar Quantization: Each coordinate independently quantized with a precomputed optimal codebook
  3. QJL Residual Correction (Prod mode): 1-bit Quantized Johnson-Lindenstrauss sketch removes inner-product bias

Theoretical Distortion Bounds (unit vectors)

Bits MSE Distortion Inner Product Distortion
1 ≈ 0.36 ≈ 1.57/d
2 ≈ 0.117 ≈ 0.56/d
3 ≈ 0.03 ≈ 0.18/d
4 ≈ 0.009 ≈ 0.047/d

API Reference

quantize_mse(vectors, bits, dim=None, seed=None, device=None)

  • vectors: Tensor of shape (..., d)
  • bits: int in [1, 8]
  • seed: int or None
  • device: torch.device or None
  • Returns: QuantizedMSE

dequantize_mse(quantized) → Tensor

quantize_prod(vectors, bits, dim=None, seed=None, device=None)

  • bits: int in [2, 8]
  • Returns: QuantizedProd

dequantize_prod(quantized) → Tensor

CPU Version

For a NumPy/SciPy-based CPU-only package:

pip install pyturboquant-cpu

Citation

@article{zandieh2025turboquant,
  title={TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate},
  author={Zandieh, Amir and Daliri, Majid and Hadian, Majid and Mirrokni, Vahab},
  journal={arXiv preprint arXiv:2504.19874},
  year={2025}
}

License

Apache 2.0

Release files for pyturboquant-gpu 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyturboquant-gpu 0.1.0
File Size Uploaded
pyturboquant_gpu-0.1.0.tar.gz 18.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyturboquant-gpu 0.1.0
File Interpreter ABI Platform
pyturboquant_gpu-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 34.3 kB

Release files / pyturboquant_gpu-0.1.0.tar.gz

Download URL pyturboquant_gpu-0.1.0.tar.gz
Size 18.8 kB
Tags Source
SHA-256 checksum
How to use checksums
dbfb2860dfc226d11ebe074fc0b35d398979d92b5eea07ba7baac80fa0394ba0
BLAKE2b-256 checksum
How to use checksums
dc878df932c3dcf2a405f82328c842225e8ab73605464419f0d083c0aa22d102
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / pyturboquant_gpu-0.1.0-py3-none-any.whl

Download URL pyturboquant_gpu-0.1.0-py3-none-any.whl
Size 15.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c2f74263a2cd8c117a28d13293207357b80768c33e3e0aff798f6f44dbb73fb6
BLAKE2b-256 checksum
How to use checksums
b52f54e4ee04fffde3a2830cdf198d2a47225f7da3838d680c1ac6f42d6e564a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page