PCA-Matryoshka + TurboQuant compression for embeddings, LLM KV caches, pgvector, and NATS — up to 27x compression
Project description
TurboQuant Pro
Consumer-aware compression for embedding indexes and LLM KV caches. TurboQuant Pro compresses each vector by the metric its downstream consumer actually uses — retrieval recall for indexes, attention/generation quality for KV caches — not reconstruction cosine alone, which is repeatedly shown here to be blind, or even anti-correlated, with quality.
pip install turboquant-pro
tqp replay embedding_glove_recall --small # reproduce the headline retrieval claim — CI-gated, runs in seconds
- Embedding retrieval: 32× compression at recall@10 0.9992 (after matched reranking, on real LaBSE/Gutenberg data) — outperforms RaBitQ and ties OPQ under a matched-byte public benchmark protocol, at 4–20× lower index-build cost.
- KV caches: architecture-aware key quantization avoids a failure that is invisible to reconstruction metrics — PolarQuant keys read 0.995 cosine yet blow perplexity to ≈10⁴; per-channel keys keep it near fp16.
- At scale & in production: compressed-domain search, persisted / larger-than-RAM sharded + memory-mapped indexes, distribution-free rank certificates, one-command replay, and drift monitoring.
Every headline number — with its reproduction status, dataset, one-click notebook, and hardware — is a row in
CLAIMS.md. The acceptance signal everywhere is rank fidelity / a certificate / the consumer's metric — never reconstruction cosine.
The current release is 1.9.0 (larger-than-RAM search + index format v3); the tqp CLI and certification platform shipped in 1.8.0. Full notes: CHANGELOG.md.
Installation
pip install turboquant-pro # core (numpy only) + the `tqp` CLI
pip install turboquant-pro[torch] # + operator tracer (`tqp trace`)
pip install turboquant-pro[fast] # + AVX2 ADC kernel (pybind11)
pip install turboquant-pro[gpu] # + CuPy CUDA 12.x
pip install turboquant-pro[all] # everything (pgvector, FAISS, NATS, …)
30-second embedding compression
The central, best-validated contribution — compress a corpus and search the codes directly:
from turboquant_pro import PCAMatryoshka, ADCIndex
pca = PCAMatryoshka(input_dim=768, output_dim=256).fit(train_vectors)
pipeline = pca.with_quantizer(bits=3) # PCA rotate/truncate + 3-bit TurboQuant
index = ADCIndex(pipeline).add(corpus) # compressed-domain index (~63 B/vec)
ids, scores = index.search(queries, k=10) # single-pass, fast
ids, scores = index.search(queries, k=10, rerank=5, originals=corpus) # exact rerank → ~0.9997
PCAMatryoshka.suggest_output_dim(corpus, target_variance=0.95) picks the truncation dim from the data's spectrum. See the user guide.
Or compress an LLM KV cache
Architecture-aware by quantizer, not just bit-width — per-channel keys + PolarQuant values:
from turboquant_pro import TurboQuantKVCache
cache = TurboQuantKVCache.robust(head_dim=128, n_heads=32, hot_window=512) # asym-NF4 keys + 2% outliers, 4-bit K/V
# or auto-configure from a model name:
from turboquant_pro import AutoConfig
cache = AutoConfig.from_pretrained("llama-3-8b", target="balanced").build_cache() # K4/V3
robust() is one codebook that stays near-fp16 across every architecture tested (including high-GQA models where symmetric NF4 silently collapses). See the KV keys finding.
Choose your workflow
| Goal | Start here |
|---|---|
| Compress a vector index and search it | User guide · fast ADC design |
| Keep an index larger than RAM (memmap / shards) | Production lifecycle |
| Compress an LLM KV cache correctly | KV keys finding · operator-aware quantization |
| Compress model weights | Model-weight guide |
| Certify & third-party-verify a deployment | Certification |
| Reproduce a headline number yourself | CLAIMS.md · claim replay |
| Integrate (pgvector, FAISS, NATS, vLLM, …) | Integrations |
| Drive it from an agent (LangChain / DSPy / MCP / GPT) | Agent tools · agent_tools |
Why consumer-aware compression?
One governing principle ties the whole toolkit together:
Compress a tensor by the metric its consumer uses. Accept or reject on that metric — recall, perplexity, a rank certificate, an expert-set flip rate — never reconstruction cosine on its own.
The sharpest illustration is KV-cache keys. PolarQuant normalizes each key and quantizes its direction, discarding the per-channel scale that softmax(Q·Kᵀ) depends on. On Qwen2.5 that reads a reassuring 0.995 key cosine while perplexity explodes to ≈10⁴; per-channel key quantization at the same width keeps it near fp16 (≈15). A reconstruction-only benchmark cannot see this. Full write-up: docs/KV_KEYS_FINDING.md.
That boundary is now instrumented, so the principle ships as tooling rather than advice:
rank_certificate— turns a measured distortion κ + the corpus's distance-ratio concentration μ̂ into a distribution-free rank floor (Kendall τ ≥ 1−2μ̂); a vacuous floor is the per-corpus "exact reranking required" signal. Emit withtqp certify, re-check withtqp verify(a third party re-hashes the inputs and reproduces the math).a2_probe— selects the quantizer family against the declared consumer (cosine / L2 / attention logits) at calibration time; it reproduces the keys catastrophe as a unit test.operator_trace/operator_sensitivity— infer each tensor's consumer (softmax score / residual / MoE gate / SSM decay) and apply the discipline that operator needs, validated on real Mixtral, OLMoE, and Mamba models.
Backed by the companion theory papers: the-angular-observer (the rank-certificate and (A2) transfer theory) and geometric-observation — the evidence repository home of Paper III (Observation Theory: consumer-relative rate–distortion and the omission floor) and Paper IV (the consumer-relative flip). TurboQuant Pro is Paper II of that series, the compression-as-observation work.
The strategic bet
As models and vector databases scale, the binding constraint shifts from storing the vector to preserving what its consumer reads with it. Reconstruction fidelity — the objective essentially every quantizer optimizes — is increasingly the wrong one: it can show a reassuring 0.995 cosine while the downstream task collapses. TurboQuant Pro is the production embodiment of the alternative: measure the consumer's read operator, spend bits against it, and ship a certificate that the ranking survives — turning a theory program (Paper I's transfer/rank theory, Paper IV's consumer-relative flip) into instruments you run in CI. The bet is that certified, consumer-aware compression becomes table stakes as ratios climb and silent quality regressions get more expensive to miss. That is the axis this project competes on — not one more point on the compression-vs-reconstruction curve, but the certificate that the compression preserved the thing that mattered.
How it works
A per-vector flow — extract L2 norm → unit-normalize → random-orthogonal rotate → Lloyd-Max scalar-quantize → bit-pack — compresses embeddings and KV-cache values near-losslessly (the TurboQuant algorithm, Zandieh et al., ICLR 2026). KV-cache keys take the per-channel path instead (above).
flowchart LR
A["Raw vector<br/>(float32)"] --> B["PCA-Matryoshka<br/>rotate + truncate"]
B --> C["Random orthogonal<br/>rotation"]
C --> D["TurboQuant<br/>Lloyd-Max SQ"]
D --> E["Bit-pack<br/>8×3-bit = 3 B"]
E --> F["Compressed code"]
A -. "L2 norm (kept alongside)" .-> F
classDef out fill:#c8e6c9,stroke:#2e7d32,stroke-width:2px;
class F out;
Benchmark snapshot
At 32× compression, recall@10 on real LaBSE / multilingual-Gutenberg embeddings — all methods reranked identically:
| method | recall@10 (single) | recall@10 (+rerank) | index build |
|---|---|---|---|
| PQ | 0.467 | 0.827 | 142 s |
| RaBitQ (2024 SOTA) | 0.630 | 0.962 | 0.3 s |
| OPQ | 0.780 | 0.999 | 632 s |
| turboquant-pro | 0.784 | 0.9992 | 31 s |
Holds at 1M scale (0.989 +rerank, tying OPQ). Full tables — the 15-method BGE-M3 comparison, the rerank frontier, KV-cache generation quality & memory, the RaBitQ estimator-isolated head-to-head — are in docs/benchmarks/embeddings.md and docs/benchmarks/kv.md. Reproduce end-to-end on public data: notebooks/turboquant_benchmark.ipynb · Colab.
Reading compression ratios. Ratios vary with source dimension, PCA truncation, code width, retained metadata, and whether exact originals are kept for reranking — so distinguish compressed payload vs all-in index storage vs full retrieval-pipeline storage. The canonical headline is 32× at recall@10 0.9992 above; other figures in the benchmark docs (e.g. 27.7× single-vector, 114× pipeline-storage) are labeled by their accounting basis.
At scale & in production
Larger-than-RAM search (1.9.0). TQEIndex persists an index and memory-maps it; a block-streamed path keeps peak RAM at O(n_queries × block) at any corpus size. ShardedIndex splits a corpus into shards that share one PCA basis (scores stay comparable) behind a JSON manifest and fans search across them (parallel across cores; distributed.py partitions shards across machines). On disk, index format v3 bit-packs sub-byte codes — a lossless re-encoding (rankings bit-identical to v2) at 24.1 B/row vs 41 B/row in v2 (2M rows / 4-bit / --no-originals).
from turboquant_pro import TQEIndex, ShardedIndex
idx = TQEIndex.open("index.tqe", mmap=True) # memory-mapped, read/search only
ids, scores = idx.search(queries, k=10, block=100_000) # bounded-RAM, block-streamed
ShardedIndex.create(corpus, "shards/", shard_size=500_000, bits=3) # one shared PCA basis
ids, scores = ShardedIndex.open("shards/manifest.json").search(queries, k=10)
The tqp CLI covers the whole lifecycle — trace → plan → compress → certify → verify → replay → monitor, plus a persisted-index workflow:
tqp plan embeddings --embeddings corpus.npy --target "recall@10 >= 0.90" # recipe on the Pareto frontier
tqp certify --original corpus.npy --reconstructed corpus_q.npy --min-tau 0.8 \
--task "recall@10 >= 0.995" --environment --html report.html # rank floor + provenance envelope
tqp verify certificate.json --original corpus.npy --reconstructed corpus_q.npy # a third party re-checks it
tqp index create --embeddings corpus.npy --out shards/ --bits 3 --shard-size 500000
tqp index search shards/manifest.json --queries q.npy --k 10 --mmap --block 100000
tqp query "SELECT id, score FROM 'x.tqe' ORDER BY COSINE(:q) LIMIT 10 WITH (RECALL >= 0.95)" \
--queries q.npy # declare the target; the planner meets it (1.9.1)
tqp anatomy --npy corpus.npy --k 10 # hub anatomy: what your hubs ARE (1.9.1)
tqp hubdiff --original corpus.npy --reconstructed corpus_q.npy --min-anti-recall 0.9 \
# the tail mean recall hides (1.9.1)
New to hubness and anti-hubs? docs/HUBNESS_PRIMER.md
— the ten-minute primer on why aggregate recall can stay green while your
hardest queries collapse, and how anatomy/hubdiff catch it. Trust the
tail, not the mean.
Full command reference: docs/CLI.md. Also here: QualityMonitor (cosine + (A2) tangential drift, Prometheus metrics), behavioral_agreement (decision-level flip rate + noise floor), hardware-aware profiles (Volta→Blackwell), a portable Triton fused-decode kernel, and cross-framework export (FAISS / Milvus / Qdrant / Weaviate / Pinecone) — see Integrations.
Agents & tool use
Autonomous systems can consume the whole pipeline as tools. turboquant_pro.agent_tools is a small JSON-in/JSON-out surface with docstrings written for tool-calling models — wrapped for LangChain, DSPy, an MCP server, and custom-GPT Actions in examples/agentic/.
from turboquant_pro import best_compression_at_recall, certify_ranking
plan = best_compression_at_recall(corpus, k=10, min_recall=0.99) # "best ratio at 0.99 recall" — accepts on recall, not cosine
cert = certify_ranking(corpus, reconstructed) # the distribution-free rank receipt
The goal is a runtime input: the agent declares the target recall (or the consumer metric, or k) per task, and the tool accepts and certifies against that goal — never reconstruction cosine. That is the project's one rule expressed as an API, and it is why cosine can't be the gate: the coordinate worth keeping is the one that carries the currently-declared goal's geometry. Full guide: examples/agentic/README.md.
Feature & stability matrix
The full table is in docs/api-stability.md (the source of truth); component reference in docs/API.md.
| Tier | Components |
|---|---|
| Stable | PCAMatryoshka, embedding compression pipeline, basic TurboQuantKV, TQE1 format |
| Beta | ADCIndex, TQEIndex (memmap + format v3), ShardedIndex, TurboQuantKVCache, the rank certificate (tqp certify/verify), the (A2) probe + quality monitor, the tqp index lifecycle, the runtime safe-fallback policy, FAISS / pgvector wrappers |
| Experimental | agent tool surface (agent_tools + examples/agentic), tqp query (SQL-ish workload interface), hub anatomy + anti-hub oracle (tqp anatomy/hubdiff), vLLM V1 KV connector (turboquant_pro.connectors — 2.0 roadmap), quantizer plugin registry + conformance kit, CUDA/Triton fused decode, multi-node shard server (distributed.py), vLLM manager, model-weight compressor, PostgreSQL extension, NATS transport |
Scope & honesty: results are strongest on text embeddings and LLM workloads; multimodal APIs/presets exist but are less validated. "Beats RaBitQ" means under our matched-byte public protocol; "robust across every architecture" means every architecture tested. All 4-bit KV quant (asym-NF4 included) still degrades on very-long-generation tasks. Negative results and caveats are kept first-class in docs/claims.md and the soundness audit.
Not to be confused with the similarly-named
turboquant(the HuggingFace KV-cache implementation of the original ICLR TurboQuant algorithm). TurboQuant Pro is a broader, retrieval-first platform that uses that quantizer as one component.
Documentation & reproducibility
- Documentation hub — guides, reference, and the 15-minute reviewer path.
- Agents & MCP:
examples/agentic/— LangChain / DSPy / MCP / custom-GPT wrappers overturboquant_pro.agent_tools. - Reproduce the claims:
CLAIMS.md(claim → notebook → hardware → status) · claim replay guide · evidence ladder. - Benchmarks: embeddings · KV cache · release/library growth.
- Formats: FORMATS.md (TQE1 / TQIX / certificates at a glance) · FORMAT_SPEC.md · CERTIFICATE_SPEC.md.
- Citation:
CITATION.cff(GitHub "Cite this repository") · full BibTeX + acknowledgments indocs/CITATION.md.
License
MIT License. See LICENSE. Author: Andrew H. Bond, San Jose State University.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file turboquant_pro-2.0.0a1.tar.gz.
File metadata
- Download URL: turboquant_pro-2.0.0a1.tar.gz
- Upload date:
- Size: 5.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4062a659aa78d48304fe7ca5282a93283e3d5c8b2ab1ee2e319ae039c8108f24
|
|
| MD5 |
98fa1e2b051c7b8e825ec165058f181f
|
|
| BLAKE2b-256 |
4cff4206fabcd30480dac844d96d522bde8cf8f272758e60f7e15d0790d5634a
|
Provenance
The following attestation bundles were made for turboquant_pro-2.0.0a1.tar.gz:
Publisher:
publish.yml on ahb-sjsu/turboquant-pro
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
turboquant_pro-2.0.0a1.tar.gz -
Subject digest:
4062a659aa78d48304fe7ca5282a93283e3d5c8b2ab1ee2e319ae039c8108f24 - Sigstore transparency entry: 2224528776
- Sigstore integration time:
-
Permalink:
ahb-sjsu/turboquant-pro@ed86648e905825569b3c0acfc84e55b8426db8d8 -
Branch / Tag:
refs/tags/v2.0.0a1 - Owner: https://github.com/ahb-sjsu
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ed86648e905825569b3c0acfc84e55b8426db8d8 -
Trigger Event:
release
-
Statement type:
File details
Details for the file turboquant_pro-2.0.0a1-py3-none-any.whl.
File metadata
- Download URL: turboquant_pro-2.0.0a1-py3-none-any.whl
- Upload date:
- Size: 311.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
221f01f9a8b3f07c1e0a006beb0dc7c0c54247706b50269c3cd724cada39f87f
|
|
| MD5 |
5dd30e398b13f3eb55496ea5c25114d3
|
|
| BLAKE2b-256 |
63ed8a5d88264679f50db63cab5e64a3428cdf9664a128a4126fac46c7bcd225
|
Provenance
The following attestation bundles were made for turboquant_pro-2.0.0a1-py3-none-any.whl:
Publisher:
publish.yml on ahb-sjsu/turboquant-pro
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
turboquant_pro-2.0.0a1-py3-none-any.whl -
Subject digest:
221f01f9a8b3f07c1e0a006beb0dc7c0c54247706b50269c3cd724cada39f87f - Sigstore transparency entry: 2224529874
- Sigstore integration time:
-
Permalink:
ahb-sjsu/turboquant-pro@ed86648e905825569b3c0acfc84e55b8426db8d8 -
Branch / Tag:
refs/tags/v2.0.0a1 - Owner: https://github.com/ahb-sjsu
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ed86648e905825569b3c0acfc84e55b8426db8d8 -
Trigger Event:
release
-
Statement type: