Skip to main content

High-performance RAG vector compression

Project description

VectorZip

PyPI version License

VectorZip is a high-performance Python library designed to optimize vector database storage and query latency in Retrieval-Augmented Generation (RAG) pipelines. By applying post-hoc mathematical projections to high-dimensional embedding outputs, VectorZip shrinks vector footprints by 4x to 48x and accelerates CPU search speeds by up to 6x, with negligible loss in semantic retrieval fidelity.


Resources


Installation

pip install vectorzip

Compression Modes & Performance Matrix

VectorZip integrates multiple compression backends under a single unified, easy-to-use API. Choose the method that best matches your target model and database constraints:

Compression Mode Method String Target Model Type Compression Factor Search Speedup (CPU) Retrieval Fidelity
Spectral DCT (Flagship) "dct" Standard / Non-Matryoshka 4x - 12x 6.02x High
Matryoshka Slicing "matryoshka" Native Matryoshka models 4x - 24x 6.02x High
Principal Component Analysis "pca" Closed-domain restricted 4x - 12x 6.02x High (Local)
Product Quantization "pq" General dense vectors 8x - 16x 4.20x Moderate
Scalar Quantization "sq8" General dense vectors 4x 1.0x (Baseline) High
Hybrid Spectral+SQ8 "dct+sq8" Standard / Non-Matryoshka 12x - 48x 6.02x High

Note: Search Speedup is measured over a database of 100,000 document vectors on CPU compared to raw high-dimensional scalar quantized (SQ8) search.


1-Minute Quickstart

VectorZip provides a high-level model wrapper, VectorZipModel, that acts as a direct, drop-in replacement for standard SentenceTransformer models, automating offline calibration and online query compression:

1. Calibrate and Compress Database Index

from vectorzip import VectorZipModel

# Wrap your model, specify target dimensions, and choose your method
model = VectorZipModel("BAAI/bge-m3", n_components=128, method="dct+sq8")

# Calibrate and compress your corpus (calibration runs automatically on first batch)
documents = ["VectorZip enables high-performance compression.", "Vector search is now 6x faster."]
compressed_embeddings = model.encode(documents)

# Save the calibrated JSON configuration for production
model.save_pretrained("./vectorzip_calibrated_model")

2. Compress Online Queries in Production

from vectorzip import VectorZipModel

# Load the pre-calibrated JSON configuration in your production environment
production_model = VectorZipModel.from_pretrained("./vectorzip_calibrated_model")

# Compress query embedding instantly to uint8 for database search
query_vector = production_model.encode(["Fast vector search"])

License

VectorZip is released under the Apache 2.0 License.

Performance Benchmarks

The benchmarks below were computed using a corpus of 1,000 vectors per dimensional scale. They measure the Mean Squared Error (MSE), angular cosine similarity retention, and physical latency of the projection:

Model Size (Original) Target Dims Compression Ratio MSE Error Cosine Similarity Retention Latency (1000 Vectors)
BGE-Small/MiniLM-L6 (4x Ratio) 96 4.0x 0.000301 99.97% 11.60 ms
BGE-Base/Nomic-v1.5 (4x Ratio) 192 4.0x 0.000310 99.97% 12.42 ms
BGE-Large/BGE-M3 (4x Ratio) 256 4.0x 0.001628 99.84% 8.33 ms
GTE-Qwen2/OpenAI-Small (4x Ratio) 384 4.0x 0.001392 99.87% 15.27 ms
OpenAI-Large/Cohere-v3 (4x Ratio) 768 4.0x 0.000470 99.95% 32.72 ms

[!NOTE] Fidelity Invariant: For all dimensional scales up to a 4.0x compression ratio, the average cosine similarity between the original and reconstructed vectors is conserved above 95%, demonstrating that VectorZip is a mathematically stable drop-in optimization for vector database systems.

Downstream Retrieval Quality Evaluation

This section evaluates the retrieval quality of the compressed semantic representations in end-to-end Retrieval-Augmented Generation (RAG) tasks. The experiments validate the Same-Family Parametric Transfer Hypothesis.

Theoretical Basis: Post-Hoc Projection vs. Low-Capacity Model Training

Deploying lower-dimensional embedding representations (e.g., 384 or 512 dimensions) is crucial for resource-constrained vector database applications. However, native low-dimensional models in a given family (e.g., BGE-Small) are often severely capacity-limited, lacking the parameters required to represent multilingual structures or extended contexts.

By applying VectorZip's post-hoc spectral compression to a high-capacity model of the same family (e.g., BGE-M3 or Qwen2.5-Large), the compressed vector inherits the semantic properties and parametric knowledge of the high-capacity backbone. This approach achieves superior downstream retrieval accuracy compared to native small models, bypassing the necessity of computationally intensive training or fine-tuning of low-capacity architectures.

Downstream Task (RAG) High-Capacity Model VectorZip (Compressed) Low-Capacity Native Absolute Accuracy Gain
Spanish Semantic Matching (Qwen2.5 1536 to 512) 100.00% (Qwen2.5-Large 1536) 100.00% (VectorZip 512) 33.33% (Qwen2.5-Small 512) +66.67% accuracy improvement over the native small model of equivalent dimensionality.
Spanish Multilingual Sovereignty (BGE 1024 to 384) 100.00% (BGE-M3 1024) 100.00% (VectorZip 384) 66.67% (BGE-Small 384) +33.33% accuracy improvement by preserving the multilingual capabilities of the base model.
Long-Context PDF Retrieval (BGE 1024 to 384) 100.00% (BGE-M3 1024) 100.00% (VectorZip 384) 0.00% (BGE-Small 384) +100.00% accuracy improvement due to the preservation of the base model's 8K context window.
MTEB SciFact (BEIR Suite) (BGE 1024 to 384) 73.46% (BGE-Large 1024) 71.45% (VectorZip 384) 72.00% (BGE-Small 384) 97.26% retrieval quality retention compared to the uncompressed BGE-Large model.

[!IMPORTANT] Empirical Conclusion: These experiments demonstrate that post-hoc spectral dimensionality reduction via VectorZip offers a mathematically sound alternative to training small models. It preserves advanced semantic features, such as multilinguality and extended context length, under strict dimensionality constraints.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vectorzip-0.1.0.tar.gz (178.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vectorzip-0.1.0-py3-none-any.whl (176.5 kB view details)

Uploaded Python 3

File details

Details for the file vectorzip-0.1.0.tar.gz.

File metadata

  • Download URL: vectorzip-0.1.0.tar.gz
  • Upload date:
  • Size: 178.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for vectorzip-0.1.0.tar.gz
Algorithm Hash digest
SHA256 649fffdc940e902ac95300a7fb81ad1e5ecf9cf5c28041af380af0b60666ffb1
MD5 c81a47d21b7fb33245df936a8a9c83b7
BLAKE2b-256 00d5768dd928284c2dbdec13618fcb2971b964b7ed665c6fe7f6623d2a713a50

See more details on using hashes here.

File details

Details for the file vectorzip-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: vectorzip-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 176.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for vectorzip-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e84cbb127190d1cebbacfaabff1204e5c424b00117767a4ff6443740fa3b862b
MD5 085a36b799559701504803d395fdbac3
BLAKE2b-256 e528494fc77528c0a6d0d0ddd26cef22f743e8c334a7b6fe2f5980926241d0b4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page