Skip to main content

madhava-l2

Deterministic vector search with mathematical guarantees.

Every document excluded from the results carries a proof that it could not be in the top-K — by the Cauchy-Schwarz inequality. Zero bound violations by construction.

PyPI Python C++ License: BSL 1.1 Benchmark


madhava-l2 is a real, pip-installable Python package with a native C++20 core. It answers a question no approximate index (HNSW, IVF, PQ) can answer:

"Prove that your search did not miss a relevant document."

The proof is per-document and mathematical: a Cauchy-Schwarz upper bound on the inner product, which converts into a lower bound on L2². If the bound says a vector cannot be in the top-K, that vector is not in the top-K. No heuristics, no random graphs, no "we think it's fine."

Verified on the official BIGANN-100M L2 ground truth — see Benchmarks.

Table of contents


Installation

pip install madhava-l2

Requirements: Python ≥ 3.8, NumPy. The C++ core ships pre-built in the wheel (manylinux); a C++20 compiler + CMake ≥ 3.20 are needed only when building from source.

Not published to PyPI yet? Install straight from this repo:

pip install git+https://github.com/winnex-ai/madhava-l2.git

Quick start

import numpy as np
import madhava_l2

# 1. Build an engine over your corpus (uint8, shape (n, dim)).
corpus = np.random.randint(0, 256, size=(100_000, 128), dtype=np.uint8)

engine = madhava_l2.build_engine(corpus, dim=128, k=10)
print(f"indexed {engine.num_vectors()} vectors in {engine.build_seconds():.2f}s")

# 2. Search.
query = corpus[0].astype(np.float32)   # (128,) float32
result = engine.search(query)

print(result.indices)                  # top-K dataset ids, ascending L2²
print(result.latency_ms)               # milliseconds
print(result.bound_violations)         # always 0 — the guarantee

That's it. Same query + same data → same result, every time. Deterministic.

Why?

Modern vector search is a heuristic gamble. HNSW builds a random proximity graph and hopes it didn't prune a relevant neighbor; IVF picks clusters and hopes the right one was probed. When the search is a legal discovery, a medical record lookup, or a compliance audit, "hope" is not a defensible answer.

madhava-l2 replaces hope with proof. Each excluded document carries its upper bound; if the bound is below the threshold, exclusion is a theorem, not a guess.

Property HNSW / IVF / PQ madhava-l2
Deterministic (same input → same output) No (random graph) Yes
Proves every exclusion No Yes (Cauchy-Schwarz)
Bound violations Not measurable 0
Rebuild speed (1M) ~40 s ~2.6 s
Exact-recall ceiling reachable No Yes (post-filter)

API

madhava_l2.build_engine(corpus, *, dim=None, stage1_dim=64, k=10, k1_fraction=0.05, postfilter=True, seed=42) -> MadhavaL2

Build an index over a (n, dim) uint8 array.

  • dim — vector dimensionality (defaults to corpus.shape[1]).
  • stage1_dim — dimensionality of the Stage-1 QR projection (64 by default).
  • k — number of results to return.
  • k1_fraction — fraction of the corpus kept after Stage-1 pruning (0.05 = 5%).
  • postfilter — when True, exact L2 is computed on the survivors so the result matches the exact top-K of the surviving set.

engine.search(query: np.ndarray) -> SearchResult

Returns indices, latency_ms, k1, k3, bound_pairs, bound_violations.

engine.search_exact(query: np.ndarray) -> SearchResult

Exhaustive L2 scan over all N vectors — the recall ceiling of your corpus. Use it to measure how close an approximate index gets to the physical limit.

madhava_l2.benchmark_vs_groundtruth(engine, queries, gt_ids, *, query_alignment=1, k=None) -> dict

Evaluate against ground-truth id lists. Returns recall_at_k, ndcg_at_k, latency_ms, and per-query detail.

Metrics

  • madhava_l2.recall_at_k(result, gt_set, k)
  • madhava_l2.ndcg_at_k(result, gt_set, k)
  • madhava_l2.read_bigann_groundtruth(path, n_queries)

The mathematics

For any query q and candidate vector v, the Cauchy-Schwarz inequality bounds the raw inner product:

⟨v, q⟩  ≤  ⟨Pv, Pq⟩  +  ‖v − PᵀPv‖ · ‖q − PᵀPq‖

where P is a QR-orthogonalized (Modified Gram-Schmidt) random projection. Because

‖v − q‖²  =  ‖v‖² + ‖q‖² − 2·⟨v, q⟩

the bound on ⟨v, q⟩ becomes a lower bound on L2²:

‖v − q‖²  ≥  ‖v‖² + ‖q‖² − 2·UB(⟨v, q⟩)

Stage 1 computes this lower bound for every vector and keeps the top-k1 by smallest L2². Any vector pruned here is mathematically proven not to be in the exact top-K. Bound violations = 0 by construction.

Post-filter (optional) computes the exact L2² on the surviving top-k1 and returns the true top-K. Because Stage 1 never prunes a real neighbor, the post-filter recovers everything a perfect scan would find.

The residual ‖v − PᵀPv‖ is computed on the real float32 projection, not the int8-quantized one — this is what the inequality requires, and it is what makes the bound exact rather than approximate.

Benchmarks

Verified 2026-08-04 against the official BIGANN-100M L2 ground truth on a CPU-only machine (28 threads, AVX2+FMA).

Scale Exact-scan ceiling
(R@10)
madhava-l2
(R@10)
Efficiency
10M 0.430 0.430 100%
100M 0.788 0.745 94%

The ceiling column is search_exact — a perfect exhaustive scan over the same subset. madhava-l2 reaches 100% of that ceiling at 10M and 94% at 100M, with 0 bound violations at every scale.

Why is the ceiling not 1.0?

The official BIGANN L2 ground truth was generated in the full 1B space. In a 100M subset, the exact top-K by L2² differs, so even a perfect scan caps at R@10 ≈ 0.79. No index — exact or approximate — can do better on this subset against this ground truth. madhava-l2 gets essentially all of it.

Reproduce:

python -m madhava_l2.benchmark --n 10000000 --nq 50

Or via the C++ executable:

./build/madhava_l2_bench bigann_data/base.u8bin \
    bigann_data/unif_query_10k.u8bin \
    bigann_data/unif_groundtruth_10k.bin 100000000 50 0.05

Honest comparison

We are explicit about where madhava-l2 does not win:

Use case Best tool Why
Lowest latency (sub-ms) HNSW HNSW ≈ 0.45 ms vs madhava ≈ 2.7 ms at 50K×1536D
Provable completeness madhava-l2 Only engine with 0 bound violations + per-doc proof
Frequent index rebuilds madhava-l2 Build ≈ 2.6 s (1M) vs HNSW ≈ 40 s
Regulated / auditable retrieval madhava-l2 Deterministic, per-document audit trail

If you need raw speed, use HNSW — it is excellent. madhava-l2 is for the regions where "fast but unprovable" is a liability: legal discovery, medical records, financial compliance, government audits, and RAG systems that must not silently drop a relevant document.

Build from source

# Wheel + sdist (pip-installable)
python -m build

# C++ library only
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
ctest --test-dir build        # C++ unit tests

# Python tests
python -m pytest tests/python/

License

Business Source License 1.1 — same as the Winnex stack. Free to use for evaluation and non-production work. Commercial use requires a license.

pay@winnex.ai · Winnex Brasil Soluções Empresariais LTDA-ME · Goiânia, Brazil

Release files for madhava-l2 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for madhava-l2 1.0.0
File Size Uploaded
madhava_l2-1.0.0.tar.gz 28.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for madhava-l2 1.0.0
File Interpreter ABI Platform
madhava_l2-1.0.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.28+ x86-64, Linux glibc 2.27+ x86-64 Details

Total release size: 318.3 kB

Release files / madhava_l2-1.0.0.tar.gz

Download URL madhava_l2-1.0.0.tar.gz
Size 28.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0d29c66db22f73c8c636755fa2d3af103834b364300a8855d48b7382f599c2c8
BLAKE2b-256 checksum
How to use checksums
cc617b21e611bf419a741db7b5e0fa2e29a6877af61dfe20cc3711d941924512
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 4, 2026.

Transparency log

Release files / madhava_l2-1.0.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Download URL madhava_l2-1.0.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Size 289.7 kB
Tags CPython 3.12 Linux glibc 2.27+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
a07eb8f1725bf39b53fda9f2fdf554377d78e5f7faab2fe4d449be18b3266a07
BLAKE2b-256 checksum
How to use checksums
5d3c0f50955c070bed34a645fcf6c8e675d373b36a7d326f6909fa3d8bb5ff1a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page