circuit-rag
exact top-k vector search on the gpu. clusters get skipped only when a geometric bound proves none of their vectors can beat the current kth best score, so the results are identical to brute force, not approximate.
fast try (reviewers start here)
pip install -e . --no-build-isolation # 1. build (~1 min, needs nvcc + torch w cuda)
python tests/test_exact.py # 2. ~5s, prints exact: 25/25 queries match brute force
python examples/real_proof.py # 3. ~1 min, real scifact text, downloads its own data
pass = step 2 prints exact: 25/25 ... and step 3 prints 300/300.
try it (details)
needs an nvidia gpu, cuda toolkit (nvcc) and pytorch w cuda.
git clone <this repo> && cd circuit-rag
pip install -e . --no-build-isolation # compiles the cuda ext, ~1 min
python tests/test_exact.py # circuit == brute force
python examples/real_proof.py # real scifact text, downloads its own data
real_proof.py pulls scifact from the public beir server, embeds it w BAAI/bge-small-en-v1.5, runs 300 labelled queries and checks every result against brute force. nothing precomputed. on my rtx 5050 laptop:
exact top-10 match vs brute force: 300/300
leaf groups pruned by bounds (avg): 0%
recall@10 vs SciFact labels: 0.840
the 0% thing
yes thats real. on real text embeddings (384 dims) the bounds cant rule out any clusters, so circuit ends up checking everything. results are still exact, but theres no speedup from pruning there. i tried more roots / higher rank, best case skipped ~7%.
on clustered data it works: tests/test_exact.py (fake clustered vecs) skips ~97% of the child clusters. so the math holds, real embeddings are just too spread out in high dims for a rank-8 bound to separate them. thats whats next.
speed (honest)
- native c++ bench (
circuit_bench, scifact index, rtx 5050 laptop): ~38us p50 per query - python path is about the same as a plain cublas gemv (~370us vs ~330us on 200k vecs), per-query python overhead + 1 host sync for the seed threshold
- under 10k vecs it just does a flat gemv
so ship #1 is about being exact and provably so, not about being faster yet.
whats in here
- 5 cuda kernels in
native/kernels/: root bounds, parent projection, child bounds, leaf scoring, topk - torch extension
circuit_cuda(native/bindings/torch_extension.cpp) - python:
build_circuit_index,circuit_search,CircuitRAG(text in, text out),.circuitfile save/load w mmap native/bin/circuit_bench: standalone benchmark (cmake -S native -B build && cmake --build build)
usage
from circuit_rag import CircuitRAG
rag = CircuitRAG() # bge-small-en-v1.5
rag.add_documents(["mRNA vaccines make cells build an antigen protein.", "CRISPR edits dna."])
for r in rag.search("how do mrna vaccines work?", k=2):
print(r.score, r.text)
or the raw index:
from circuit_rag import build_circuit_index, circuit_search
index = build_circuit_index(X, num_roots=64, children_per_root=16, rank=8) # X: float32 [N, D], unit norm
scores, ids, trace = circuit_search(query, index, k=10) # query: torch tensor [D]
index.save("my.circuit")
scores are inner products so normalise ur embeddings if u want cosine.
wiki
full docs (install, usage, how it works, file format, faq) are in wiki/.
known limits
rankhas to be <= 32 (higher crashes the kernel launch, havent found why yet)add_documentsrebuilds the whole index every call- no deletes, filters, reranking or hybrid search yet
Metadata
Release files for circuit-rag 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| circuit_rag-0.1.0.tar.gz | 102.8 kB | Details |
Release files / circuit_rag-0.1.0.tar.gz
| Download URL | circuit_rag-0.1.0.tar.gz |
|---|---|
| Size | 102.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a9e093023588b40c37effa0608735dfd3abcc97c347c41a287f68333afd8612f
|
|
BLAKE2b-256 checksum How to use checksums |
835b86f7c57b64304b587c63ebeaaccd95ca6204f6fead5adaed7fd859b7caed
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|