⚡ NanoVector
The SQLite of Vector Search & Episodic Memory for AI Agents
Bare-metal C99 · AVX2+FMA · ARM NEON · FASM x64 · Zero Dependencies · ~120 KB
Quickstart • Why NanoVector? • Benchmarks • Architecture • Python API • Ecosystem
🚀 Why NanoVector?
Modern AI agents and local LLM pipelines are plagued by vector database bloat:
- ChromaDB, Pinecone clients, and FAISS pull hundreds of megabytes of dependencies (
torch,onnxruntime,pydantic,fastapi,duckdb). - Cold Start Penalty: Importing Chroma takes 1.5 to 2.5 seconds, crippling CLI tools, serverless workers (AWS Lambda), and autonomous agent loops.
- The Small-to-Medium Vector Trap: Over 95% of AI agents store between 50 and 50,000 vectors (conversation turns, tool execution history, episodic facts). At this scale, graph traversal (HNSW) incurs heavy pointer indirection, high memory overhead, and non-deterministic recall.
NanoVector solves this by delivering exact, sub-millisecond, brute-force SIMD search directly in CPU cache with zero external dependencies.
| Feature | NanoVector ⚡ | ChromaDB 🐢 | FAISS ⚖️ |
|---|---|---|---|
| Distribution Wheel Size | 38 KB (~120 KB unpacked) | ~120 MB+ | ~50 MB+ |
| External Dependencies | 0 (Zero) | 35+ packages | OpenMP, BLAS |
| Python Cold Import Overhead | < 1 ms (3,000x faster) | ~1,850 ms | ~120 ms |
| Search Latency (N=2,000, 384D) | 0.13 ms (7,478 QPS) | 8.2 ms | 0.22 ms |
| Batch Ingestion Throughput | 1,414,000 vectors/sec | ~25,000 vectors/sec | ~400,000 vectors/sec |
| Storage Format | Single file (.nvec) |
SQLite + DuckDB dirs | Custom binary |
| Zero-Copy NumPy | Yes (Buffer Protocol) | No (copies memory) | Partial |
| GIL Release during Search | Yes (Py_BEGIN_ALLOW_THREADS) |
Partial | Partial |
⚡ Installation
Install the zero-dependency pre-compiled binary wheel in under 1 second:
pip install nanovector
🏁 Quickstart
import nanovector
import numpy as np
# 1. Initialize an index (dim=384 for all-MiniLM-L6-v2, 768 for BERT, 1536 for OpenAI)
index = nanovector.Index(dim=384, metric="cosine")
# 2. Add single embeddings with optional metadata strings
vec = np.random.randn(384).astype(np.float32)
index.add("doc_1", vec, metadata='{"author": "eminsk", "tag": "ai"}')
# 3. Batch addition (Zero-Copy directly from 2D NumPy array)
batch_vecs = np.random.randn(5000, 384).astype(np.float32)
batch_ids = [f"turn_{i}" for i in range(5000)]
index.add_batch(batch_ids, batch_vecs)
# 4. Search top-k nearest neighbors (returns in ~0.15 ms)
query = np.random.randn(384).astype(np.float32)
results = index.search(query, top_k=5)
for r in results:
print(f"[{r.id}] Score: {r.score:.4f} | Metadata: {r.metadata}")
# 5. Single-file instant persistence (.nvec)
index.save("agent_memory.nvec")
# 6. Instant reload from disk
loaded_index = nanovector.load("agent_memory.nvec")
print(f"Reloaded {len(loaded_index)} vectors in {loaded_index.dim}D")
AI Agent Episodic Memory Pattern
Give your LLM agents lightning-fast, persistent long-term memory:
import nanovector
import numpy as np
class AgentEpisodicMemory:
def __init__(self, filepath="agent_brain.nvec", dim=384):
self.filepath = filepath
try:
self.index = nanovector.load(filepath)
except Exception:
self.index = nanovector.Index(dim=dim, metric="cosine")
def remember(self, fact_id: str, embedding: np.ndarray, fact_text: str):
self.index.add(fact_id, embedding, metadata=fact_text)
self.index.save(self.filepath)
def recall(self, query_embedding: np.ndarray, top_k=3):
return self.index.search(query_embedding, top_k=top_k)
# Usage in Agent Loop
memory = AgentEpisodicMemory(filepath="agent_brain.nvec")
# Store facts if brain is empty
if len(memory.index) == 0:
memory.remember("mem_1", np.random.randn(384).astype(np.float32), "User prefers Python, C, and FASM.")
memory.remember("mem_2", np.random.randn(384).astype(np.float32), "NanoVector achieves sub-millisecond search.")
memory.remember("mem_3", np.random.randn(384).astype(np.float32), "Episodic memory saves state in single .nvec file.")
query_vec = np.random.randn(384).astype(np.float32)
recalled_facts = memory.recall(query_vec, top_k=3)
for match in recalled_facts:
print(f"Score: {match.score:.4f} -> Memory: {match.metadata}")
📊 Benchmarks
Real-world benchmarks measured on Intel/AMD x86_64 CPU (AVX2+FMA) using standard 384-dimensional sentence embeddings (all-MiniLM-L6-v2) against NumPy 2.x / OpenBLAS:
Single-Threaded Exact Search Latency
| Dataset Size ($N$) | Metric | NanoVector Latency | NanoVector QPS | NumPy Baseline | Speedup |
|---|---|---|---|---|---|
| 500 vectors | Cosine | 0.0347 ms (34.7 µs) | 28,854 QPS | 0.0828 ms | 2.39x faster |
| 2,000 vectors | Cosine | 0.1337 ms (133.7 µs) | 7,478 QPS | 0.1876 ms | 1.40x faster |
| 10,000 vectors | Cosine | 1.4021 ms | 713 QPS | 1.1617 ms | Comparable (1 thread vs multi-core OpenBLAS) |
| 50,000 vectors | Cosine | 6.7479 ms | 148 QPS | 4.8132 ms | Exact 100% Recall |
High-Throughput Batch Ingestion & Persistence
- Ingestion Throughput: 1,414,447 vectors/sec (20,000 512D vectors ingested in 14.14 ms via Zero-Copy Buffer Protocol).
- Multi-Threaded Concurrency (8 threads): 14,300 QPS (400 concurrent queries executed in 27.97 ms with zero lock contention).
- Persistence Serialization: Save 2,000 vectors in 1.71 ms, load in 3.92 ms (single binary
.nvecfile).
🏛️ Architecture & Acceleration
NanoVector is written in standard C99 with a multi-tiered hardware acceleration pipeline:
┌───────────────────────────────┐
│ Python C-API │
│ (Buffer Protocol / No-GIL) │
└───────────────┬───────────────┘
│
┌───────────────▼───────────────┐
│ NanoVector C99 Core │
│ Top-K In-Place Heap $O(N\log K)$ │
└───────────────┬───────────────┘
│
┌────────────────────────┼────────────────────────┐
│ │ │
┌────────▼────────┐ ┌────────▼────────┐ ┌────────▼────────┐
│ x86_64 AVX2 │ │ ARM64 NEON │ │ FASM x64 │
│ 256-bit FMA │ │ 128-bit FMA │ │ Bare-Metal ASM │
│ (32 floats/iter)│ │ (16 floats/iter)│ │ (Windows x64) │
└─────────────────┘ └─────────────────┘ └─────────────────┘
- 256-bit AVX2 + FMA (
src/nanovector_avx2.c):- 4-way unrolled kernel processing 32 single-precision floats per loop iteration across 4 YMM accumulators.
- Fused multiply-accumulate (
_mm256_fmadd_ps) eliminates intermediate register spills. - Tail handling handles arbitrary vector dimensions with zero padding penalties.
- ARM NEON (
src/nanovector_neon.c):- 128-bit vectorization for Apple Silicon (M1/M2/M3/M4) and AWS Graviton processors.
- 4-way unrolling processing 16 floats per iteration using
vfmaq_f32andvaddvq_f32.
- Pure FASM Assembly (
src/asm/nanovector_x64.asm):- Hand-crafted Windows x64 assembly routines adhering strictly to Microsoft x64 ABI calling conventions (volatile register allocation
ymm0..ymm5, shadow store handling). - Assembles cleanly into a 629-byte object file using Flat Assembler (FASM).
- Hand-crafted Windows x64 assembly routines adhering strictly to Microsoft x64 ABI calling conventions (volatile register allocation
- In-Place Top-$K$ Heap:
- Min-heap / Max-heap maintains the best $K$ matches in $O(N \log K)$.
- Branch-predicted pruning: candidate items with scores worse than the current $K$-th element are discarded in a single CPU clock cycle.
.nvecBinary Specification:- 64-byte aligned header with magic bytes
NVEC\x01. - Contiguous $N \times D \times 4$ raw float block (zero-copy memory-mappable).
- Compact length-prefixed ID and JSON metadata string tables.
- 64-byte aligned header with magic bytes
🐍 Python API Reference
nanovector.Index(dim: int, metric: str = "cosine", normalize: bool = False)
Initializes an embedded vector index.
dim(int): Vector dimensionality (e.g. 384, 768, 1536).metric(str): Distance metric:"cosine": Cosine similarity ($\frac{u \cdot v}{|u| |v|}$), higher is closer. Range $[-1.0, 1.0]$."dot"or"ip": Inner Product ($u \cdot v$), higher is closer."l2"or"euclidean": Squared Euclidean distance ($\sum (u_i - v_i)^2$), lower is closer.
normalize(bool): IfTrue, vectors are automatically L2-normalized upon insertion and search.
Methods
| Method | Description |
|---|---|
add(id: str, vector: Any, metadata: Optional[str] = None) |
Adds a single 1D vector (NumPy array, list, or buffer) with unique ID and optional metadata string. |
add_batch(ids: List[str], vectors: Any, metadatas: Optional[List[str]] = None) |
Adds multiple vectors in batch directly from 2D numpy.ndarray (Zero-Copy). Releases GIL. |
search(query: Any, top_k: int = 10) -> List[Match] |
Searches Top-$K$ nearest neighbors for query vector. Releases GIL during search. |
save(filepath: str) -> None |
Serializes the entire index to a single .nvec binary file on disk. |
load(filepath: str) -> Index |
Classmethod / function loading an index from a .nvec file in sub-millisecond time. |
Properties
index.dim(int): Dimensionality of indexed vectors.index.count(int) orlen(index): Total number of indexed vectors.index.metric(str): Active distance metric.nanovector.version()(str): Library version string (e.g."0.1.0").nanovector.simd_backend()(str): Active hardware acceleration backend ("AVX2+FMA (x86_64)","ARM NEON", etc.).
🌐 High-Performance Systems Ecosystem
nanovector is developed by @eminsk as part of an open-source performance ecosystem:
- ⚡ NanoGEMM — Bare-metal AVX2+FMA SIMD matrix multiplication engine in ~100KB for sub-microsecond CPU neural network inference (
pip install nanogemm). - 📈 yfinance-ta-patterns — Institutional-grade technical pattern scanner with AI Confluence Scoring and LLM prompt generation (
pip install yfinance-ta-patterns). - 🎥 screenvideo — Desktop screen recorder with WASAPI audio and standalone pure x64 FASM edition.
- 📊 xlsx_vievers — Desktop spreadsheet processor with SSE2 SIMD hardware math engine.
- 🔍 StackOverflowAPI — Bilingual desktop client with native FASM x64 search client.
📄 License
MIT License. See LICENSE for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file nanovector-0.1.1.tar.gz.
File metadata
- Download URL: nanovector-0.1.1.tar.gz
- Upload date:
- Size: 34.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ae40fa8e970e233bdc47f96aab70b980e343d2ee6f839121928954281bbbd2f5
|
|
| MD5 |
94f8547c4ba78929ed8347ba24dafbf7
|
|
| BLAKE2b-256 |
769ac20cd9066e0d8b2fe413dd0f096be64dadc75bbbef0e911ed35d06a1b3aa
|
File details
Details for the file nanovector-0.1.1-cp314-cp314-win_amd64.whl.
File metadata
- Download URL: nanovector-0.1.1-cp314-cp314-win_amd64.whl
- Upload date:
- Size: 25.0 kB
- Tags: CPython 3.14, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
053c8c04c689bfdd0492386f1c9a0366a9288b8c5cbbfe558c5c4e27255d1c35
|
|
| MD5 |
5d20600e2150fe0a94c2abe36bd2b736
|
|
| BLAKE2b-256 |
9709a3fe81218aeb6af524eb33519681460d0622e72b8771d6252b966fbd8849
|