Skip to main content

⚡ NanoVector

The SQLite of Vector Search & Episodic Memory for AI Agents

Bare-metal C99 · AVX2+FMA · ARM NEON · FASM x64 · Zero Dependencies · ~120 KB

PyPI Version Python Versions GitHub Release Open In Colab License: MIT SIMD Zero Dependencies

QuickstartGoogle ColabWhy NanoVector?BenchmarksArchitecturePython APIEcosystem


🚀 Why NanoVector?

Modern AI agents and local LLM pipelines are plagued by vector database bloat:

  • ChromaDB, Pinecone clients, and FAISS pull hundreds of megabytes of dependencies (torch, onnxruntime, pydantic, fastapi, duckdb).
  • Cold Start Penalty: Importing Chroma takes 1.5 to 2.5 seconds, crippling CLI tools, serverless workers (AWS Lambda), and autonomous agent loops.
  • The Small-to-Medium Vector Trap: Over 95% of AI agents store between 50 and 50,000 vectors (conversation turns, tool execution history, episodic facts). At this scale, graph traversal (HNSW) incurs heavy pointer indirection, high memory overhead, and non-deterministic recall.

NanoVector solves this by delivering exact, sub-millisecond, brute-force SIMD search directly in CPU cache with zero external dependencies.

Feature NanoVector ChromaDB 🐢 FAISS ⚖️
Distribution Wheel Size 38 KB (~120 KB unpacked) ~120 MB+ ~50 MB+
External Dependencies 0 (Zero) 35+ packages OpenMP, BLAS
Python Cold Import Overhead < 1 ms (3,000x faster) ~1,850 ms ~120 ms
Search Latency (N=2,000, 384D) 0.13 ms (7,478 QPS) 8.2 ms 0.22 ms
Batch Ingestion Throughput 1,414,000 vectors/sec ~25,000 vectors/sec ~400,000 vectors/sec
Storage Format Single file (.nvec) SQLite + DuckDB dirs Custom binary
Zero-Copy NumPy Yes (Buffer Protocol) No (copies memory) Partial
GIL Release during Search Yes (Py_BEGIN_ALLOW_THREADS) Partial Partial

⚡ Installation

Install the zero-dependency pre-compiled binary wheel in under 1 second:

pip install nanovector

🏁 Quickstart

import nanovector
import numpy as np

# 1. Initialize an index (dim=384 for all-MiniLM-L6-v2, 768 for BERT, 1536 for OpenAI)
index = nanovector.Index(dim=384, metric="cosine")

# 2. Add single embeddings with optional metadata strings
vec = np.random.randn(384).astype(np.float32)
index.add("doc_1", vec, metadata='{"author": "eminsk", "tag": "ai"}')

# 3. Batch addition (Zero-Copy directly from 2D NumPy array)
batch_vecs = np.random.randn(5000, 384).astype(np.float32)
batch_ids = [f"turn_{i}" for i in range(5000)]
batch_metas = [f'{{"turn_id": {i}, "role": "agent"}}' for i in range(5000)]
index.add_batch(batch_ids, batch_vecs, metadatas=batch_metas)

# 4. Search top-k nearest neighbors (returns in ~0.15 ms)
query = np.random.randn(384).astype(np.float32)
results = index.search(query, top_k=5)

for r in results:
    print(f"[{r.id}] Score: {r.score:.4f} | Metadata: {r.metadata}")

# 5. Single-file instant persistence (.nvec)
index.save("agent_memory.nvec")

# 6. Instant reload from disk
loaded_index = nanovector.load("agent_memory.nvec")
print(f"Reloaded {len(loaded_index)} vectors in {loaded_index.dim}D")

AI Agent Episodic Memory Pattern

Give your LLM agents lightning-fast, persistent long-term memory:

import nanovector
import numpy as np

class AgentEpisodicMemory:
    def __init__(self, filepath="agent_brain.nvec", dim=384):
        self.filepath = filepath
        try:
            self.index = nanovector.load(filepath)
        except Exception:
            self.index = nanovector.Index(dim=dim, metric="cosine")

    def remember(self, fact_id: str, embedding: np.ndarray, fact_text: str):
        self.index.add(fact_id, embedding, metadata=fact_text)
        self.index.save(self.filepath)

    def recall(self, query_embedding: np.ndarray, top_k=3):
        return self.index.search(query_embedding, top_k=top_k)

# Usage in Agent Loop
memory = AgentEpisodicMemory(filepath="agent_brain.nvec")

# Store facts if brain is empty
if len(memory.index) == 0:
    memory.remember("mem_1", np.random.randn(384).astype(np.float32), "User prefers Python, C, and FASM.")
    memory.remember("mem_2", np.random.randn(384).astype(np.float32), "NanoVector achieves sub-millisecond search.")
    memory.remember("mem_3", np.random.randn(384).astype(np.float32), "Episodic memory saves state in single .nvec file.")

query_vec = np.random.randn(384).astype(np.float32)
recalled_facts = memory.recall(query_vec, top_k=3)

for match in recalled_facts:
    print(f"Score: {match.score:.4f} -> Memory: {match.metadata}")

🚀 Interactive Google Colab Demo

Run NanoVector interactively in your browser with zero local setup:

Open In Colab

The Interactive Colab Notebook demonstrates:

  • Zero-Setup Installation & Hardware SIMD Detection: Compiles native C/AVX2 on Colab CPU in seconds.
  • 10-line Cosine Similarity Search: Indexing and querying embeddings with JSON metadata.
  • Real-World AI Agent Episodic Memory: Recalling instructions and preferences using sentence-transformers embeddings (all-MiniLM-L6-v2).
  • Single-File .nvec Brain Persistence: Instant binary save and zero-overhead reload.
  • Live 50,000-Vector Benchmark: Measuring ingestion throughput (1M+ vectors/sec) and search latency (~0.1 ms) directly on Colab VM hardware.

📊 Benchmarks

Real-world benchmarks measured on Intel/AMD x86_64 CPU (AVX2+FMA) using standard 384-dimensional sentence embeddings (all-MiniLM-L6-v2) against NumPy 2.x / OpenBLAS:

Single-Threaded Exact Search Latency

Dataset Size ($N$) Metric NanoVector Latency NanoVector QPS NumPy Baseline Speedup
500 vectors Cosine 0.0347 ms (34.7 µs) 28,854 QPS 0.0828 ms 2.39x faster
2,000 vectors Cosine 0.1337 ms (133.7 µs) 7,478 QPS 0.1876 ms 1.40x faster
10,000 vectors Cosine 1.4021 ms 713 QPS 1.1617 ms Comparable (1 thread vs multi-core OpenBLAS)
50,000 vectors Cosine 6.7479 ms 148 QPS 4.8132 ms Exact 100% Recall

High-Throughput Batch Ingestion & Persistence

  • Ingestion Throughput: 1,414,447 vectors/sec (20,000 512D vectors ingested in 14.14 ms via Zero-Copy Buffer Protocol).
  • Multi-Threaded Concurrency (8 threads): 14,300 QPS (400 concurrent queries executed in 27.97 ms with zero lock contention).
  • Persistence Serialization: Save 2,000 vectors in 1.71 ms, load in 3.92 ms (single binary .nvec file).

🏛️ Architecture & Acceleration

NanoVector is written in standard C99 with a multi-tiered hardware acceleration pipeline:

                  ┌───────────────────────────────┐
                  │       Python C-API            │
                  │  (Buffer Protocol / No-GIL)   │
                  └───────────────┬───────────────┘
                                  │
                  ┌───────────────▼───────────────┐
                  │      NanoVector C99 Core      │
                  │   Top-K In-Place Heap $O(N\log K)$  │
                  └───────────────┬───────────────┘
                                  │
         ┌────────────────────────┼────────────────────────┐
         │                        │                        │
┌────────▼────────┐      ┌────────▼────────┐      ┌────────▼────────┐
│   x86_64 AVX2   │      │   ARM64 NEON    │      │    FASM x64     │
│   256-bit FMA   │      │   128-bit FMA   │      │ Bare-Metal ASM  │
│ (32 floats/iter)│      │ (16 floats/iter)│      │  (Windows x64)  │
└─────────────────┘      └─────────────────┘      └─────────────────┘
  1. 256-bit AVX2 + FMA (src/nanovector_avx2.c):
    • 4-way unrolled kernel processing 32 single-precision floats per loop iteration across 4 YMM accumulators.
    • Fused multiply-accumulate (_mm256_fmadd_ps) eliminates intermediate register spills.
    • Tail handling handles arbitrary vector dimensions with zero padding penalties.
  2. ARM NEON (src/nanovector_neon.c):
    • 128-bit vectorization for Apple Silicon (M1/M2/M3/M4) and AWS Graviton processors.
    • 4-way unrolling processing 16 floats per iteration using vfmaq_f32 and vaddvq_f32.
  3. Pure FASM Assembly (src/asm/nanovector_x64.asm):
    • Hand-crafted Windows x64 assembly routines adhering strictly to Microsoft x64 ABI calling conventions (volatile register allocation ymm0..ymm5, shadow store handling).
    • Assembles cleanly into a 629-byte object file using Flat Assembler (FASM).
  4. In-Place Top-$K$ Heap:
    • Min-heap / Max-heap maintains the best $K$ matches in $O(N \log K)$.
    • Branch-predicted pruning: candidate items with scores worse than the current $K$-th element are discarded in a single CPU clock cycle.
  5. .nvec Binary Specification:
    • 64-byte aligned header with magic bytes NVEC\x01.
    • Contiguous $N \times D \times 4$ raw float block (zero-copy memory-mappable).
    • Compact length-prefixed ID and JSON metadata string tables.

🐍 Python API Reference

nanovector.Index(dim: int, metric: str = "cosine", normalize: bool = False)

Initializes an embedded vector index.

  • dim (int): Vector dimensionality (e.g. 384, 768, 1536).
  • metric (str): Distance metric:
    • "cosine": Cosine similarity ($\frac{u \cdot v}{|u| |v|}$), higher is closer. Range $[-1.0, 1.0]$.
    • "dot" or "ip": Inner Product ($u \cdot v$), higher is closer.
    • "l2" or "euclidean": Squared Euclidean distance ($\sum (u_i - v_i)^2$), lower is closer.
  • normalize (bool): If True, vectors are automatically L2-normalized upon insertion and search.

Methods

Method Description
add(id: str, vector: Any, metadata: Optional[str] = None) Adds a single 1D vector (NumPy array, list, or buffer) with unique ID and optional metadata string.
add_batch(ids: List[str], vectors: Any, metadatas: Optional[List[str]] = None) Adds multiple vectors in batch directly from 2D numpy.ndarray (Zero-Copy). Releases GIL.
search(query: Any, top_k: int = 10) -> List[Match] Searches Top-$K$ nearest neighbors for query vector. Releases GIL during search.
save(filepath: str) -> None Serializes the entire index to a single .nvec binary file on disk.
load(filepath: str) -> Index Classmethod / function loading an index from a .nvec file in sub-millisecond time.

Properties

  • index.dim (int): Dimensionality of indexed vectors.
  • index.count (int) or len(index): Total number of indexed vectors.
  • index.metric (str): Active distance metric.
  • nanovector.version() (str): Library version string (e.g. "0.1.0").
  • nanovector.simd_backend() (str): Active hardware acceleration backend ("AVX2+FMA (x86_64)", "ARM NEON", etc.).

🌐 High-Performance Systems Ecosystem

nanovector is developed by @eminsk as part of an open-source performance ecosystem:

  • NanoGEMM — Bare-metal AVX2+FMA SIMD matrix multiplication engine in ~100KB for sub-microsecond CPU neural network inference (pip install nanogemm).
  • 📈 yfinance-ta-patterns — Institutional-grade technical pattern scanner with AI Confluence Scoring and LLM prompt generation (pip install yfinance-ta-patterns).
  • 🎥 screenvideo — Desktop screen recorder with WASAPI audio and standalone pure x64 FASM edition.
  • 📊 xlsx_vievers — Desktop spreadsheet processor with SSE2 SIMD hardware math engine.
  • 🔍 StackOverflowAPI — Bilingual desktop client with native FASM x64 search client.

📄 License

MIT License. See LICENSE for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nanovector-0.1.2.tar.gz (35.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nanovector-0.1.2-cp314-cp314-win_amd64.whl (25.5 kB view details)

Uploaded CPython 3.14Windows x86-64

File details

Details for the file nanovector-0.1.2.tar.gz.

File metadata

  • Download URL: nanovector-0.1.2.tar.gz
  • Upload date:
  • Size: 35.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for nanovector-0.1.2.tar.gz
Algorithm Hash digest
SHA256 040ba032203fcc3146d23e24da37d4d2a8a9b5649959756fe84cc5f05e508855
MD5 d18110a97bf6a1bf3dc50014964517da
BLAKE2b-256 bad80decdc8e17b639b15e82eb85cb0ffafd8b794150208efa6ceb8c03a211f5

See more details on using hashes here.

File details

Details for the file nanovector-0.1.2-cp314-cp314-win_amd64.whl.

File metadata

  • Download URL: nanovector-0.1.2-cp314-cp314-win_amd64.whl
  • Upload date:
  • Size: 25.5 kB
  • Tags: CPython 3.14, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for nanovector-0.1.2-cp314-cp314-win_amd64.whl
Algorithm Hash digest
SHA256 09ec1734a71164e75fb3ba0bccfbcb086875d72f9a5ada9d2eed67f89ca29099
MD5 8da5a61780a6c2a335be6e7d5bde5bea
BLAKE2b-256 140db899d89bdf07877fab4a500b3906f7a163441ed0ef69ac1e13837e711ebe

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.3

44 files

This release

0.1.2 This release

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page