Skip to main content

vxdb

PyPI PyPI: vxdb-server CI Python License

The vector database that fits in your pocket. Fast enough to be memory in the loop.

In-process. Rust-powered. Python-native. One pip install away.

pip install vxdb

Use it as a database

import vxdb

db = vxdb.Database(path="./my_data")  # persistent — data survives restarts
collection = db.create_collection("docs", dimension=384)

embed = your_embedding_function  # OpenAI, Sentence Transformers, Cohere, etc.

collection.upsert(
    ids=["a", "b"],
    vectors=[embed("how to train a model"), embed("best pasta recipe")],
    documents=["how to train a model", "best pasta recipe"],
)

collection.query(vector=embed("machine learning"), top_k=5)

embed() is any function that turns text into vectors — see examples/ for OpenAI, Sentence Transformers, LangChain, and Cohere.

That's it. No Docker. No config files. No cloud account. No 500 MB of dependencies.

Use it as agent working memory

vxdb answers in about 100 microseconds in-process, three orders of magnitude below a networked store, so an agent can read and write it on every step of its loop. scratch() allocates a semantic scratchpad, and two tools hand it to the model. The model decides what earns storage; the scratchpad itself rejects near-duplicates:

from agents import Agent, function_tool  # OpenAI Agents SDK; any tool-calling framework works
from vxdb.agent import scratch

wm = scratch(embed)  # ephemeral scratchpad for this run

@function_tool
def remember(fact: str) -> str:
    """Save one durable fact, preference, or constraint."""
    if wm.seen(fact, threshold=0.9):  # loop guard: near-duplicates are rejected
        return "already known"
    wm.add(fact)
    return "stored"

@function_tool
def recall(query: str) -> list[str]:
    """Fetch the stored facts most relevant to the query."""
    return [hit.text for hit in wm.recall(query, k=5)]

agent = Agent(
    name="assistant",
    instructions=(
        "The moment the user states a durable fact, preference, or constraint, "
        "call remember() with it. Never store small talk. Call recall() before "
        "answering anything about earlier context."
    ),
    tools=[remember, recall],
)

Every remember call is the model's own retention decision, and each one costs microseconds of store time, so consulting memory on every step is effectively free. The notebook runs this live, plus a with/without A/B where a bounded-window agent fails without memory and succeeds with it.

Why developers choose vxdb

Stupid fast

The entire hot path — distance computation, HNSW traversal, BM25 scoring, mmap I/O — is pure Rust. Search releases the GIL, so concurrent queries run in parallel across cores instead of serializing. Your Python code calls directly into compiled native code via PyO3. No serialization overhead. No REST round-trips. No subprocess.

Stupid light

A single native wheel under 5 MB with zero Python dependencies. Starts in under 10 ms. No numpy. No scipy. No protobuf. No grpcio version conflicts. Just pip install vxdb and you're done.

Runs anywhere

Laptop. CI pipeline. Raspberry Pi. AWS Lambda. Docker container. Air-gapped server. Anywhere Python runs, vxdb runs. No infrastructure required to get started — scale up to a standalone server when you need it.

Hybrid search built-in

Vector similarity + BM25 keyword matching fused via Reciprocal Rank Fusion. One API call. Tunable alpha parameter. No separate search engine needed. No Elasticsearch sidecar.

Other databases like Qdrant, Milvus, and Zvec support hybrid search too — but they require you to run a separate sparse encoder (BM25 or SPLADE) yourself and pass pre-computed sparse vectors. vxdb computes BM25 internally from the documents you already upserted. One call: hybrid_query(vector=..., query="text", alpha=0.5). No extra step.

Dual-mode: embedded + server

Many databases now offer an "embedded" mode — but the implementations vary widely. Qdrant's local mode is a Python reimplementation (not their Rust engine). Weaviate embedded downloads a Go binary and runs it as a subprocess. Milvus Lite works but is limited to Linux/macOS and recommended for <1M vectors.

vxdb's embedded mode is the real Rust engine compiled directly into a Python extension via PyO3. No serialization. No subprocess. No network. And the same engine powers the standalone REST server — start in a notebook, scale to multi-client HTTP when you're ready. No rewrite.

The full picture

vxdb architecture

Quick Start

import vxdb

# Persistent (data survives restarts)
db = vxdb.Database(path="./my_data")

# Or in-memory (ephemeral, great for prototyping)
# db = vxdb.Database()

collection = db.create_collection("docs", dimension=384, metric="cosine")

Insert vectors

collection.upsert(
    ids=["a", "b", "c"],
    vectors=[[0.1, 0.2, ...], [0.3, 0.4, ...], [0.5, 0.6, ...]],
    metadata=[{"type": "article"}, {"type": "blog"}, {"type": "article"}],
    documents=["intro to ML", "my favorite recipes", "deep learning guide"],
)

Search — four ways

# 1. Vector similarity
results = collection.query(vector=[0.1, 0.2, ...], top_k=5)

# Trade recall for latency on HNSW: raise ef_search for more accurate results,
# lower it for faster ones. Defaults to the index setting when omitted.
results = collection.query(vector=[0.1, ...], top_k=5, ef_search=200)

# 2. Filtered (metadata constraints)
results = collection.query(
    vector=[0.1, ...], top_k=5,
    filter={"type": {"$eq": "article"}}
)

# 3. Hybrid (vector + keyword — the sweet spot)
results = collection.hybrid_query(
    vector=[0.1, ...],
    query="machine learning",
    top_k=5,
    alpha=0.5,  # 0=keyword only, 1=vector only
)

# 4. Keyword only (BM25)
results = collection.keyword_search(query="machine learning", top_k=5)

Every result returns {"id", "score", "metadata", "document"}.

Installation

pip install vxdb

That's the whole thing. Works on macOS, Linux, Windows. Python 3.11+.

For the HTTP client (talking to a remote vxdb server):

pip install 'vxdb[server]'

Embedding Providers

vxdb stores pre-computed vectors — bring any embedding model you want. We have step-by-step notebooks for each:

Provider Install API Key? Notebook
OpenAI pip install openai Yes examples/openai_embeddings.ipynb
Sentence Transformers pip install sentence-transformers No (local) examples/sentence_transformers.ipynb
LangChain (any provider) pip install langchain-openai Depends examples/langchain_integration.ipynb
Cohere pip install cohere Yes examples/cohere_embeddings.ipynb
Ollama (local LLMs) pip install ollama No (local)

Or let vxdb embed for you. Attach an embedding_function to a collection and work in text — pass documents to upsert and query_text to query:

from vxdb import Database, EmbeddingFunction

class MyEmbedder(EmbeddingFunction):
    def embed(self, texts: list[str]) -> list[list[float]]:
        return your_model.encode(texts)

db = Database()
docs = db.create_collection("docs", embedding_function=MyEmbedder())  # dimension inferred

docs.upsert(ids=["a", "b"], documents=["how to train a model", "best pasta recipe"])
docs.query(query_text="machine learning", top_k=5)

The embedding_function can be an EmbeddingFunction subclass or any callable list[str] -> list[list[float]]. Passing vectors/vector explicitly always works and bypasses embedding — vxdb never requires or imports your model library.

Server Mode

Same engine, accessed over HTTP. Deploy it as a standalone service.

The server ships as a separate, optional packagepip install vxdb-server adds the vxdb-server binary without touching the lean core vxdb wheel:

# Install the standalone server (separate package, no extra deps)
pip install vxdb-server

# Start it
vxdb-server --host 0.0.0.0 --port 8080

The Python Client lives in the core package — install it with the server extra (which pulls in httpx):

pip install 'vxdb[server]'

Note: server mode is currently in-memory only — data does not persist across restarts. For persistence, use embedded mode (vxdb.Database(path=...)).

Python client:

from vxdb import Client

client = Client("http://localhost:8080")
coll = client.create_collection("docs", dimension=384)
coll.upsert(ids=["a"], vectors=[[0.1, ...]], documents=["hello world"])
results = coll.hybrid_query(vector=[0.1, ...], query="hello", top_k=5)

cURL:

# Create collection
curl -X POST localhost:8080/collections \
  -H "Content-Type: application/json" \
  -d '{"name": "docs", "dimension": 384}'

# Upsert
curl -X POST localhost:8080/collections/docs/upsert \
  -H "Content-Type: application/json" \
  -d '{"ids": ["a"], "vectors": [[0.1, 0.2]], "documents": ["hello world"]}'

# Query
curl -X POST localhost:8080/collections/docs/query \
  -H "Content-Type: application/json" \
  -d '{"vector": [0.1, 0.2], "top_k": 5}'

Docker:

docker build -t vxdb .
docker run -p 8080:8080 vxdb    # ~145 MB Debian-based image

Most vector databases give you vector search OR keyword search. vxdb gives you both, fused intelligently in a single call.

How it works:

  1. You upsert with documents — raw text is tokenized into a built-in BM25 index alongside your vectors
  2. At query time — vector search and BM25 run in parallel, then Reciprocal Rank Fusion merges both ranked lists
  3. You control the blendalpha=1.0 (pure vector) → alpha=0.5 (balanced) → alpha=0.0 (pure keyword)

When to use it: Specific product names. Error codes. Proper nouns. Anything where exact terms matter alongside semantic meaning. See examples/hybrid_search.ipynb for a deep dive with side-by-side comparisons.

results = collection.hybrid_query(
    vector=embed("lightweight laptop for students"),
    query="MacBook Air M4",
    top_k=5,
    alpha=0.5,
)

How vxdb compares

vxdb Zvec (Alibaba) ChromaDB Qdrant Pinecone Milvus Weaviate FAISS
Language Rust C++ (Proxima) Rust (v1.0+) Rust Proprietary Go/C++ Go C++
Embedded mode PyO3, true in-process In-process In-process Python-only local mode No Milvus Lite Subprocess (downloads Go binary) SWIG bindings
Server mode Yes No Yes Yes Cloud only Yes Yes No
pip install just works Yes Yes Yes Yes (local mode) N/A (SaaS) Yes (Milvus Lite) Yes (Linux/macOS) Yes
Python dependencies None (zero) DashText SDK Several numpy, grpcio, etc. N/A grpcio, protobuf, etc. grpcio, etc. numpy
Wheel size ~5 MB ~30 MB ~20 MB ~50 MB N/A ~50 MB+ ~100 MB+ (downloads binary) ~20 MB
Startup time <10 ms <100 ms <500 ms ~1-3 s (server) N/A ~5-10 s (server) ~3-5 s (server) <10 ms
Hybrid search Built-in BM25 + RRF BM25 + RRF + weighted RRF (dense+sparse) RRF, DBSF Sparse+dense Sparse vectors BM25 + RRF No
BM25 without external encoder Yes (automatic) Requires DashText SDK Yes Requires sparse encoder No Requires sparse encoder Yes No
Sparse vectors No Yes Yes Yes Yes Yes No No
Multi-vector queries No Yes No Yes No No No No
Metadata filtering 10 operators Structured filters Yes Yes Yes Yes Yes No
Persistence mmap + SQLite + WAL Custom engine SQLite Gridstore Cloud RocksDB LSM Manual
Crash recovery WAL Yes Yes (v1.0) Yes Yes Yes Yes No
Quantization No (planned) FP16, INT8, INT4, RaBitQ No Scalar/PQ Yes Yes PQ/BQ PQ/SQ
Docker image ~145 MB N/A (no server) ~200 MB+ ~100 MB No ~1 GB+ ~300 MB+ No
Runs offline Yes Yes Yes Yes No Yes Yes Yes
License Apache 2.0 Apache 2.0 Apache 2.0 Apache 2.0 Proprietary Apache 2.0 BSD-3 MIT

API Reference

Python (Embedded)

# Database
db = vxdb.Database()                  # in-memory (ephemeral)
db = vxdb.Database(path="./my_data")  # persistent (data survives restarts)
db.create_collection(name, dimension, metric="cosine", index="flat")
db.get_collection(name)
db.list_collections()
db.delete_collection(name)

# Collection
collection.upsert(ids, vectors, metadata=None, documents=None)
collection.query(vector, top_k=10, filter=None, ef_search=None)
collection.hybrid_query(vector, query, top_k=10, alpha=0.5)
collection.keyword_search(query, top_k=10)
collection.delete(ids)
collection.count()

vectors accepts a list[list[float]] or a 2-D float32 NumPy array — NumPy arrays are read zero-copy via the buffer protocol, and NumPy is never imported or required.

REST API

Method Endpoint Description
POST /collections Create collection
GET /collections List collections
DELETE /collections/{name} Delete collection
POST /collections/{name}/upsert Upsert vectors (+ optional documents)
POST /collections/{name}/query Vector search (+ optional filter)
POST /collections/{name}/hybrid Hybrid vector + keyword search
POST /collections/{name}/keyword BM25 keyword search
POST /collections/{name}/delete Delete vectors by ID
GET /collections/{name}/count Count vectors

Parameters

Parameter Values Default
metric "cosine", "euclidean", "dot" "cosine"
index "flat" (exact), "hnsw" (approximate) "flat"
filter $eq $ne $gt $gte $lt $lte $in $nin $and $or
alpha 0.0 (keyword) to 1.0 (vector) 0.5
ef_search HNSW candidates explored per query — higher lifts recall, costs latency 150 (index default)

Examples

Interactive Jupyter notebooks with step-by-step walkthroughs:

Notebook What you'll build
quickstart.ipynb Every feature in 5 min (no API keys)
openai_embeddings.ipynb Semantic search with OpenAI embeddings
sentence_transformers.ipynb Free, local embeddings (no API key)
langchain_integration.ipynb LangChain + RAG pipeline
cohere_embeddings.ipynb Multilingual search with Cohere
hybrid_search.ipynb Deep dive: vector vs keyword vs hybrid

Development

git clone https://github.com/getmykhan/vxdb.git && cd vxdb

# Rust
cargo build --all
cargo test --all        # 120+ tests

# Python
uv venv .venv && source .venv/bin/activate
uv pip install maturin pytest httpx
maturin develop
PYTHONPATH=python pytest tests/ -v

The codebase is a Cargo workspace:

vxdb/
├── crates/
│   ├── vxdb-core/       # Engine: indexes, distance, storage, hybrid search
│   ├── vxdb-python/     # PyO3 bindings
│   └── vxdb-server/     # Axum REST API server
├── python/vxdb/         # Python package (client SDK, embedding interface)
├── examples/             # Jupyter notebooks
└── tests/                # Python integration tests

Roadmap

  • Persistent collections (mmap + SQLite + WAL) Done
  • SIMD-accelerated distance computation Done (v0.5.1: NEON on arm64, AVX2 on x86_64)
  • Quantization (int8/binary) for reduced memory
  • GPU acceleration (CUDA/Metal)
  • HNSW graph serialization (fast restart for large indexes)
  • Streaming upsert for large datasets
  • Sparse vector support
  • gRPC API
  • Official LangChain VectorStore integration
  • Kubernetes Helm chart
  • Benchmarks suite vs Qdrant, ChromaDB, Zvec, FAISS

License

Apache 2.0

Release files for vxdb 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vxdb 0.5.1
File Size Uploaded
vxdb-0.5.1.tar.gz 71.5 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for vxdb 0.5.1
File
vxdb-0.5.1-cp311-abi3-win_amd64.whl CPython 3.11 abi3 Windows x86-64 Details
vxdb-0.5.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.11 abi3 Linux glibc 2.17+ x86-64 Details
vxdb-0.5.1-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl CPython 3.11 abi3 Linux glibc 2.17+ ARM64 Details
vxdb-0.5.1-cp311-abi3-macosx_11_0_arm64.whl CPython 3.11 abi3 macOS 11.0+ ARM64 Details
vxdb-0.5.1-cp311-abi3-macosx_10_12_x86_64.whl CPython 3.11 abi3 macOS 10.12+ x86-64 Details

Total release size: 7.3 MB

Release files / vxdb-0.5.1.tar.gz

Download URL vxdb-0.5.1.tar.gz
Size 71.5 kB
Tags Source
SHA-256 checksum
How to use checksums
6d96f92b000b2afa5c5ef5d68bebf150c9d0a90eb550e428f50ff93ec46351fa
BLAKE2b-256 checksum
How to use checksums
074a46c6c463c8466f903b3135373e03df9a35adf2615bfc0755e11004cb48b6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release files / vxdb-0.5.1-cp311-abi3-win_amd64.whl

Download URL vxdb-0.5.1-cp311-abi3-win_amd64.whl
Size 1.3 MB
Tags CPython 3.11 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
20ad8b640015322368fa151d3d2e74c4c5c967abd9ceb28a59340dc18b0a21e2
BLAKE2b-256 checksum
How to use checksums
031f36e77617ffe8c30142f3d7fdc8c9429f72b21728d5ebc158843c06a6843d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release files / vxdb-0.5.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL vxdb-0.5.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 1.6 MB
Tags CPython 3.11 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
40640825f9a5866bd220f827f7ada0f5f276fd0bed9b959ab746770edcb7ec41
BLAKE2b-256 checksum
How to use checksums
5cde8bbd593c49333fa7baa92ea52bc9f433457a099d32574631095df0ddf2dc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release files / vxdb-0.5.1-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl

Download URL vxdb-0.5.1-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Size 1.4 MB
Tags CPython 3.11 Linux glibc 2.17+ ARM64 abi3
SHA-256 checksum
How to use checksums
bfb7adc952374b2a8c1795a5b315f927487c250548000188a61e6734cc20160d
BLAKE2b-256 checksum
How to use checksums
37945476eb7a84d9ad35e85c0b262d04e1e9878c791df166ffcdb6dc61983b37
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release files / vxdb-0.5.1-cp311-abi3-macosx_11_0_arm64.whl

Download URL vxdb-0.5.1-cp311-abi3-macosx_11_0_arm64.whl
Size 1.4 MB
Tags CPython 3.11 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
fdafdfc7920df21c002c2e57802804d26c79f708f4f2f3e24d430a5a0c5b2a6e
BLAKE2b-256 checksum
How to use checksums
945744d2580c38666d561b0cb86a106c8796fa8d19b01d00a3985d90e837cc55
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release files / vxdb-0.5.1-cp311-abi3-macosx_10_12_x86_64.whl

Download URL vxdb-0.5.1-cp311-abi3-macosx_10_12_x86_64.whl
Size 1.5 MB
Tags CPython 3.11 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
9ebdc4a413fc6bd19d0da08197ef16654f6d1b25a673b1eb4bb96e5962c56173
BLAKE2b-256 checksum
How to use checksums
f77ee28b0bbef58d0f09f095ed1c4649d935845917a6147f38a311b7247f2752
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release history Release notifications | RSS feed

This release

0.5.1 This release

6 release files

0.5.0

6 release files

0.4.1

6 release files

0.4.0

6 release files

0.3.3

6 release files

0.3.2

6 release files

0.3.1

6 release files

0.3.0

6 release files

0.2.1

6 release files

0.2.0

6 release files

0.1.0

6 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page