A local-first vector database engine with pluggable indexes, collection management, metadata filtering, and vector similarity search.
Project description
Quantara
A lightweight, local-first vector database built in pure Python with support for collections, metadata filtering, persistence, pluggable indexes, and custom index registration.
Quantara is designed for developers, researchers, and students who want a simple yet extensible vector database that runs entirely on their local machine without requiring external services.
Features
Core Database
- Collection-based organization
- CRUD operations
- Batch document insertion
- Metadata support
- Metadata filtering during search
- Collection cloning
- Collection renaming
- Collection deletion
- Collection statistics
Vector Search
- Cosine Similarity
- Dot Product Similarity
- Euclidean Distance
- Top-K nearest neighbor search
- Configurable search metrics
Persistence
- Local
.dbstorage - Automatic persistence
- JSON import/export
- Database configuration persistence
- Index persistence and restoration
Indexing
- Pluggable indexing architecture
- Built-in Brute Force index
- Custom index registration API
- Collection-specific indexes
- Index save/load support
Developer Experience
- Type hints throughout the codebase
- Dataclass-based records
- Extensible architecture
- Local-first design
- No external database dependencies
Installation
pip install quantara
Or install from source:
git clone https://github.com/<your-username>/quantara.git
cd quantara
pip install -e .
Quick Start
import ollama
from quantara import Database
EMBEDDING_MODEL = "snowflake-arctic-embed:335m"
db = Database(
"my_database",
auto_persist=True
)
text = "I want to build AI agents."
embedding = ollama.embeddings(
model=EMBEDDING_MODEL,
prompt=text
)["embedding"]
db.insert_doc(
name=text,
vector=embedding,
metadata={
"category": "ai"
}
)
query = "How can I build autonomous AI systems?"
query_embedding = ollama.embeddings(
model=EMBEDDING_MODEL,
prompt=query
)["embedding"]
results = db.search_doc(
query_embedding,
top_k=3
)
print(results)
Collections
Create a collection:
db.create_collection(
"research"
)
Insert into a collection:
db.insert_doc(
collection="research",
name="Transformers Paper",
vector=embedding,
metadata={
"author": "Vaswani"
}
)
List collections:
db.list_collections()
Rename a collection:
db.rename_collection(
"research",
"papers"
)
Clone a collection:
db.clone_collection(
"papers",
"papers_backup"
)
Delete a collection:
db.delete_collection(
"papers_backup"
)
Metadata Filtering
Search only documents matching specific metadata:
results = db.search_doc(
query_embedding,
collection="research",
filters={
"author": "Vaswani"
}
)
Batch Insertion
db.batch_insert_docs(
objects=[
(
"Document 1",
embedding_1,
{
"type": "paper"
}
),
(
"Document 2",
embedding_2,
{
"type": "article"
}
)
],
collection="research"
)
Similarity Metrics
Quantara supports multiple similarity metrics:
db.search_doc(
query_embedding,
metric="cosine"
)
db.search_doc(
query_embedding,
metric="dot"
)
db.search_doc(
query_embedding,
metric="euclidean"
)
Persistence
Persist database state:
db.persist_doc()
Export database:
db.export_to_json(
"backup.json"
)
Import database:
db.import_from_json(
"backup.json"
)
Indexes
Saving an Index
db.save_index(
collection="research",
path="research.index"
)
Loading an Index
db = Database(
"my_database",
index_paths={
"research": "research.index"
}
)
Creating Custom Indexes
Quantara provides a pluggable indexing architecture.
Create a custom index:
from quantara import BaseIndex
class MyIndex(BaseIndex):
def build(self, vectors):
...
def add(self, record_id, vector):
...
def remove(self, record_id):
...
def update(self, record_id, vector):
...
def search(
self,
query_vector,
top_k=3,
**kwargs
):
...
def clear(self):
...
def size(self):
...
def save(self, path):
...
def load(self, path):
...
Register the index:
from quantara import register_index
register_index(
"myindex",
MyIndex
)
Use it in a collection:
db.create_collection(
"research",
index_type="myindex"
)
Statistics
Database statistics:
db.stats()
Collection statistics:
db.collection_stats(
"research"
)
Architecture
Database
│
├── Collections
│ ├── Records
│ └── Indexes
│
├── Persistence Layer
│
├── Metadata Layer
│
└── Search Layer
Project Goals
Quantara aims to provide:
- A lightweight local vector database
- A simple developer experience
- Extensible indexing support
- Fast experimentation for AI and RAG applications
- An educational implementation of vector database concepts
Roadmap
Version 1.x
- Brute Force Index
- Metadata Filtering
- Collection Management
- Persistent Storage
- Custom Index Registration
Version 2.x
- HNSW Index
- Approximate Nearest Neighbor Search
- Additional Search Optimizations
Version 3.x
- Vector Quantization
- Compressed Vector Storage
- Advanced Indexing Strategies
License
MIT License
Author
Abhijeet Rajhans
Built as a learning-focused vector database project exploring vector search, indexing systems, persistence, and database architecture in Python.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file quantara-0.1.4.tar.gz.
File metadata
- Download URL: quantara-0.1.4.tar.gz
- Upload date:
- Size: 18.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a4c199f2cd00e531e53ffa8326699ac93ae429b70a80216c2e1f0970ddc759b1
|
|
| MD5 |
9d81242432d3d316ee948d62dca32e7e
|
|
| BLAKE2b-256 |
ecee3e1606b5efb906632ad86be04b0ccba663af55b34a2d52f04cf0e07b3023
|
File details
Details for the file quantara-0.1.4-py3-none-any.whl.
File metadata
- Download URL: quantara-0.1.4-py3-none-any.whl
- Upload date:
- Size: 18.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f11d1ceaee94c47e5e2f77cbff8ced4a04970c90a50651fb97b065b7dd77c116
|
|
| MD5 |
c67ef4332fff6a22eeab2e45bc6af6a8
|
|
| BLAKE2b-256 |
9f1cb7de58847d70bdc771fd4420002efe52e0916ebf0f943673b6d447d37e0c
|