Skip to main content

A local-first vector database engine with pluggable indexes, collection management, metadata filtering, and vector similarity search.

Project description

Quantara

A lightweight, local-first vector database built in pure Python with support for collections, metadata filtering, persistence, pluggable indexes, and custom index registration.

Quantara is designed for developers, researchers, and students who want a simple yet extensible vector database that runs entirely on their local machine without requiring external services.


Features

Core Database

  • Collection-based organization
  • CRUD operations
  • Batch document insertion
  • Metadata support
  • Metadata filtering during search
  • Collection cloning
  • Collection renaming
  • Collection deletion
  • Collection statistics

Vector Search

  • Cosine Similarity
  • Dot Product Similarity
  • Euclidean Distance
  • Top-K nearest neighbor search
  • Configurable search metrics

Persistence

  • Local .db storage
  • Automatic persistence
  • JSON import/export
  • Database configuration persistence
  • Index persistence and restoration

Indexing

  • Pluggable indexing architecture
  • Built-in Brute Force index
  • Custom index registration API
  • Collection-specific indexes
  • Index save/load support

Developer Experience

  • Type hints throughout the codebase
  • Dataclass-based records
  • Extensible architecture
  • Local-first design
  • No external database dependencies

Installation

pip install quantara

Or install from source:

git clone https://github.com/<your-username>/quantara.git

cd quantara

pip install -e .

Quick Start

import ollama

from quantara import Database

EMBEDDING_MODEL = "snowflake-arctic-embed:335m"

db = Database(
    "my_database",
    auto_persist=True
)

text = "I want to build AI agents."

embedding = ollama.embeddings(
    model=EMBEDDING_MODEL,
    prompt=text
)["embedding"]

db.insert_doc(
    name=text,
    vector=embedding,
    metadata={
        "category": "ai"
    }
)

query = "How can I build autonomous AI systems?"

query_embedding = ollama.embeddings(
    model=EMBEDDING_MODEL,
    prompt=query
)["embedding"]

results = db.search_doc(
    query_embedding,
    top_k=3
)

print(results)

Collections

Create a collection:

db.create_collection(
    "research"
)

Insert into a collection:

db.insert_doc(
    collection="research",
    name="Transformers Paper",
    vector=embedding,
    metadata={
        "author": "Vaswani"
    }
)

List collections:

db.list_collections()

Rename a collection:

db.rename_collection(
    "research",
    "papers"
)

Clone a collection:

db.clone_collection(
    "papers",
    "papers_backup"
)

Delete a collection:

db.delete_collection(
    "papers_backup"
)

Metadata Filtering

Search only documents matching specific metadata:

results = db.search_doc(
    query_embedding,
    collection="research",
    filters={
        "author": "Vaswani"
    }
)

Batch Insertion

db.batch_insert_docs(
    objects=[
        (
            "Document 1",
            embedding_1,
            {
                "type": "paper"
            }
        ),
        (
            "Document 2",
            embedding_2,
            {
                "type": "article"
            }
        )
    ],
    collection="research"
)

Similarity Metrics

Quantara supports multiple similarity metrics:

db.search_doc(
    query_embedding,
    metric="cosine"
)
db.search_doc(
    query_embedding,
    metric="dot"
)
db.search_doc(
    query_embedding,
    metric="euclidean"
)

Persistence

Persist database state:

db.persist_doc()

Export database:

db.export_to_json(
    "backup.json"
)

Import database:

db.import_from_json(
    "backup.json"
)

Indexes

Saving an Index

db.save_index(
    collection="research",
    path="research.index"
)

Loading an Index

db = Database(
    "my_database",
    index_paths={
        "research": "research.index"
    }
)

Creating Custom Indexes

Quantara provides a pluggable indexing architecture.

Create a custom index:

from quantara import BaseIndex


class MyIndex(BaseIndex):

    def build(self, vectors):
        ...

    def add(self, record_id, vector):
        ...

    def remove(self, record_id):
        ...

    def update(self, record_id, vector):
        ...

    def search(
        self,
        query_vector,
        top_k=3,
        **kwargs
    ):
        ...

    def clear(self):
        ...

    def size(self):
        ...

    def save(self, path):
        ...

    def load(self, path):
        ...

Register the index:

from quantara import register_index

register_index(
    "myindex",
    MyIndex
)

Use it in a collection:

db.create_collection(
    "research",
    index_type="myindex"
)

Statistics

Database statistics:

db.stats()

Collection statistics:

db.collection_stats(
    "research"
)

Architecture

Database
│
├── Collections
│   ├── Records
│   └── Indexes
│
├── Persistence Layer
│
├── Metadata Layer
│
└── Search Layer

Project Goals

Quantara aims to provide:

  • A lightweight local vector database
  • A simple developer experience
  • Extensible indexing support
  • Fast experimentation for AI and RAG applications
  • An educational implementation of vector database concepts

Roadmap

Version 1.x

  • Brute Force Index
  • Metadata Filtering
  • Collection Management
  • Persistent Storage
  • Custom Index Registration

Version 2.x

  • HNSW Index
  • Approximate Nearest Neighbor Search
  • Additional Search Optimizations

Version 3.x

  • Vector Quantization
  • Compressed Vector Storage
  • Advanced Indexing Strategies

License

MIT License


Author

Abhijeet Rajhans

Built as a learning-focused vector database project exploring vector search, indexing systems, persistence, and database architecture in Python.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quantara-0.1.2.tar.gz (17.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

quantara-0.1.2-py3-none-any.whl (17.5 kB view details)

Uploaded Python 3

File details

Details for the file quantara-0.1.2.tar.gz.

File metadata

  • Download URL: quantara-0.1.2.tar.gz
  • Upload date:
  • Size: 17.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for quantara-0.1.2.tar.gz
Algorithm Hash digest
SHA256 0471584cd2b196618db8bb97d9a2ab51d2c4fd5de0dbc157df90a5d28a27290d
MD5 bd6627fb8df0e7e4c78c66a2d0e5f436
BLAKE2b-256 d5e7318e02294a9afcb3201b196d2b01f3737679f2d5fd86aba034f3ff502d91

See more details on using hashes here.

File details

Details for the file quantara-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: quantara-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 17.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for quantara-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 8b9eed6c4da1efa7006c6caf7b86395ac42f335830626bea1dc8ad908c80b69e
MD5 e8c72a3c084c44f7337c4916e06b276d
BLAKE2b-256 391a6be7cf377aa618c2d012985561347e9a16b88f72cd9490a1764841b8eb91

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page