Skip to main content

A local-first vector database engine with pluggable indexes, collection management, metadata filtering, and vector similarity search.

Project description

Quantara

A lightweight, local-first vector database built in pure Python with support for collections, metadata filtering, persistence, pluggable indexes, and custom index registration.

Quantara is designed for developers, researchers, and students who want a simple yet extensible vector database that runs entirely on their local machine without requiring external services.


Features

Core Database

  • Collection-based organization
  • CRUD operations
  • Batch document insertion
  • Metadata support
  • Metadata filtering during search
  • Collection cloning
  • Collection renaming
  • Collection deletion
  • Collection statistics

Vector Search

  • Cosine Similarity
  • Dot Product Similarity
  • Euclidean Distance
  • Top-K nearest neighbor search
  • Configurable search metrics

Persistence

  • Local .db storage
  • Automatic persistence
  • JSON import/export
  • Database configuration persistence
  • Index persistence and restoration

Indexing

  • Pluggable indexing architecture
  • Built-in Brute Force index
  • Custom index registration API
  • Collection-specific indexes
  • Index save/load support

Developer Experience

  • Type hints throughout the codebase
  • Dataclass-based records
  • Extensible architecture
  • Local-first design
  • No external database dependencies

Installation

pip install quantara

Or install from source:

git clone https://github.com/<your-username>/quantara.git

cd quantara

pip install -e .

Quick Start

import ollama

from quantara import Database

EMBEDDING_MODEL = "snowflake-arctic-embed:335m"

db = Database(
    "my_database",
    auto_persist=True
)

text = "I want to build AI agents."

embedding = ollama.embeddings(
    model=EMBEDDING_MODEL,
    prompt=text
)["embedding"]

db.insert_doc(
    name=text,
    vector=embedding,
    metadata={
        "category": "ai"
    }
)

query = "How can I build autonomous AI systems?"

query_embedding = ollama.embeddings(
    model=EMBEDDING_MODEL,
    prompt=query
)["embedding"]

results = db.search_doc(
    query_embedding,
    top_k=3
)

print(results)

Collections

Create a collection:

db.create_collection(
    "research"
)

Insert into a collection:

db.insert_doc(
    collection="research",
    name="Transformers Paper",
    vector=embedding,
    metadata={
        "author": "Vaswani"
    }
)

List collections:

db.list_collections()

Rename a collection:

db.rename_collection(
    "research",
    "papers"
)

Clone a collection:

db.clone_collection(
    "papers",
    "papers_backup"
)

Delete a collection:

db.delete_collection(
    "papers_backup"
)

Metadata Filtering

Search only documents matching specific metadata:

results = db.search_doc(
    query_embedding,
    collection="research",
    filters={
        "author": "Vaswani"
    }
)

Batch Insertion

db.batch_insert_docs(
    objects=[
        (
            "Document 1",
            embedding_1,
            {
                "type": "paper"
            }
        ),
        (
            "Document 2",
            embedding_2,
            {
                "type": "article"
            }
        )
    ],
    collection="research"
)

Similarity Metrics

Quantara supports multiple similarity metrics:

db.search_doc(
    query_embedding,
    metric="cosine"
)
db.search_doc(
    query_embedding,
    metric="dot"
)
db.search_doc(
    query_embedding,
    metric="euclidean"
)

Persistence

Persist database state:

db.persist_doc()

Export database:

db.export_to_json(
    "backup.json"
)

Import database:

db.import_from_json(
    "backup.json"
)

Indexes

Saving an Index

db.save_index(
    collection="research",
    path="research.index"
)

Loading an Index

db = Database(
    "my_database",
    index_paths={
        "research": "research.index"
    }
)

Creating Custom Indexes

Quantara provides a pluggable indexing architecture.

Create a custom index:

from quantara import BaseIndex


class MyIndex(BaseIndex):

    def build(self, vectors):
        ...

    def add(self, record_id, vector):
        ...

    def remove(self, record_id):
        ...

    def update(self, record_id, vector):
        ...

    def search(
        self,
        query_vector,
        top_k=3,
        **kwargs
    ):
        ...

    def clear(self):
        ...

    def size(self):
        ...

    def save(self, path):
        ...

    def load(self, path):
        ...

Register the index:

from quantara import register_index

register_index(
    "myindex",
    MyIndex
)

Use it in a collection:

db.create_collection(
    "research",
    index_type="myindex"
)

Statistics

Database statistics:

db.stats()

Collection statistics:

db.collection_stats(
    "research"
)

Architecture

Database
│
├── Collections
│   ├── Records
│   └── Indexes
│
├── Persistence Layer
│
├── Metadata Layer
│
└── Search Layer

Project Goals

Quantara aims to provide:

  • A lightweight local vector database
  • A simple developer experience
  • Extensible indexing support
  • Fast experimentation for AI and RAG applications
  • An educational implementation of vector database concepts

Roadmap

Version 1.x

  • Brute Force Index
  • Metadata Filtering
  • Collection Management
  • Persistent Storage
  • Custom Index Registration

Version 2.x

  • HNSW Index
  • Approximate Nearest Neighbor Search
  • Additional Search Optimizations

Version 3.x

  • Vector Quantization
  • Compressed Vector Storage
  • Advanced Indexing Strategies

License

MIT License


Author

Abhijeet Rajhans

Built as a learning-focused vector database project exploring vector search, indexing systems, persistence, and database architecture in Python.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quantara-0.1.3.tar.gz (18.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

quantara-0.1.3-py3-none-any.whl (18.5 kB view details)

Uploaded Python 3

File details

Details for the file quantara-0.1.3.tar.gz.

File metadata

  • Download URL: quantara-0.1.3.tar.gz
  • Upload date:
  • Size: 18.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for quantara-0.1.3.tar.gz
Algorithm Hash digest
SHA256 2ca0ed9119f7bdb3e42f79a70346ade6afa55a6c47a565bc4f68b8faacaba3da
MD5 263b450003791172f291dc40515672fe
BLAKE2b-256 ff7897ff7731d2d14f10c7b179a307ea7a19d0e935ae8a70fc6601b70461dc29

See more details on using hashes here.

File details

Details for the file quantara-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: quantara-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 18.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for quantara-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 c7473feed85bbdbeb3dc3ec01e0f5ebf2de23f83ef4ebf6b4d128f40be186015
MD5 be70199cbf5b8057d46bf68add9db07e
BLAKE2b-256 7a3ad10b40ac1ef76794ebb363e2181db867d89b24401581c78e5be17cf7bb47

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page