Skip to main content

LangChain HyperspaceDB Integration

PyPI version License

Official LangChain integration for HyperspaceDB - a hyperbolic vector database with Edge-Cloud Federation.

Features

  • 🌐 Hyperbolic Geometry: Poincaré ball model for hierarchical embeddings
  • 🔄 Edge-Cloud Federation: Offline-first with automatic sync
  • 🌳 Merkle Tree Sync: Efficient data replication and verification
  • 🗜️ 1-bit Quantization: 64x memory reduction with minimal accuracy loss
  • 🔍 Built-in Deduplication: Content-based hashing prevents duplicates
  • High Performance: Written in Rust for maximum speed

Installation

pip install langchain-hyperspace

Quick Start

Basic Usage

from langchain_hyperspace import HyperspaceVectorStore
from langchain_openai import OpenAIEmbeddings
from langchain.text_splitter import CharacterTextSplitter
from langchain.document_loaders import TextLoader

# Initialize embeddings
embeddings = OpenAIEmbeddings()

# Create vector store
vectorstore = HyperspaceVectorStore(
    host="localhost",
    port=50051,
    collection_name="my_documents",
    embedding_function=embeddings,
    api_key="your_api_key"  # Optional
)

# Load and split documents
loader = TextLoader("path/to/document.txt")
documents = loader.load()
text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)
docs = text_splitter.split_documents(documents)

# Add documents to vector store
vectorstore.add_documents(docs)

# Search for similar documents
query = "What is the main topic?"
results = vectorstore.similarity_search(query, k=4)

for doc in results:
    print(doc.page_content)

RAG (Retrieval-Augmented Generation) Example

from langchain_hyperspace import HyperspaceVectorStore
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain.chains import RetrievalQA

# Setup
embeddings = OpenAIEmbeddings()
vectorstore = HyperspaceVectorStore(
    host="localhost",
    port=50051,
    collection_name="knowledge_base",
    embedding_function=embeddings
)

# Create RAG chain
llm = ChatOpenAI(model_name="gpt-4")
qa_chain = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="stuff",
    retriever=vectorstore.as_retriever(search_kwargs={"k": 3})
)

# Ask questions
response = qa_chain.run("What are the key features of HyperspaceDB?")
print(response)

Content Deduplication

HyperspaceDB automatically deduplicates content using SHA-256 hashing:

vectorstore = HyperspaceVectorStore(
    host="localhost",
    port=50051,
    collection_name="deduplicated_docs",
    embedding_function=embeddings,
    enable_deduplication=True  # Default
)

# Adding the same text twice will only store it once
vectorstore.add_texts([
    "This is a unique document",
    "This is a unique document",  # Duplicate - will be skipped
    "This is another document"
])

Sync Verification (Edge-Cloud Federation)

Check synchronization status using Merkle Tree digest:

# Get collection digest
digest = vectorstore.get_digest()

print(f"Logical Clock: {digest['logical_clock']}")
print(f"State Hash: {digest['state_hash']}")
print(f"Vector Count: {digest['count']}")
print(f"Bucket Hashes: {len(digest['buckets'])} buckets")

Configuration

Connection Options

vectorstore = HyperspaceVectorStore(
    host="localhost",          # Server host
    port=50051,                # gRPC port
    collection_name="default", # Collection name
    embedding_function=embeddings,
    api_key=None,              # Optional API key
    dimension=1536,            # Vector dimension (must match embeddings)
    metric="l2",               # Distance metric: 'l2', 'cosine', 'dot'
    enable_deduplication=True  # Enable content-based deduplication
)

Distance Metrics

  • l2: Euclidean distance (default)
  • cosine: Cosine similarity
  • dot: Dot product

Advanced Usage

Metadata Filtering

# Add documents with metadata
vectorstore.add_texts(
    texts=["Document 1", "Document 2"],
    metadatas=[
        {"source": "web", "category": "tech"},
        {"source": "pdf", "category": "science"}
    ]
)

# Search with metadata filter (coming soon)
results = vectorstore.similarity_search(
    "technology trends",
    k=5,
    filter={"category": "tech"}
)

Batch Operations

# Add large batches efficiently
texts = [f"Document {i}" for i in range(10000)]
metadatas = [{"index": i} for i in range(10000)]

vectorstore.add_texts(texts, metadatas=metadatas)

Running HyperspaceDB Server

Using Docker

docker run -p 50051:50051 -p 50050:50050 \
  -e HYPERSPACE_API_KEY=your_secret_key \
  hyperspacedb/hyperspace-server:latest

From Source

git clone https://github.com/yourusername/hyperspace-db
cd hyperspace-db
cargo build --release
HYPERSPACE_API_KEY=your_secret_key ./target/release/hyperspace-server

Development

Setup

git clone https://github.com/yourusername/hyperspace-db
cd hyperspace-db/integrations/langchain-python

# Install in development mode
pip install -e ".[dev]"

# Generate protobuf files
./generate_proto.sh

Running Tests

pytest tests/

Code Quality

# Format code
black src/ tests/

# Lint
ruff check src/ tests/

# Type check
mypy src/

Examples

See the examples/ directory for complete examples:

  • rag_chatbot.py: RAG chatbot with memory
  • document_qa.py: Document Q&A system
  • semantic_search.py: Semantic search engine
  • edge_sync.py: Edge-Cloud synchronization demo

Documentation

Performance

HyperspaceDB is optimized for:

  • Insert: 10K+ vectors/second
  • Search: <10ms p99 latency
  • Memory: 64x reduction with 1-bit quantization
  • Sync: Merkle Tree-based differential sync

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.

License

Apache License 2.0 - see LICENSE for details.

Support

Citation

If you use HyperspaceDB in your research, please cite:

@software{hyperspacedb2024,
  title = {HyperspaceDB: Hyperbolic Vector Database with Edge-Cloud Federation},
  author = {HyperspaceDB Team},
  year = {2024},
  url = {https://github.com/yourusername/hyperspace-db}
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

langchain_hyperspace-3.1.6.tar.gz (27.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

langchain_hyperspace-3.1.6-py3-none-any.whl (24.5 kB view details)

Uploaded Python 3

File details

Details for the file langchain_hyperspace-3.1.6.tar.gz.

File metadata

  • Download URL: langchain_hyperspace-3.1.6.tar.gz
  • Upload date:
  • Size: 27.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.4

File hashes

Hashes for langchain_hyperspace-3.1.6.tar.gz
Algorithm Hash digest
SHA256 c4e45836905f45433ae4c9f52695206c3a556e541acd9437a645da78f46e3a4b
MD5 c552cc5b83c51bab4be298241b0f0001
BLAKE2b-256 8d563582a259672dd95eee0743c926918b4d68ffc30c294bdbb27561f487b842

See more details on using hashes here.

File details

Details for the file langchain_hyperspace-3.1.6-py3-none-any.whl.

File metadata

File hashes

Hashes for langchain_hyperspace-3.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 988729ad92b12531a1853181c990ab26f87ef5c8993e9068f7b1ce1509e3e924
MD5 8530d7e23cdb65a6a73561834464a3f9
BLAKE2b-256 10f3ca2d7795c4b7bd41f5957457a7f58161af7eea3cf775b55fafdbb3b8863b

See more details on using hashes here.

Release history Release notifications | RSS feed

3.1.7

2 files

This release

3.1.6 This release

2 files

3.1.5

2 files

3.1.4

2 files

3.1.3

2 files

3.1.2

2 files

3.0.5

2 files

3.0.4

2 files

3.0.3

2 files

3.0.2

2 files

3.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page