An integration package connecting PolarDB-X and LangChain
Project description
🦜️🔗 LangChain PolarDB-X
A powerful integration between LangChain and PolarDB-X, enabling native vector search capabilities for AI applications.
Overview
LangChain PolarDB-X provides seamless integration between LangChain, a framework for building applications with large language models (LLMs), and PolarDB-X with native vector search support. This integration enables efficient vector storage and retrieval for AI applications like semantic search, recommendation systems, and RAG (Retrieval Augmented Generation).
PolarDB-X is a cloud-native distributed database system developed by Alibaba Cloud, featuring native HNSW-based vector index support that delivers high-performance approximate nearest neighbor (ANN) search directly within the database engine.
Requirements
- Python 3.9+
- PolarDB-X with vector index support
- MySQL connector:
mysql-connector-python>=8.0.0(included in package dependencies) - Async support:
pip install langchain-polardbx[async](optional) - MMR search:
pip install langchain-polardbx[mmr](optional)
Enable Vector Index
PolarDB-X disables the vector index feature by default (vidx_disabled = ON). You need to enable it before using this package:
-- Enable vector index (run as admin/root on DN node)
SET GLOBAL vidx_disabled = OFF;
This setting takes effect immediately for new connections. No restart required.
All transaction isolation levels (READ-COMMITTED, REPEATABLE-READ, SERIALIZABLE) are supported — choose according to your business needs.
Note: Some advanced features (e.g., inner product distance, index monitoring,
EF_CONSTRUCTIONparameter) require newer PolarDB-X versions. The package automatically detects available capabilities and adapts accordingly.
Features
- Native Vector Storage: Store embeddings using PolarDB-X's native
VECTOR(N)data type - HNSW Index: Efficient approximate nearest neighbor search with configurable
MandEF_CONSTRUCTIONparameters - Multiple Distance Metrics: Support for Cosine, Euclidean, and Inner Product distance
- Similarity Search: Perform efficient similarity searches with score thresholds
- MMR Search: Maximal Marginal Relevance search for diverse results
- Metadata Filtering: Filter search results by metadata with rich operators (
$eq,$ne,$gt,$gte,$lt,$lte,$in,$nin,$like) - Dynamic Index Management: Create, drop, and rebuild vector indexes at runtime without recreating tables
- Search Mode Control: Switch between ANN (index-accelerated) and KNN (full-scan) modes per query
- Per-Query Tuning: Adjust
ef_searchon a per-query basis for accuracy/latency trade-offs - Index Health Monitoring: Runtime statistics, index health diagnostics, and preload checks
- Batch Operations: Efficient batch insert and bulk upsert with configurable batch size
- Full Async Support: All public methods have async equivalents (
aadd_texts,asimilarity_search, etc.) - Dual-Version Compatibility: Automatically detects database capabilities and adapts SQL accordingly
- Partitioned Table Support: Create partitioned vector tables with HASH/KEY partitioning
- Connection Pooling: Built-in connection pool with automatic retry logic
Installation
pip install langchain-polardbx
Optional Dependencies
For async support:
pip install langchain-polardbx[async]
For MMR search support:
pip install langchain-polardbx[mmr]
For using OpenAI embeddings:
pip install langchain-openai
For using DashScope embeddings (Alibaba Cloud):
pip install langchain-community dashscope
Quick Start
Basic Usage
from langchain_polardbx import PolarDBXVectorStore
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings()
vectorstore = PolarDBXVectorStore(
host="your-polardbx-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
embedding=embeddings,
table_name="my_vectors",
distance_strategy="cosine", # or "euclidean", "inner_product"
hnsw_m=16, # HNSW index M parameter (3-200)
)
# Add texts
ids = vectorstore.add_texts(["Hello world", "PolarDB-X is great"])
# Similarity search
results = vectorstore.similarity_search("Hello", k=2)
for doc in results:
print(f"- {doc.page_content}")
# Search with scores
results_with_scores = vectorstore.similarity_search_with_score("Hello", k=2)
for doc, score in results_with_scores:
print(f"[Score: {score:.4f}] {doc.page_content}")
Using DashScope Embeddings
from langchain_polardbx import PolarDBXVectorStore
from langchain_community.embeddings import DashScopeEmbeddings
embeddings = DashScopeEmbeddings(
model="text-embedding-v4",
dashscope_api_key="your-dashscope-api-key",
)
vectorstore = PolarDBXVectorStore(
host="your-polardbx-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
embedding=embeddings,
table_name="langchain_vectors",
)
Usage Examples
Creating from Documents
from langchain_core.documents import Document
documents = [
Document(page_content="Hello world", metadata={"source": "greeting"}),
Document(page_content="LangChain is great", metadata={"source": "review"}),
]
vectorstore = PolarDBXVectorStore.from_documents(
documents=documents,
embedding=embeddings,
host="your-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
)
Search with Metadata Filter
# Add texts with metadata
texts = ["Apple is a fruit", "Banana is yellow", "Car is a vehicle"]
metadatas = [
{"category": "fruit", "price": 5},
{"category": "fruit", "price": 3},
{"category": "vehicle", "price": 20000},
]
vectorstore.add_texts(texts, metadatas=metadatas)
# Simple equality filter
results = vectorstore.similarity_search("yellow things", k=2, filter={"category": "fruit"})
# Operator filter: price > 2 AND category = fruit
results = vectorstore.similarity_search(
"fresh fruit",
k=2,
filter={"category": "fruit", "price": {"$gt": 2}},
)
Supported filter operators:
| Operator | SQL | Description |
|---|---|---|
$eq |
= |
Equal (default for simple values) |
$ne |
!= |
Not equal |
$gt |
> |
Greater than |
$gte |
>= |
Greater than or equal |
$lt |
< |
Less than |
$lte |
<= |
Less than or equal |
$in |
IN |
In a list of values |
$nin |
NOT IN |
Not in a list of values |
$like |
LIKE |
Pattern match |
Score Threshold Filtering
# Only return results with distance <= 0.8 (lower distance = more similar)
results = vectorstore.similarity_search_with_score(
"Hello",
k=10,
score_threshold=0.8,
)
Maximal Marginal Relevance (MMR) Search
# MMR search for diverse results
results = vectorstore.max_marginal_relevance_search(
"technology",
k=4,
fetch_k=20,
lambda_mult=0.5, # 0 = max diversity, 1 = max relevance
)
Search Mode Control
# Force ANN (use vector index for HNSW acceleration)
results = vectorstore.similarity_search("query", k=10, search_type="ann")
# Force KNN (full table scan, bypass vector index)
results = vectorstore.similarity_search("query", k=10, search_type="knn")
# Let the optimizer decide (default)
results = vectorstore.similarity_search("query", k=10, search_type="auto")
# Tune ef_search per query (higher = more accurate, slower)
results = vectorstore.similarity_search("query", k=10, ef_search=100)
Dynamic Vector Index Management
# Create a vector index at runtime
vectorstore.apply_vector_index(
index_name="my_vi",
m=16,
distance="COSINE",
ef_construction=200,
)
# Drop the vector index
vectorstore.drop_vector_index()
# Rebuild the index to reclaim space and improve recall
vectorstore.optimize()
Index Monitoring
# Get runtime statistics
stats = vectorstore.get_stats()
print(stats) # e.g. {"Vidx_query_count": 100, "Vidx_load_node_hits": 950, "Vidx_load_node_misses": 50, ...}
# Preload HNSW index into memory cache to eliminate cold-start latency
vectorstore.preload_index()
# Check if preloading would fit in cache
check_result = vectorstore.preload_check()
print(check_result)
# Diagnose index health (combines VECTOR_INDEXES view + EXPLAIN)
health = vectorstore.explain_index_health()
print(health)
Delete and Manage Vectors
# Delete by IDs
vectorstore.delete(ids=["id1", "id2"])
# Get documents by IDs
docs = vectorstore.get_by_ids(["id1", "id2"])
# Count vectors
count = vectorstore.count()
# Clear all data (TRUNCATE TABLE)
vectorstore.clear()
# Drop the entire table
vectorstore.drop_table()
# Search documents by metadata only (no vector similarity)
docs = vectorstore.search_by_metadata(filter={"category": "fruit"}, limit=10)
# Delete documents matching metadata conditions
deleted_count = vectorstore.delete_by_metadata(filter={"status": {"$eq": "deleted"}})
Bulk Upsert
# Upsert multiple texts with pre-computed embeddings, metadata, and custom IDs
texts = ["doc1", "doc2", "doc3"]
embeddings = [[0.1, 0.2, ...], [0.3, 0.4, ...], [0.5, 0.6, ...]]
metadatas = [{"src": "web"}, {"src": "pdf"}, {"src": "api"}]
ids = ["a1", "a2", "a3"]
vectorstore.bulk_upsert(
texts=texts,
embeddings=embeddings,
ids=ids,
metadatas=metadatas,
batch_size=100,
)
Async API
All public methods have async equivalents:
import asyncio
async def main():
# Add texts
ids = await vectorstore.aadd_texts(["Hello", "World"])
# Search
results = await vectorstore.asimilarity_search("Hello", k=2)
# MMR search
results = await vectorstore.amax_marginal_relevance_search("Hello", k=4)
# Delete
await vectorstore.adelete(ids=["id1"])
# Get by IDs
docs = await vectorstore.aget_by_ids(["id1", "id2"])
# Count
count = await vectorstore.acount()
# Clear
await vectorstore.aclear()
# Dynamic index management
await vectorstore.aapply_vector_index(index_name="vi", m=16)
await vectorstore.adrop_vector_index()
await vectorstore.aoptimize()
# Index monitoring
stats = await vectorstore.aget_stats()
await vectorstore.apreload_index()
health = await vectorstore.aexplain_index_health()
asyncio.run(main())
Partitioned Table
# Create a partitioned vector table (HASH partitioning, 8 partitions)
# Note: Not available on all PolarDB-X versions
vectorstore = PolarDBXVectorStore(
host="your-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
embedding=embeddings,
table_name="partitioned_vectors",
partition_by="HASH", # or "KEY"
partitions=8,
)
Configuration Options
| Parameter | Type | Default | Description |
|---|---|---|---|
host |
str | - | PolarDB-X host address |
port |
int | - | PolarDB-X port number |
user |
str | - | Username |
password |
str | - | Password |
database |
str | - | Database name |
embedding |
Embeddings | - | LangChain embedding model |
table_name |
str | "polardbx_vectors" |
Table name for vector storage |
distance_strategy |
str | "cosine" |
Distance function: "cosine", "euclidean", or "inner_product" |
hnsw_m |
int | 6 | HNSW index M parameter (3-200) |
pool_size |
int | 5 | Connection pool size |
pre_delete_table |
bool | False | Drop table before creating |
embedding_dimension |
int | None | Embedding dimension (auto-inferred if not provided) |
ef_construction |
int | None | HNSW build-time candidate list size (5-1000) |
connection_retries |
int | 3 | Number of connection retry attempts |
retry_delay |
float | 1.0 | Delay between retries in seconds |
vector_index_name |
str | None | Vector index name for FORCE INDEX hints (auto-detected if None) |
partition_by |
str | None | Partition strategy: "HASH" or "KEY" |
partitions |
int | 0 | Number of partitions (only effective with partition_by) |
**kwargs |
- | - | Additional connection arguments (e.g. ssl_ca, ssl_cert, ssl_key, ssl_disabled) |
PolarDB-X Vector Functions Used
This integration uses PolarDB-X's native vector functions:
VECTOR(N)— Vector column data type with N dimensionsVEC_FROMTEXT('[1,2,3]')— Convert JSON array string to vectorVEC_TOTEXT(vector)— Convert vector to JSON array stringVEC_DISTANCE(v1, v2)— Auto-inferred distance functionVEC_DISTANCE_COSINE(v1, v2)— Cosine distanceVEC_DISTANCE_EUCLIDEAN(v1, v2)— Euclidean distanceVEC_DISTANCE_INNER_PRODUCT(v1, v2)— Inner product distanceVECTOR_DIM(v)— Get vector dimensionVECTOR INDEX (col) M=N DISTANCE=COSINE— HNSW vector index DDLEF_CONSTRUCTION=N— HNSW build-time parameter in DDLSET SESSION vidx_hnsw_ef_search = N— Per-session search width tuningSHOW GLOBAL STATUS LIKE 'Vidx%'— Runtime index statisticsCALL dbms_vidx.preload(db, table, col)— Preload index into cacheCALL dbms_vidx.preload_check(db, table, col)— Check preload feasibilityinformation_schema.VECTOR_INDEXES— Vector index metadata view
Development
This package uses uv for dependency management.
# Install dependencies
uv sync --group test --group test_integration
# Run unit tests
make test
# Run integration tests (requires a running PolarDB-X instance)
make integration_tests
# Lint
make lint
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file langchain_polardbx-0.1.1.tar.gz.
File metadata
- Download URL: langchain_polardbx-0.1.1.tar.gz
- Upload date:
- Size: 259.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6b53427fffcb6df248ed1629c495fc8a969f11bc84c00c03c01b2360baabbefe
|
|
| MD5 |
af37f957ea8fadf4363432dd970c35fc
|
|
| BLAKE2b-256 |
a4f856438c81628d93874bbe9b5a1ea2e856f1de7c950a04165194a56c47f6b8
|
File details
Details for the file langchain_polardbx-0.1.1-py3-none-any.whl.
File metadata
- Download URL: langchain_polardbx-0.1.1-py3-none-any.whl
- Upload date:
- Size: 26.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b9498240a25808c9c6a97c9cc74724ffc534e73451a1f33b212408a0319e227c
|
|
| MD5 |
4c92bbb41762cb41e072d7446256ed29
|
|
| BLAKE2b-256 |
fa5bf637742be67e48c5b0cd99425ce575bb29bff07a6c2e0e2318e98544d3f6
|