🦙 LlamaIndex Vector Stores PolarDB-X
A powerful integration between LlamaIndex and PolarDB-X, enabling native vector search capabilities for AI applications.
Overview
LlamaIndex PolarDB-X provides seamless integration between LlamaIndex, a framework for building context-augmented LLM applications, and PolarDB-X with native vector search support. This integration enables efficient vector storage and retrieval for AI applications like semantic search, recommendation systems, and RAG (Retrieval Augmented Generation).
PolarDB-X is a cloud-native distributed database system developed by Alibaba Cloud, featuring native HNSW-based vector index support that delivers high-performance approximate nearest neighbor (ANN) search directly within the database engine.
Requirements
- Python 3.10+
- PolarDB-X with vector index support
- LlamaIndex core:
llama-index-core>=0.13.0,<0.15(included in package dependencies) - SQLAlchemy:
sqlalchemy>=2.0.0(included in package dependencies) - Async support:
aiomysql>=0.2.0(included in package dependencies) - MySQL driver:
pymysql>=1.0.0(included in package dependencies)
Enable Vector Index
PolarDB-X disables the vector index feature by default (vidx_disabled = ON). You need to enable it before using this package:
-- Enable vector index (run as admin/root on DN node)
SET GLOBAL vidx_disabled = OFF;
This setting takes effect immediately for new connections. No restart required.
All transaction isolation levels (READ-COMMITTED, REPEATABLE-READ, SERIALIZABLE) are supported — choose according to your business needs.
Features
- Native Vector Storage: Store embeddings using PolarDB-X's native
VECTOR(N)data type - HNSW Index: Efficient approximate nearest neighbor search with configurable
MandEF_CONSTRUCTIONparameters - Multiple Distance Metrics: Support for Cosine, Euclidean, and Inner Product distance (v3)
- Similarity Search: Perform efficient similarity searches with configurable top-k
- Metadata Filtering: Filter search results by metadata with rich operators (
$eq,$ne,$gt,$gte,$lt,$lte,$in,$nin) - Dynamic Index Management: Create, drop, and rebuild vector indexes at runtime without recreating tables
- Search Mode Control: Switch between ANN (index-accelerated) and KNN (full-scan) modes per query
- Per-Query Tuning: Adjust
ef_searchon a per-query basis for accuracy/latency trade-offs - Index Health Monitoring: Runtime statistics, index health diagnostics, and preload checks (v3)
- Batch Operations: Batch insert with UPSERT support within a single transaction
- Full Async Support: All public methods have async equivalents (
async_add,aquery, etc.) - Dual-Version Compatibility: Automatically detects database capabilities and adapts SQL accordingly
- Connection Pooling: Built-in SQLAlchemy Engine with connection pooling
Installation
pip install -U llama-index-vector-stores-polardbx
Optional Dependencies
For using OpenAI embeddings:
pip install llama-index-embeddings-openai
For using DashScope embeddings (Alibaba Cloud):
pip install llama-index-embeddings-dashscope
Quick Start
Basic Usage
from llama_index.vector_stores.polardbx import PolarDBXVectorStore
from llama_index.core import StorageContext, VectorStoreIndex, Settings
from llama_index.embeddings.openai import OpenAIEmbedding
# Configure embedding model
Settings.embed_model = OpenAIEmbedding()
# Create vector store
vector_store = PolarDBXVectorStore(
host="your-polardbx-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
table_name="my_vectors",
embed_dim=1536,
distance_method="COSINE", # or "EUCLIDEAN", "INNER_PRODUCT" (v3 only)
)
# Create index from documents
storage_context = StorageContext.from_defaults(vector_store=vector_store)
index = VectorStoreIndex.from_documents(
documents,
storage_context=storage_context,
)
# Query
query_engine = index.as_query_engine()
response = query_engine.query("What is PolarDB-X?")
print(response)
Using from_params Factory Method
from llama_index.vector_stores.polardbx import PolarDBXVectorStore
vector_store = PolarDBXVectorStore.from_params(
host="your-polardbx-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
table_name="my_vectors",
embed_dim=1536,
distance_method="COSINE",
ssl=True, # Enable TLS
ssl_ca="/path/to/ca.pem", # CA certificate
)
Using DashScope Embeddings
from llama_index.vector_stores.polardbx import PolarDBXVectorStore
from llama_index.embeddings.dashscope import DashScopeEmbedding
embed_model = DashScopeEmbedding(
model_name="text-embedding-v4",
api_key="your-dashscope-api-key",
)
vector_store = PolarDBXVectorStore(
host="your-polardbx-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
table_name="my_vectors",
embed_dim=1024,
)
Usage Examples
Direct Vector Store Operations
from llama_index.vector_stores.polardbx import PolarDBXVectorStore
from llama_index.core.schema import TextNode
vector_store = PolarDBXVectorStore(
host="your-host",
port=3306,
user="your-user",
password="your-password",
database="your-database",
table_name="my_vectors",
embed_dim=1536,
)
# Add nodes
nodes = [
TextNode(text="Hello world", embedding=[0.1, 0.2, ...]),
TextNode(text="PolarDB-X is great", embedding=[0.3, 0.4, ...]),
]
ids = vector_store.add(nodes)
# Query
from llama_index.core.vector_stores import VectorStoreQuery
query = VectorStoreQuery(
query_embedding=[0.1, 0.2, ...],
similarity_top_k=5,
)
result = vector_store.query(query)
for node, score in zip(result.nodes, result.similarities):
print(f"[Score: {score:.4f}] {node.text}")
Search with Metadata Filter
from llama_index.core.vector_stores import (
MetadataFilter,
MetadataFilters,
FilterOperator,
)
# Add nodes with metadata
nodes = [
TextNode(text="Apple is a fruit", embedding=[...], metadata={"category": "fruit", "price": 5}),
TextNode(text="Banana is yellow", embedding=[...], metadata={"category": "fruit", "price": 3}),
TextNode(text="Car is a vehicle", embedding=[...], metadata={"category": "vehicle", "price": 20000}),
]
vector_store.add(nodes)
# Filter: category = "fruit" AND price > 2
filters = MetadataFilters(
filters=[
MetadataFilter(key="category", value="fruit", operator=FilterOperator.EQ),
MetadataFilter(key="price", value=2, operator=FilterOperator.GT),
]
)
query = VectorStoreQuery(query_embedding=[...], similarity_top_k=5, filters=filters)
result = vector_store.query(query)
Search Mode Control
# Force ANN (use vector index for HNSW acceleration)
result = vector_store.query(query, search_type="ann")
# Force KNN (full table scan, bypass vector index)
result = vector_store.query(query, search_type="knn")
# Let the optimizer decide (default)
result = vector_store.query(query, search_type="auto")
# Tune ef_search per query (higher = more accurate, slower)
result = vector_store.query(query, ef_search=100)
Dynamic Vector Index Management
# Create a vector index at runtime
vector_store.apply_vector_index(
index_name="my_vi",
m=16,
distance="COSINE",
ef_construction=200, # v3 only, ignored on old versions
)
# Drop the vector index
vector_store.drop_vector_index()
# Rebuild the index to reclaim space and improve recall
vector_store.optimize()
Index Monitoring (v3 only)
# Get runtime statistics
stats = vector_store.get_stats()
print(stats) # e.g. {"Vidx_query_count": 100, "Vidx_load_node_hits": 950, ...}
# Preload HNSW index into memory cache to eliminate cold-start latency
vector_store.preload_index()
# Check if preloading would fit in cache
check_result = vector_store.preload_check()
print(check_result)
# Diagnose index health
health = vector_store.explain_index_health()
print(health)
Delete and Manage Nodes
# Delete by ref_doc_id
vector_store.delete(ref_doc_id="doc-001")
# Delete by node_ids
vector_store.delete_nodes(node_ids=["node-1", "node-2"])
# Get nodes by node_ids
nodes = vector_store.get_nodes(node_ids=["node-1", "node-2"])
# Count vectors
count = vector_store.count()
# Clear all data (TRUNCATE TABLE)
vector_store.clear()
# Drop the entire table
vector_store.drop()
Async API
All public methods have async equivalents:
import asyncio
async def main():
# Add nodes
ids = await vector_store.async_add(nodes)
# Query
result = await vector_store.aquery(query)
# Delete
await vector_store.adelete(ref_doc_id="doc-001")
# Delete nodes
await vector_store.adelete_nodes(node_ids=["node-1"])
# Count
count = await vector_store.acount()
# Clear
await vector_store.aclear()
# Dynamic index management
await vector_store.aapply_vector_index(index_name="vi", m=16)
await vector_store.adrop_vector_index()
await vector_store.aoptimize()
# v3 monitoring
stats = await vector_store.aget_stats()
await vector_store.apreload_index()
check = await vector_store.apreload_check()
health = await vector_store.aexplain_index_health()
asyncio.run(main())
Configuration Options
| Parameter | Type | Default | Description |
|---|---|---|---|
host |
str | - | PolarDB-X host address |
port |
int | - | PolarDB-X port number |
user |
str | - | Username |
password |
str | - | Password |
database |
str | - | Database name |
table_name |
str | "llama_index_table" |
Table name for vector storage |
embed_dim |
int | 1536 | Embedding dimension |
distance_method |
str | "COSINE" |
Distance function: "COSINE", "EUCLIDEAN", or "INNER_PRODUCT" (v3) |
default_m |
int | 6 | HNSW index M parameter (DB allows 3-200; client validates positive int only) |
perform_setup |
bool | True | Whether to auto-create table on init |
debug |
bool | False | Enable SQLAlchemy echo mode |
ef_construction |
int | None | HNSW build-time candidate list size (DB allows 5-1000, v3 only; client validates positive int only) |
ssl |
bool | False | Enable TLS/SSL encryption |
ssl_ca |
str | None | Path to CA certificate for SSL verification (only effective when ssl=True) |
vector_index_name |
str | None | Vector index name for FORCE INDEX hints (auto-detected if None) |
PolarDB-X Vector Functions Used
This integration uses PolarDB-X's native vector functions:
VECTOR(N)— Vector column data type with N dimensionsVEC_FROMTEXT('[1,2,3]')— Convert JSON array string to vectorVEC_TOTEXT(vector)— Convert vector to JSON array stringVEC_DISTANCE(v1, v2)— Auto-inferred distance function (v3)VEC_DISTANCE_COSINE(v1, v2)— Cosine distance (old versions)VEC_DISTANCE_EUCLIDEAN(v1, v2)— Euclidean distance (old versions)VEC_DISTANCE_INNER_PRODUCT(v1, v2)— Inner product distance (used when v3 auto-inference unavailable; INNER_PRODUCT distance itself requires v3)VECTOR_DIM(v)— Get vector dimension (v3)VECTOR INDEX (col) M=N DISTANCE=COSINE— HNSW vector index DDLEF_CONSTRUCTION=N— HNSW build-time parameter in DDL (v3)SET SESSION vidx_hnsw_ef_search = N— Per-session search width tuningSHOW GLOBAL STATUS LIKE 'Vidx%'— Runtime index statisticsCALL dbms_vidx.preload(db, table, col)— Preload index into cache (v3)CALL dbms_vidx.preload_check(db, table, col)— Check preload feasibility (v3)information_schema.VECTOR_INDEXES— Vector index metadata view (v3)
Error Handling
When using features that require PolarDB-X v3 (e.g. INNER_PRODUCT distance, preload_index, explain_index_health), a NotSupportedError is raised on older versions:
from llama_index.vector_stores.polardbx import PolarDBXVectorStore, NotSupportedError
try:
vector_store.preload_index()
except NotSupportedError as e:
print(f"Feature not supported: {e}")
Development
This package uses uv for dependency management.
# Install dependencies
uv sync --group dev
# Run tests
pytest -v
# Lint
ruff check .
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llama_index_vector_stores_polardbx-0.1.0.tar.gz.
File metadata
- Download URL: llama_index_vector_stores_polardbx-0.1.0.tar.gz
- Upload date:
- Size: 19.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
518bafb09e095dddd7858d61c65271810636c810d152a943ab2be00cfa56f5ed
|
|
| MD5 |
2dafb6bc0d77929442961dce51a499da
|
|
| BLAKE2b-256 |
de26f362ae12548209b0076a5adb56b9f410f818b87b3915d2e9f392e1800b18
|
File details
Details for the file llama_index_vector_stores_polardbx-0.1.0-py3-none-any.whl.
File metadata
- Download URL: llama_index_vector_stores_polardbx-0.1.0-py3-none-any.whl
- Upload date:
- Size: 20.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
16829b934bdd68440d65c9aafa01944e09af616e43728c3cb88142b0f5f45eee
|
|
| MD5 |
a5d89dae8cabad40b15034025fc73458
|
|
| BLAKE2b-256 |
daa981f9cb8e52510becb4bf03a1a970d7c7156c65a84853b5904fdf520cb790
|