langchain-vastdb
LangChain VectorStore integration for VAST Database.
langchain-vastdb provides a VastDBVectorStore class that implements the
LangChain VectorStore interface, enabling similarity search, document storage,
and retrieval-augmented generation (RAG) workflows backed by VAST Database's
native vector indexing.
Compatibility: Python 3.10 - 3.13 | langchain-core >= 1.0, < 2 | vastdb >= 2.0.3 | Query Engine paths require VAST 5.4+
Status: Alpha (v0.0.1). API may change between minor releases.
License: Apache-2.0
Requirements
-
Python 3.10+
-
A running VAST Database cluster. The store reads the cluster version from the SDK session and picks paths per operation:
VAST search get_by_idsdelete notes 5.3 SDK in-memory scan ( VASTDB_ALLOW_FALLBACK=1required)SDK SDK no Query Engine 5.4 Query Engine, brute force Query Engine SDK not live-tested; Query Engine DML unverified 5.5 Query Engine, vector index when built Query Engine Query Engine verified on 5.5.1 unknown Query Engine Query Engine SDK session reports no version; one warning per store Configuring ADBC on a 5.3 cluster is harmless: lookup and delete stay on the SDK.
-
vastdbSDK >= 2.0.3 -
langchain-core>= 1.0, < 2 -
An
Embeddingsmodel (e.g., OpenAI, HuggingFace, or any LangChain-compatible embeddings)
Installation
pip install langchain-vastdb
Or with uv:
uv add langchain-vastdb
Quickstart
Option 1: Pass a pre-built session
import vastdb
from langchain_vastdb import VastDBVectorStore
session = vastdb.connect(
endpoint="http://vast-cluster:8070",
access="YOUR_ACCESS_KEY",
secret="YOUR_SECRET_KEY",
)
store = VastDBVectorStore(
embedding=my_embeddings,
session=session,
bucket="my-bucket",
schema="my-schema",
table_name="my-table",
)
# Add documents and search
ids = store.add_texts(["Paris is the capital of France."])
results = store.similarity_search("capital city", k=1)
print(results[0].page_content)
Option 2: Use the convenience factory
from langchain_vastdb import VastDBVectorStore
store = VastDBVectorStore.from_connection_params(
embedding=my_embeddings,
endpoint="http://vast-cluster:8070",
access_key="YOUR_ACCESS_KEY",
secret_key="YOUR_SECRET_KEY",
bucket="my-bucket",
schema="my-schema",
table_name="my-table",
)
Credentials are passed to vastdb.connect() for the SDK session. They are also
kept on the instance as private attributes (_access_key, _secret_key) so the
ADBC Query Engine connection can reuse them; they are never exposed publicly.
Option 3: Create a store and add texts in one call
import vastdb
from langchain_vastdb import VastDBVectorStore
session = vastdb.connect(
endpoint="http://vast-cluster:8070",
access="YOUR_ACCESS_KEY",
secret="YOUR_SECRET_KEY",
)
store = VastDBVectorStore.from_texts(
texts=["Paris is the capital of France.", "Berlin is the capital of Germany."],
embedding=my_embeddings,
session=session,
bucket="my-bucket",
schema="my-schema",
table_name="my-table",
)
CRUD Operations
# Add documents with metadata
ids = store.add_texts(
["Some text", "More text"],
metadatas=[{"source": "wiki"}, {"source": "blog"}],
)
# Similarity search by text query
docs = store.similarity_search("capital city", k=2)
# Similarity search with distance scores
scored = store.similarity_search_with_score("capital city", k=2)
for doc, score in scored:
print(f"{doc.page_content} (distance: {score})")
# Search with a pre-computed vector
docs = store.similarity_search_by_vector([0.1, 0.2, ...], k=2)
# Retrieve documents by ID
docs = store.get_by_ids(ids)
# Delete by ID
store.delete(ids=ids)
Using as a retriever
VastDBVectorStore integrates directly with LangChain's retriever interface:
retriever = store.as_retriever(search_kwargs={"k": 3})
docs = retriever.invoke("What is the capital of France?")
This works seamlessly in LCEL RAG chains:
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
retriever = store.as_retriever(search_kwargs={"k": 3})
prompt = ChatPromptTemplate.from_template(
"Answer based on context:\n{context}\n\nQuestion: {question}"
)
def format_docs(docs):
return "\n".join(d.page_content for d in docs)
chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm # any LangChain-compatible LLM
| StrOutputParser()
)
answer = chain.invoke("What is the capital of France?")
Cache management
VastDBVectorStore caches table metadata after the first access to avoid
repeated bucket/schema/table round trips. If you alter the table structure
externally, invalidate the cache:
store.invalidate_table_cache()
Configuration Reference
Constructor: VastDBVectorStore(...)
| Parameter | Type | Default | Description |
|---|---|---|---|
embedding |
Embeddings |
required | The embeddings model used to generate vectors. |
session |
vastdb.Session |
required | A pre-built session connected to the VAST cluster. |
bucket |
str |
required | The VAST bucket name containing the target table. |
schema |
str |
required | The schema name within the bucket. |
table_name |
str |
required | The table name for vector operations. |
id_column |
str |
"id" |
Column name for document IDs. |
text_column |
str |
"text" |
Column name for document text. |
vector_column |
str |
"vector" |
Column name for embedding vectors. |
metadata_column |
str |
"metadata" |
Column name for document metadata (stored as JSON). |
adbc_driver_path |
str | None |
None |
Path to libadbc_driver_vastdb.so. Enables Query Engine search, lookup and delete. |
adbc_endpoint |
str | None |
None |
Full Query Engine URL (e.g. http://host:80), separate from the SDK endpoint. |
access_key |
str | None |
None |
Access key for ADBC connection. |
secret_key |
str | None |
None |
Secret key for ADBC connection. |
Custom column names
Column names default to id, text, vector, and metadata. Override them at
construction time:
store = VastDBVectorStore(
embedding=my_embeddings,
session=session,
bucket="my-bucket",
schema="my-schema",
table_name="my-table",
id_column="doc_id",
text_column="content",
vector_column="emb",
metadata_column="meta",
)
Factory classmethod: from_connection_params(...)
Creates a VastDBVectorStore by building a vastdb.Session internally from
connection parameters.
| Parameter | Type | Default | Description |
|---|---|---|---|
embedding |
Embeddings |
required | The embeddings model. |
endpoint |
str |
required | The VAST cluster HTTP endpoint URL. |
access_key |
str |
required | Access key for authentication. |
secret_key |
str |
required | Secret key for authentication. |
bucket |
str |
required | The VAST bucket name. |
schema |
str |
required | The schema name within the bucket. |
table_name |
str |
required | The table name for vector operations. |
adbc_driver_path |
str | None |
None |
Path to ADBC driver shared library. |
adbc_endpoint |
str | None |
None |
ADBC/QueryEngine endpoint. |
**kwargs |
Additional keyword arguments forwarded to the constructor (e.g., custom column names). |
ADBC Query Engine operations
With the ADBC driver, endpoint and credentials configured, vector search fetches
ranked documents in one Query Engine SQL query. get_by_ids and delete also
use the Query Engine; upsert joins the SQL delete and SDK Arrow insert in one
transaction. Insertion remains SDK Arrow. A vector index is optional: without
one (or on VAST 5.4) the Query Engine brute-forces the distance; with one, the
distance function comes from the index metadata (array_distance for l2sq,
array_inner_product for ip). A freshly created indexed table may brute-force
until the index is built. See the version table under Requirements for which
operations use the Query Engine on 5.3, 5.4 and 5.5. Without ADBC, lookup and
delete retain their SDK paths. With VASTDB_ALLOW_FALLBACK=1, search, lookup
and delete all fall back to the SDK on ADBC errors; otherwise errors propagate.
Search and lookup without a transaction or per-call overrides reuse one
autocommit connection per store per thread, reconnecting after an error; calls
given tx and delete use a dedicated connection joined to the transaction.
Unfiltered count() uses cached table stats, which may lag recent writes or
over-count while a table settles. Use count(predicate) for an exact count;
it scans matching IDs rather than using Query Engine COUNT(*).
store = VastDBVectorStore(
embedding=my_embeddings,
session=session,
bucket="my-bucket",
schema="my-schema",
table_name="my-table",
adbc_driver_path="/usr/lib/libadbc_driver_vastdb.so",
adbc_endpoint="http://query-engine.example.com:80",
access_key="YOUR_ACCESS_KEY",
secret_key="YOUR_SECRET_KEY",
)
Subclassing Guide
VastDBVectorStore uses the Template Method pattern. Public methods like
add_texts and similarity_search handle embedding, filter conversion, and
result formatting, then delegate storage operations to five protected hook
methods. Override these hooks to customize behavior without reimplementing the
full LangChain interface.
Hook methods
| Hook | Purpose | Returns |
|---|---|---|
_insert_vectors |
Customize record insertion | list[str] (IDs) |
_build_metadata_columns |
Customize column layout for metadata | dict[str, list] |
_select_columns / _typed_metadata_columns |
Customize document columns projected during search and lookup | list[str] / mapping |
_vector_search |
Customize similarity search | list[tuple[dict, float]] |
_delete_by_ids |
Customize document deletion | bool |
_get_by_ids |
Customize ID lookup (not used by ADBC search) | list[dict] |
_row_to_document |
Customize row-to-Document conversion | Document |
Hook signatures
def _insert_vectors(
self,
texts: list[str],
embeddings: list[list[float]],
metadatas: list[dict],
ids: list[str],
*,
tx: Transaction | None = None,
) -> list[str]: ...
def _vector_search(
self,
query_vector: list[float],
k: int,
predicate: ibis.Expr | None = None,
*,
tx: Transaction | None = None,
**kwargs: Any,
) -> list[tuple[dict, float]]: ...
def _delete_by_ids(
self,
ids: list[str],
*,
tx: Transaction | None = None,
) -> bool: ...
def _get_by_ids(
self,
ids: list[str],
*,
tx: Transaction | None = None,
) -> list[dict]: ...
def _row_to_document(
self,
row: dict,
score: float | None = None,
) -> Document: ...
Override _open_adbc_connection(self, **kwargs) with **kwargs even if your
subclass currently ignores them: joined operations pass
adbc_conn_kwargs_overrides={"vast.db.external_txid": str(tx.active_txid)}.
The base implementation caches a connection when no overrides are passed; an
override that opens a fresh connection per call bypasses that cache.
Overrides of _do_vector_search and _do_vector_search_adbc receive
tx=None when the caller passes no transaction.
Transaction reuse
Write hooks open a transaction by default. The optional tx parameter lets
subclasses pass in an existing transaction for multi-step atomic operations
(ADBC delete joins it):
with self._session.transaction() as tx:
self._insert_vectors(texts, embeddings, metadatas, ids, tx=tx)
# additional operations in the same transaction
Example: typed metadata columns
The base class stores metadata as a single JSON string column. If you need typed
columns for performance-critical filtering, set _typed_metadata_columns:
from langchain_vastdb import TypedColumn, VastDBVectorStore
class TypedMetadataStore(VastDBVectorStore):
"""Store with typed 'category' and 'priority' metadata columns."""
_typed_metadata_columns = {
"category": TypedColumn(),
"priority": TypedColumn(),
}
This automatically extracts category and priority into separate typed columns
on insert, preserves any extra metadata in the JSON column, and merges everything
back together on read. The public LangChain interface (add_texts,
similarity_search, etc.) stays unchanged.
Use TypedColumn fields for custom defaults, PyArrow type coercion, or
controlling which columns are backfilled on read
(see the Migration Guide for details).
Examples
See the examples/ directory for runnable scripts:
basic_usage.py-- add texts, search, retrieverag_pipeline.py--as_retriever()+ LCEL RAG chainsubclassing.py-- declarative typed metadata columnsfiltered_search.py-- metadata filtering patterns
Migration Guide
Migrating an existing VectorStore subclass to VastDBVectorStore? See the
Migration Guide for step-by-step instructions,
a hook mapping table, and a before/after code comparison.
Development
Clone the repository and install dependencies with uv:
uv sync
Run the linter:
uv run ruff check .
Run unit tests:
uv run pytest tests/unit_tests/
Run integration tests (requires a VAST cluster):
uv run pytest tests/integration_tests/
License
Apache-2.0 -- see LICENSE for details. test sync
Metadata
Release files for langchain-vastdb 0.0.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_vastdb-0.0.6.tar.gz | 170.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_vastdb-0.0.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 200.7 kB
Release files / langchain_vastdb-0.0.6.tar.gz
| Download URL | langchain_vastdb-0.0.6.tar.gz |
|---|---|
| Size | 170.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0bdf1c0ff95cb3b0f543d2076301dd44ab8b8a090107c2dd3db43a1ce72b4f13
|
|
BLAKE2b-256 checksum How to use checksums |
c6c4ef9bb2deb61547bd5cb8447dd830d2cf63453edee5ed22d627bbc7d16e7c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / langchain_vastdb-0.0.6-py3-none-any.whl
| Download URL | langchain_vastdb-0.0.6-py3-none-any.whl |
|---|---|
| Size | 30.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
edce254044081978ccb8ff7bb9f172b19f1c61497fc789f900d83fe61a7845f0
|
|
BLAKE2b-256 checksum How to use checksums |
ec22ae82abf5e0bff38aeb59d143a5990784682b22e2cf1ff926194b96cca810
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|