Vespa Plugin for Search Toolkit
Vespa integration plugin for mistralai-search-toolkit.
This plugin provides a production-ready Vespa search backend implementation for the Search Toolkit, enabling powerful vector, keyword, and hybrid search capabilities.
Installation
pip install mistralai-search-toolkit-plugins-vespa
Or as an optional dependency of the core package:
pip install mistralai-search-toolkit[vespa]
Quick Start
1. Bootstrap Your Application
Create the application structure with an initial migration:
uv run mistral-vespa generate-migration --app-dir ./vespa_app initial_schema
This creates the ./vespa_app/ directory and generates a migration file. Fill it with your schema definition:
from mistralai.search.toolkit.plugins.vespa.app.schemas.app import SearchMode
from mistralai.search.toolkit.plugins.vespa.migration import VespaMigration, create_default_schema, set_app_name
class InitialSchema(VespaMigration):
def migrate(self) -> None:
set_app_name("articles")
create_default_schema(
name="articles",
mode=SearchMode.INDEX,
embedding_dimensions=1024, # Adjust based on your embedder
schema_version=1,
)
2. Start a Local Vespa Instance
uv run mistral-vespa local up --query-port 18080 --config-port 19171 --name vespa-dev
3. Deploy Your Application
Deploy the migrations to generate the vespa_app module:
uv run mistral-vespa migrate \
--app-dir ./vespa_app \
--config-server http://localhost:19171 \
--query-port 18080
This generates the vespa_app Python module that you can now import.
4. Index Documents
import os
from mistralai.search.toolkit.ingestion.pipelines import Pipeline
from mistralai.search.toolkit.ingestion.loaders import FilesystemFileLoader
from mistralai.search.toolkit.ingestion.text_splitters import CharacterTextSplitter
from mistralai.search.toolkit.embedding import MistralEmbedder, MODEL_1024_EMBEDDING
from mistralai.client import Mistral
from mistralai.search.toolkit.plugins.vespa import VespaClientConfig
from vespa_app import app
# Setup
mistral_client = Mistral(api_key=os.environ.get("MISTRAL_API_KEY"))
vespa_config = VespaClientConfig(
endpoint=os.environ.get("VESPA_ENDPOINT", "http://localhost:18080"),
)
collection_name = "articles"
# Connect to Vespa
vector_store = app.get_search_index(vespa_config, collection_name=collection_name)
# Index documents
pipeline = Pipeline(
loader=FilesystemFileLoader(),
text_splitter=CharacterTextSplitter(chunk_size=512),
embedder=MistralEmbedder(client=mistral_client, model_name=MODEL_1024_EMBEDDING),
stores=vector_store,
)
num_chunks = await pipeline.run(documents=["doc1.pdf", "doc2.pdf"])
4. Search
from mistralai.search.toolkit.embedding import MistralEmbedder, MODEL_1024_EMBEDDING
from mistralai.search.toolkit.retrieval import QueryEngine
from mistralai.search.toolkit.retrieval.retrievers import VectorRetriever
# Setup search
embedder = MistralEmbedder(client=mistral_client, model_name=MODEL_1024_EMBEDDING)
query_engine = QueryEngine(
retriever=[VectorRetriever(client=vector_store, embedder=embedder)],
)
# Search documents
results = await query_engine.search(query="What is machine learning?", top_k=10)
# Display results
for result in results.results:
print(f"Score: {result.score}")
print(f"Content: {result.content}\n")
Configuration
Quick Setup
Use app.get_search_index() for the common case where a single endpoint serves both query and feed APIs:
import os
from mistralai.search.toolkit.plugins.vespa import VespaClientConfig
from vespa_app import app
vespa_config = VespaClientConfig(
endpoint=os.environ.get("VESPA_ENDPOINT", "http://localhost:18080"),
)
vector_store = app.get_search_index(vespa_config, collection_name="articles")
Advanced Setup
Use separate query and feed endpoints for production deployments:
from mistralai.search.toolkit.plugins.vespa import VespaClientConfig
from vespa_app import app
client_config = VespaClientConfig(
query_endpoint=os.environ.get("VESPA_QUERY_ENDPOINT", "https://query.vespa.example.com"),
feed_endpoint=os.environ.get("VESPA_FEED_ENDPOINT", "https://feed.vespa.example.com"),
)
vector_store = app.get_search_index(
client_config=client_config,
collection_name="articles",
)
License
This plugin is licensed under the Apache License 2.0.
Support
For issues related to the Search Toolkit, refer to the Search Toolkit documentation.
For Vespa-specific questions, visit Vespa documentation.
Metadata
Release files for mistralai-search-toolkit-plugins-vespa 0.0.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mistralai_search_toolkit_plugins_vespa-0.0.13.tar.gz | 203.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mistralai_search_toolkit_plugins_vespa-0.0.13-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 334.1 kB
Release files / mistralai_search_toolkit_plugins_vespa-0.0.13.tar.gz
| Download URL | mistralai_search_toolkit_plugins_vespa-0.0.13.tar.gz |
|---|---|
| Size | 203.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
198b0c2cdf2353f398c05d55a92968ea7036d703c9ba44613c8838496f7fab05
|
|
BLAKE2b-256 checksum How to use checksums |
0bee9c0e421537b2205cd5643517114df546b88fa83032715db6f28a34238af4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|
Release files / mistralai_search_toolkit_plugins_vespa-0.0.13-py3-none-any.whl
| Download URL | mistralai_search_toolkit_plugins_vespa-0.0.13-py3-none-any.whl |
|---|---|
| Size | 131.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
45477b3a6e30ac08f87433654c082f4215a2491ffaf40de46d22b22e19b4fdac
|
|
BLAKE2b-256 checksum How to use checksums |
2020ed7fcf75ed80afc205ca63528757c40392a6f4ed3f360492bc7a1f21ce2c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|