Skip to main content

Search Toolkit

Modular, backend-agnostic framework for building and evaluating Information Retrieval systems.

Overview

Search Toolkit provides plug-and-play, extensible components for building production-ready IR pipelines. Every component is swappable and customizable — build exactly what your use case needs.

What's Included

Core Components

  • Ingestion: Document loaders, extractors, text splitters, enrichment, and indexing pipelines
  • Retrieval: Vector (semantic), keyword (BM25), and hybrid search with RRF fusion
  • Query Processing: LLM reformulation and custom preprocessing
  • Reranking: Rerank results to surface the most relevant information
  • Embedders: Generate embeddings for documents and queries using Mistral's embedding models
  • Storage: Abstract object storage interface for document persistence

Backend Agnostic

The toolkit is designed to work with different search backends through plugins. You can use it with any vector database or search engine by installing the appropriate plugin.

Installation

Install the base package:

pip install mistralai-search-toolkit

Install optional components:

# Text extraction from PDFs (requires pymupdf-pro)
pip install mistralai-search-toolkit[extractor-pymupdf]

# HTML to markdown conversion
pip install mistralai-search-toolkit[html-converter-markdownify]

# Email extraction
pip install mistralai-search-toolkit[extractor-email]

# Spreadsheet parsing
pip install mistralai-search-toolkit[extractor-spreadsheet]

# LangChain text splitting
pip install mistralai-search-toolkit[text-splitter-langchain]

Quick Start

1. Load and Process Documents

import os
from mistralai.search.toolkit.ingestion.loaders import FilesystemFileLoader
from mistralai.search.toolkit.ingestion.text_splitters import CharacterTextSplitter
from mistralai.client import Mistral

# Load documents from a directory
loader = FilesystemFileLoader()
documents = loader.load(path="/path/to/documents")

# Split into chunks
splitter = CharacterTextSplitter(chunk_size=512)
chunks = splitter.split(documents)

2. Generate Embeddings

from mistralai.search.toolkit.embedding import MistralEmbedder, MODEL_1024_EMBEDDING

# Create embedder (uses Mistral's API)
mistral_client = Mistral(api_key=os.environ.get("MISTRAL_API_KEY", "your-api-key"))
embedder = MistralEmbedder(client=mistral_client, model_name=MODEL_1024_EMBEDDING)

# Embed your chunks
embedded_chunks = embedder.embed(chunks)

The toolkit supports multiple search backends through plugins. See the Vespa Plugin section below for a complete example.

Vespa Plugin: Creating a Search Index

Vespa is an open-source search engine that integrates seamlessly with the toolkit.

Prerequisites

  • The Vespa plugin: pip install mistralai-search-toolkit-plugins-vespa
  • Docker for local development

Getting Started with Vespa

Step 1: Bootstrap Your Vespa Application

First, create the application structure with an initial migration:

uv run mistral-vespa generate-migration --app-dir ./vespa_app initial_schema

This creates ./vespa_app/ and generates a migration file. Fill it with your schema definition:

from mistralai.search.toolkit.plugins.vespa.app.schemas.app import SearchMode
from mistralai.search.toolkit.plugins.vespa.migration import VespaMigration, create_default_schema, set_app_name


class InitialSchema(VespaMigration):
    def migrate(self) -> None:
        set_app_name("articles")
        create_default_schema(
            name="articles",
            mode=SearchMode.INDEX,
            embedding_dimensions=1024,  # Match your embedder's dimensions
            schema_version=1,
        )

Step 2: Start a Local Vespa Instance

uv run mistral-vespa local up --query-port 18080 --config-port 19171 --name vespa-dev

Step 3: Deploy Your Application

Deploy the migrations to the running Vespa instance:

uv run mistral-vespa migrate \
  --app-dir ./vespa_app \
  --config-server http://localhost:19171 \
  --query-port 18080

This generates the vespa_app module that you can import.

Step 4: Ingest and Search Documents

After deployment, use the generated vespa_app to index and search:

import os
from mistralai.search.toolkit.ingestion.pipelines import Pipeline
from mistralai.search.toolkit.ingestion.loaders import FilesystemFileLoader
from mistralai.search.toolkit.ingestion.text_splitters import CharacterTextSplitter
from mistralai.search.toolkit.embedding import MistralEmbedder, MODEL_1024_EMBEDDING
from mistralai.client import Mistral
from mistralai.search.toolkit.plugins.vespa import VespaClientConfig
from mistralai.search.toolkit.retrieval import QueryEngine
from mistralai.search.toolkit.retrieval.retrievers import VectorRetriever
from vespa_app import app  # Generated by migration deployment

# Configuration
mistral_client = Mistral(api_key=os.environ.get("MISTRAL_API_KEY", "your-api-key"))
vespa_config = VespaClientConfig(
    endpoint=os.environ.get("VESPA_ENDPOINT", "http://localhost:18080"),
)
collection_name = "articles"

# Connect to Vespa
vector_store = app.get_search_index(vespa_config, collection_name=collection_name)

# INGESTION: Index your documents
pipeline = Pipeline(
    loader=FilesystemFileLoader(),
    text_splitter=CharacterTextSplitter(chunk_size=512),
    embedder=MistralEmbedder(client=mistral_client, model_name=MODEL_1024_EMBEDDING),
    stores=vector_store,
)

num_chunks = await pipeline.run(documents=["doc1.pdf", "doc2.pdf"])

# RETRIEVAL: Search your documents
embedder = MistralEmbedder(client=mistral_client, model_name=MODEL_1024_EMBEDDING)
query_engine = QueryEngine(
    retriever=[VectorRetriever(client=vector_store, embedder=embedder)],
)

results = await query_engine.search(query="What is RAG?", top_k=5)

# Print results
for result in results.results:
    print(f"Score: {result.score}")
    print(f"Content: {result.content}\n")

Plugins

Extend the toolkit with specialized backends:

Plugin Package Description
Vespa Plugin mistralai-search-toolkit-plugins-vespa Vespa search backend
Postgres Plugin mistralai-search-toolkit-plugins-postgres PostgreSQL + pgvector search backend
AWS S3 Storage mistralai-search-toolkit-storage-s3 AWS S3 storage backend
Azure Blob Storage mistralai-search-toolkit-storage-azure Azure Blob Storage backend
Google Cloud Storage mistralai-search-toolkit-storage-gcs Google Cloud Storage backend

License

This package is licensed under the Apache License 2.0.

Support

For more information and examples, visit Vespa documentation.

Metadata

Release files for mistralai-search-toolkit 0.0.14

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mistralai-search-toolkit 0.0.14
File Size Uploaded
mistralai_search_toolkit-0.0.14.tar.gz 282.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mistralai-search-toolkit 0.0.14
File Interpreter ABI Platform
mistralai_search_toolkit-0.0.14-py3-none-any.whl Python 3 none any Details

Total release size: 508.3 kB

Release files / mistralai_search_toolkit-0.0.14.tar.gz

Download URL mistralai_search_toolkit-0.0.14.tar.gz
Size 282.5 kB
Tags Source
SHA-256 checksum
How to use checksums
d83201edc3dc482e000b4b10c663e69da303d62c92253b146b4856223603c282
BLAKE2b-256 checksum
How to use checksums
6c5cfd1425aa7e1c216aec2b746dfba1db1d5649817700ac5c3ef9a95813a8b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release files / mistralai_search_toolkit-0.0.14-py3-none-any.whl

Download URL mistralai_search_toolkit-0.0.14-py3-none-any.whl
Size 225.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f605c788c7f173afdaa64adf1dee8365ef0be106e520ef12c543e46c45b8f11d
BLAKE2b-256 checksum
How to use checksums
9a4371b77aaf0e2b4489ac2d07009c661c6f096856cc2cdb0b1d1e7d4034402a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.14 This release

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.6

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page