Skip to main content

Haystack component to embed strings and Documents using VoyageAI Embedding models.

Project description

Voyage Embedders and Rankers - Haystack

PyPI Downloads License Tests Coverage pre-commit.ci status Types Ruff

Custom components for Haystack for creating embeddings and reranking documents using the Voyage Models.

Voyage’s embedding models are state-of-the-art in retrieval accuracy. These models outperform top performing embedding models like intfloat/e5-mistral-7b-instruct and OpenAI/text-embedding-3-large on the MTEB Benchmark.

What's New (v2.0.0)

  • Haystack 3.x compatibility: All components now work with haystack-ai>=3.0.0.
  • Lazy initialization: Components now use warm_up() for lazy client creation. The client property auto-initializes on first access.
  • Async support: All components now have async def run() methods for Haystack 3.x async pipelines.
  • Document immutability: Document embedders now use dataclasses.replace() to return new Document instances instead of mutating them in-place.

See the full Changelog for all releases.

Requirements

Installation

pip install voyage-embedders-haystack

Usage

You can use Voyage Embedding models with multiple components:

The Voyage Reranker models can be used with the VoyageRanker component.

Multimodal Embeddings

The VoyageMultimodalEmbedder uses Voyage's multimodal embedding model (voyage-multimodal-3.5) to encode text, images, and videos into a shared vector space. This enables cross-modal similarity search where you can find images using text queries or find related content across different modalities.

Key features:

  • Supports text, images (PIL Images, ByteStream), and videos
  • Inputs can combine multiple modalities (e.g., text + image)
  • Variable output dimensions: 256, 512, 1024 (default), 2048
  • Recommended model: voyage-multimodal-3.5

Usage example:

from haystack.dataclasses import ByteStream
from haystack_integrations.components.embedders.voyage_embedders import VoyageMultimodalEmbedder
from voyageai.video_utils import Video

# Text-only embedding
embedder = VoyageMultimodalEmbedder(model="voyage-multimodal-3.5")
result = embedder.run(inputs=[["What is in this image?"]])

# Mixed text and image embedding
image_bytes = ByteStream.from_file_path("image.jpg")
result = embedder.run(inputs=[["Describe this image:", image_bytes]])

# Video embedding
video = Video.from_path("video.mp4", model="voyage-multimodal-3.5")
result = embedder.run(inputs=[["Describe this video:", video]])

Contextualized Chunk Embeddings

The VoyageContextualizedDocumentEmbedder uses Voyage's contextualized embedding models to encode document chunks "in context" with other chunks from the same document. This approach preserves semantic relationships between chunks and reduces context loss, leading to improved retrieval accuracy.

Key features:

  • Documents are grouped by a metadata field (default: source_id)
  • Chunks from the same source document are embedded together
  • Maintains semantic connections between related chunks
  • Recommended model: voyage-context-3

For detailed usage examples, see the contextualized embedder example.

Once you've selected the suitable component for your specific use case, initialize the component with the model name and VoyageAI API key. You can also set the environment variable VOYAGE_API_KEY instead of passing the API key as an argument. To get an API key, please see the Voyage AI website.

Information about the supported models, can be found on the Voyage AI Documentation.

Example

You can find all the examples in the examples folder.

Below is the example Semantic Search pipeline that uses the Simple Wikipedia Dataset from HuggingFace.

Load the dataset:

# Install HuggingFace Datasets using "pip install datasets"
from datasets import load_dataset
from haystack import Pipeline
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.components.writers import DocumentWriter
from haystack.dataclasses import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore

# Import Voyage Embedders
from haystack_integrations.components.embedders.voyage_embedders import VoyageDocumentEmbedder, VoyageTextEmbedder

# Load first 100 rows of the Simple Wikipedia Dataset from HuggingFace
dataset = load_dataset("pszemraj/simple_wikipedia", split="validation[:100]")

docs = [
    Document(
        content=doc["text"],
        meta={
            "title": doc["title"],
            "url": doc["url"],
        },
    )
    for doc in dataset
]

Index the documents to the InMemoryDocumentStore using the VoyageDocumentEmbedder and DocumentWriter:

doc_store = InMemoryDocumentStore(embedding_similarity_function="cosine")
retriever = InMemoryEmbeddingRetriever(document_store=doc_store)
doc_writer = DocumentWriter(document_store=doc_store)

doc_embedder = VoyageDocumentEmbedder(
    model="voyage-4",
    input_type="document",
)
text_embedder = VoyageTextEmbedder(model="voyage-4", input_type="query")

# Indexing Pipeline
indexing_pipeline = Pipeline()
indexing_pipeline.add_component(instance=doc_embedder, name="DocEmbedder")
indexing_pipeline.add_component(instance=doc_writer, name="DocWriter")
indexing_pipeline.connect("DocEmbedder", "DocWriter")

indexing_pipeline.run({"DocEmbedder": {"documents": docs}})

print(f"Number of documents in Document Store: {len(doc_store.filter_documents())}")
print(f"First Document: {doc_store.filter_documents()[0]}")
print(f"Embedding of first Document: {doc_store.filter_documents()[0].embedding}")

Query the Semantic Search Pipeline using the InMemoryEmbeddingRetriever and VoyageTextEmbedder:

text_embedder = VoyageTextEmbedder(model="voyage-4", input_type="query")

# Query Pipeline
query_pipeline = Pipeline()
query_pipeline.add_component(instance=text_embedder, name="TextEmbedder")
query_pipeline.add_component(instance=retriever, name="Retriever")
query_pipeline.connect("TextEmbedder.embedding", "Retriever.query_embedding")

# Search
results = query_pipeline.run({"TextEmbedder": {"text": "Which year did the Joker movie release?"}})

# Print text from top result
top_result = results["Retriever"]["documents"][0].content
print("The top search result is:")
print(top_result)

Contributing

We welcome contributions from the community! Please take a look at our contributing guide for more details on how to get started.

Pull requests are welcome. For major changes, please open an issue first to discuss the proposed changes.

License

voyage-embedders-haystack is distributed under the terms of the Apache-2.0 license.

Maintained by Ashwin Mathur.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

voyage_embedders_haystack-2.0.0.tar.gz (16.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

voyage_embedders_haystack-2.0.0-py3-none-any.whl (25.3 kB view details)

Uploaded Python 3

File details

Details for the file voyage_embedders_haystack-2.0.0.tar.gz.

File metadata

  • Download URL: voyage_embedders_haystack-2.0.0.tar.gz
  • Upload date:
  • Size: 16.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for voyage_embedders_haystack-2.0.0.tar.gz
Algorithm Hash digest
SHA256 0ea6463e1a5a95447d97699a1295d8d9ba35135d70693ef5413bd37b90dce354
MD5 177993bb525095274a52fd91548bb9a3
BLAKE2b-256 a44e72b097881db876b035064f924d5ab9b8eec39f6a3f60119508c2b6e66056

See more details on using hashes here.

File details

Details for the file voyage_embedders_haystack-2.0.0-py3-none-any.whl.

File metadata

  • Download URL: voyage_embedders_haystack-2.0.0-py3-none-any.whl
  • Upload date:
  • Size: 25.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for voyage_embedders_haystack-2.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e733bd456c86e915fdcf1b228e459b476fb515a658a12aa97c70062d89975eb9
MD5 f0620a4b1d2c7e3a259001979c821947
BLAKE2b-256 f5b179b51ba9da8e8a9b45f8475111fce672f5be7abf85e866ff002b9347c914

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page