Skip to main content
ai4rag icon

ai4rag

RAG Templates Optimization Engine

AI4RAG Python Python

RAG Builder HPO AutoML

Initializes RAG Templates with optimal parameters

Getting StartedUser GuideAPI ReferenceDevelopment


🎯 What is ai4RAG?

ai4RAG is an optimization engine for RAG Templates that is LLM and vector database provider-agnostic. It accepts a variety of RAG Templates and a search space definition, then returns an initialized RAG Template with optimal parameter values (called a RAG Pattern).

[!IMPORTANT] ai4rag is provider-agnostic. It reaches foundation and embedding models through the stock openai SDK, so any OpenAI-compatible endpoint works — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment, the integration ai4rag ships helpers for out of the box. You can also plug in your own foundation model, embedding model, or vector store by implementing the matching Base* interface. To run an experiment you'll need one foundation model and one embedding model (from any of the above), plus a vector store (Chroma, Milvus, or PostgreSQL/pgvector) connected directly via ai4rag.rag.vector_store.

Model providers

ai4rag reaches foundation and embedding models through the stock openai SDK, so it works with any OpenAI-compatible endpoint — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment. Prefer something else entirely? Implement BaseFoundationModel / BaseEmbeddingModel and pass your own models straight into an experiment.

MaaS is the integration ai4rag ships helpers for, so the walkthrough below uses it:

  • SDK: openai >= 2, < 3 (Python package used by ai4RAG; installs with this project).
  • Deployment: an OpenShift AI MaaS instance exposing at least one foundation model and one embedding model.
  • Endpoints: MaaS serves everything from a single OpenAI-compatible endpointMAAS_BASE_URL (host-only or /v1-suffixed; normalized automatically). One client lists the available models (models.list()) and serves chat/completions and embeddings for all of them. Model ids are used verbatim, exactly as models.list() reports them.

Features used by ai4rag

When using the MaaS backend, ai4rag relies on:

  • Embeddings — Text embeddings via the embeddings endpoint (e.g. for indexing and query encoding). Because models.list() carries no metadata, embedding dimension and context length are auto-detected at construction (or supplied via params).
  • Chat / completions — Foundation model integration for answer generation when evaluating RAG patterns.

Vector storage is independent of MaaS: ai4rag connects directly to Chroma, Milvus, or PostgreSQL/pgvector via the config classes in ai4rag.rag.vector_store (see Vector stores below).

Vector stores

ai4RAG talks to the vector store directly through provider-specific clients — no MaaS deployment is required for this part. Pick a provider and pass its config to AI4RAGExperiment as vector_store_config:

  • ChromaConfig — Chroma. Ephemeral in-memory by default; persistent (via persist_directory) or client/server (via host/port) modes are also supported. Vector-only search.
  • MilvusConfig — Milvus. Requires a uri; supports TLS (https:// scheme) and self-signed CAs via server_cert. Hybrid search (dense + BM25).
  • PGVectorConfig — PostgreSQL with the pgvector extension. Hybrid search (dense + tsvector full-text).

Each config is a frozen dataclass with a .from_env() constructor and an env_vars attribute listing the environment variables it reads (e.g. MILVUS_URI, PGVECTOR_HOST).

Document processing

ai4RAG uses docling-core for document representation and chunking. Documents are represented as DoclingDocument instances, and the DoclingChunker leverages docling's HybridChunker for structure-aware, token-aware chunking. docling-core, openai, and the vector store clients (chromadb, pymilvus, pgvector, asyncpg) are all installed automatically with ai4rag.

Quick start

  1. Prepare a MaaS client to integrate with your models.
  2. Prepare your knowledge base documents for the experiment.
  3. Prepare benchmark_data.json with evaluation questions and answers.
  4. Define and constrain your search space.
  5. Configure the optimizer.
  6. Create and run the experiment.

Prepare the MaaS client

To enable full integration with MaaS, build a single client that lists the available models and serves them all — ai4rag reuses it for every foundation and embedding model wrapper. The dev_utils helper create_dev_maas_client() reads MAAS_BASE_URL / MAAS_API_KEY and builds that client for you.

[!tip] Store your credentials securely in a .env file.

from dotenv import load_dotenv, find_dotenv
from dev_utils.utils import create_dev_maas_client

load_dotenv(find_dotenv())

client = create_dev_maas_client()  # reads MAAS_BASE_URL / MAAS_API_KEY

[!note] dev_utils is only available when cloning the repository. For the equivalent setup using the public API (the single OpenAI client built with create_maas_client), see the Provider-Agnostic Design guide.

Prepare knowledge base documents

Prepare a set of documents to serve as the knowledge base for retrieval. Documents are represented as DoclingDocument instances (from the docling-core library). A local folder is all you need — ai4rag does not require object storage.

Convert the folder with Docling, naming each document by its path relative to that folder:

from pathlib import Path
from docling.document_converter import DocumentConverter

documents_root = Path("<path to the documents folder>")
converter = DocumentConverter()

documents = []
for file_path in sorted(p for p in documents_root.rglob("*") if p.is_file()):
    document = converter.convert(file_path).document
    # The document's key: what benchmark data references.
    document.name = str(file_path.relative_to(documents_root))
    documents.append(document)

[!important] Each document's name is its key — the identifier carried through chunking, indexing and evaluation, and the value correct_answer_document_keys must reference. Set it explicitly: Docling otherwise derives a name from the file stem, so two files named setup.pdf in different folders would collide.

Already keeping your corpus in a bucket? discover_documents() and extract_text() name each document by its full S3 object key, so that is what correct_answer_document_keys must reference — prefix included.

Prepare benchmark_data.json

Create a benchmark_data.json file following this schema:

[
	{
		"question": "<question_1>",
		"correct_answers": [
			"<answer 1 for question 1>",
			"<answer 2 for question 1>"
		],
		"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
	},
	{
		"question": "<question_2>",
		"correct_answers": [
			"<answer 1 for question 2>",
			"<answer 2 for question 2>"
		],
		"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
	}
]

All benchmark questions and answers must be derived from your knowledge base documents, and every correct_answer_document_keys entry must equal the name of one of the documents you loaded above.

import pandas as pd

benchmark_data = pd.read_json("<path to benchmark_data.json>")

Define and constrain search space

The search space defines all possible parameter combinations, where each combination creates a unique RAG Pattern. During the experiment, the engine will optimize the RAG Pattern for the selected metric over the given search space, using an objective function to evaluate each configuration.

from ai4rag.search_space.src.parameter import Parameter
from ai4rag.search_space.src.search_space import AI4RAGSearchSpace
from dev_utils.utils import build_maas_model


search_space = AI4RAGSearchSpace(
    params=[
        Parameter(
            name="foundation_model",
            param_type="C",
            values=[build_maas_model(client, model_id="qwen3-8b-fp8-dynamic", model_type="llm")],
        ),
        Parameter(
            name="embedding_model",
            param_type="C",
            values=[
                build_maas_model(
                    client,
                    model_id="bge-m3",
                    model_type="embedding",
                    embedding_params={"embedding_dimension": 1024, "context_length": 8192},
                )
            ],
        ),
        Parameter(
            name="chunking_method",
            param_type="C",
            values=["recursive", "hybrid"],
        ),
        Parameter(
            name="chunk_size",
            param_type="C",
            values=[512, 1024, 2048],
        ),
        Parameter(
            name="chunk_overlap",
            param_type="C",
            values=[0, 128, 256],
        ),
    ]
)

[!tip] chunking_method controls the chunking strategy: "recursive" uses LangChain's RecursiveCharacterTextSplitter, while "hybrid" uses docling's structure-aware HybridChunker (requires chunk_overlap=0). When omitted, both methods are included by default.

[!tip] To validate model IDs and build a search space from a MaaS deployment in one call, use prepare_search_space_with_maas() from ai4rag.search_space.prepare, passing the MaaS client and the foundation/embedding model IDs per type.

Configure optimizer

You have full control over the optimization algorithm. Configure the GAMOptimizer by adjusting GAMOptSettings.

from ai4rag.core.hpo.gam_opt import GAMOptSettings

optimizer_settings = GAMOptSettings(
    max_evals=10, n_random_nodes=4
)

Run the experiment

Using the information from the previous steps, create an experiment and run the ai4rag optimization engine.

[!note] Select the vector store by passing a vector_store_config to AI4RAGExperiment: ChromaConfig() for a zero-config in-memory store (vector-only search), or MilvusConfig.from_env() / PGVectorConfig.from_env() for a server-backed store with hybrid (dense + keyword) search.

from ai4rag.core.experiment.experiment import AI4RAGExperiment
from ai4rag.rag.vector_store import MilvusConfig
from ai4rag.utils.event_handler import LocalEventHandler

experiment = AI4RAGExperiment(
    documents=documents,
    benchmark_data=benchmark_data,
    search_space=search_space,
    vector_store_config=MilvusConfig.from_env(),
    optimizer_settings=optimizer_settings,
    event_handler=LocalEventHandler(output_path="<local-path-to-store-your-output-files>"),
)

experiment.search()
best_eval = experiment.results.get_best_evaluations(k=1)[0]
print(best_eval)

print(f"Best pattern: {best_eval.pattern_name} (score: {best_eval.final_score})")

[!note] Each trial closes its vector store once it finishes, so EvaluationResult no longer exposes a reusable rag_pattern. Read the outcome from its fields (pattern_name, final_score, scores, rag_params); rebuild the pattern from those settings if you want to run inference.

[!tip] For production use, implement your own custom EventHandler to handle status changes and artifacts produced during the experiment. See the BaseEventHandler implementation for reference.

Contribution

Pull requests are very welcome! Make sure your patches are well tested. Ideally create a topic branch for every separate change you make.

Development setup

This project uses uv for dependency management.

# Clone the repository
git clone https://github.com/IBM/ai4rag.git
cd ai4rag

# Install all development dependencies
uv sync --extra dev

# Run tests
uv run pytest tests/unit/

# Check code style
uv run black --check ai4rag/
uv run pylint ai4rag/

# Build and serve documentation locally
uv run mkdocs serve

Pull request workflow

  1. Fork the repo
  2. Create your feature branch (git checkout -b my-new-feature)
  3. Commit your changes (git commit -s -am 'Added some feature')
  4. Push to the branch (git push origin my-new-feature)
  5. Create new Pull Request

See more details in contributing section.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ai4rag-0.16.0.tar.gz (157.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ai4rag-0.16.0-py3-none-any.whl (194.9 kB view details)

Uploaded Python 3

File details

Details for the file ai4rag-0.16.0.tar.gz.

File metadata

  • Download URL: ai4rag-0.16.0.tar.gz
  • Upload date:
  • Size: 157.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ai4rag-0.16.0.tar.gz
Algorithm Hash digest
SHA256 365d8622e7408d27de1bfd3e40e65d1c4ecd2ab4a12afa5e081e5d35c566c729
MD5 85b4e51bf15dcad4181bf2e05373d46f
BLAKE2b-256 eb658382f8195949355f25fc8008e5943c542e243e66cd05691b9b7ab6b1f40f

See more details on using hashes here.

Provenance

The following attestation bundles were made for ai4rag-0.16.0.tar.gz:

Publisher: publish-pypi.yml on IBM/ai4rag

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ai4rag-0.16.0-py3-none-any.whl.

File metadata

  • Download URL: ai4rag-0.16.0-py3-none-any.whl
  • Upload date:
  • Size: 194.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ai4rag-0.16.0-py3-none-any.whl
Algorithm Hash digest
SHA256 04cd163a25bc37e5bdcb6c2627e273aec045e73714c482bfba27cec928900ebb
MD5 115a1937e663971d51c98fc60e40b580
BLAKE2b-256 c1e2ee9cd7dd5bf60e195fd79a1f50d3c6015735f6f06857881f5b78852907e2

See more details on using hashes here.

Provenance

The following attestation bundles were made for ai4rag-0.16.0-py3-none-any.whl:

Publisher: publish-pypi.yml on IBM/ai4rag

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.16.0 This release

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.1

2 files

0.11.0

2 files

0.10.4

2 files

0.10.3

2 files

0.10.2

2 files

0.10.1

2 files

0.10.0

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.1

2 files

0.8.0

2 files

0.7.0

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.6

2 files

0.5.5

2 files

0.5.4

2 files

0.5.3

2 files

0.5.2

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

0.2.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page