🎯 What is ai4RAG?
ai4RAG is an optimization engine for RAG Templates that is LLM and vector database provider-agnostic.
It accepts a variety of RAG Templates and a search space definition, then returns an initialized RAG Template with optimal parameter values (called a RAG Pattern).
[!IMPORTANT]
ai4ragis provider-agnostic. It reaches foundation and embedding models through the stockopenaiSDK, so any OpenAI-compatible endpoint works — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment, the integrationai4ragships helpers for out of the box. You can also plug in your own foundation model, embedding model, or vector store by implementing the matchingBase*interface. To run an experiment you'll need one foundation model and one embedding model (from any of the above), plus a vector store (Chroma, Milvus, or PostgreSQL/pgvector) connected directly viaai4rag.rag.vector_store.
Model providers
ai4rag reaches foundation and embedding models through the stock openai SDK, so it works with any OpenAI-compatible endpoint — a hosted API, a self-managed server (vLLM, TGI, Ollama, …), or an OpenShift AI Models-as-a-Service (MaaS) deployment. Prefer something else entirely? Implement BaseFoundationModel / BaseEmbeddingModel and pass your own models straight into an experiment.
MaaS is the integration ai4rag ships helpers for, so the walkthrough below uses it:
- SDK: openai >= 2, < 3 (Python package used by ai4RAG; installs with this project).
- Deployment: an OpenShift AI MaaS instance exposing at least one foundation model and one embedding model.
- Endpoints: MaaS serves everything from a single OpenAI-compatible endpoint —
MAAS_BASE_URL(host-only or/v1-suffixed; normalized automatically). One client lists the available models (models.list()) and serves chat/completions and embeddings for all of them. Model ids are used verbatim, exactly asmodels.list()reports them.
Features used by ai4rag
When using the MaaS backend, ai4rag relies on:
- Embeddings — Text embeddings via the
embeddingsendpoint (e.g. for indexing and query encoding). Becausemodels.list()carries no metadata, embedding dimension and context length are auto-detected at construction (or supplied viaparams). - Chat / completions — Foundation model integration for answer generation when evaluating RAG patterns.
Vector storage is independent of MaaS: ai4rag connects directly to Chroma, Milvus, or PostgreSQL/pgvector via the config classes in ai4rag.rag.vector_store (see Vector stores below).
Vector stores
ai4RAG talks to the vector store directly through provider-specific clients — no MaaS deployment is required for this part. Pick a provider and pass its config to AI4RAGExperiment as vector_store_config:
ChromaConfig— Chroma. Ephemeral in-memory by default; persistent (viapersist_directory) or client/server (viahost/port) modes are also supported. Vector-only search.MilvusConfig— Milvus. Requires auri; supports TLS (https://scheme) and self-signed CAs viaserver_cert. Hybrid search (dense + BM25).PGVectorConfig— PostgreSQL with thepgvectorextension. Hybrid search (dense +tsvectorfull-text).
Each config is a frozen dataclass with a .from_env() constructor and an env_vars attribute listing the environment variables it reads (e.g. MILVUS_URI, PGVECTOR_HOST).
Document processing
ai4RAG uses docling-core for document representation and chunking. Documents are represented as DoclingDocument instances, and the DoclingChunker leverages docling's HybridChunker for structure-aware, token-aware chunking. docling-core, openai, and the vector store clients (chromadb, pymilvus, pgvector, asyncpg) are all installed automatically with ai4rag.
Quick start
- Prepare a MaaS client to integrate with your models.
- Prepare your knowledge base documents for the experiment.
- Prepare
benchmark_data.jsonwith evaluation questions and answers. - Define and constrain your search space.
- Configure the optimizer.
- Create and run the experiment.
Prepare the MaaS client
To enable full integration with MaaS, build a single client that lists the available models and serves them all — ai4rag reuses it for every foundation and embedding model wrapper.
The dev_utils helper create_dev_maas_client() reads MAAS_BASE_URL / MAAS_API_KEY and builds that client for you.
[!tip] Store your credentials securely in a
.envfile.
from dotenv import load_dotenv, find_dotenv
from dev_utils.utils import create_dev_maas_client
load_dotenv(find_dotenv())
client = create_dev_maas_client() # reads MAAS_BASE_URL / MAAS_API_KEY
[!note]
dev_utilsis only available when cloning the repository. For the equivalent setup using the public API (the singleOpenAIclient built withcreate_maas_client), see the Provider-Agnostic Design guide.
Prepare knowledge base documents
Prepare a set of documents to serve as the knowledge base for retrieval.
Documents are represented as DoclingDocument instances (from the docling-core library).
A local folder is all you need — ai4rag does not require object storage.
Convert the folder with Docling, naming each document by its path relative to that folder:
from pathlib import Path
from docling.document_converter import DocumentConverter
documents_root = Path("<path to the documents folder>")
converter = DocumentConverter()
documents = []
for file_path in sorted(p for p in documents_root.rglob("*") if p.is_file()):
document = converter.convert(file_path).document
# The document's key: what benchmark data references.
document.name = str(file_path.relative_to(documents_root))
documents.append(document)
[!important] Each document's
nameis its key — the identifier carried through chunking, indexing and evaluation, and the valuecorrect_answer_document_keysmust reference. Set it explicitly: Docling otherwise derives a name from the file stem, so two files namedsetup.pdfin different folders would collide.
Already keeping your corpus in a bucket? discover_documents() and extract_text() name each document by
its full S3 object key, so that is what correct_answer_document_keys must reference — prefix included.
Prepare benchmark_data.json
Create a benchmark_data.json file following this schema:
[
{
"question": "<question_1>",
"correct_answers": [
"<answer 1 for question 1>",
"<answer 2 for question 1>"
],
"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
},
{
"question": "<question_2>",
"correct_answers": [
"<answer 1 for question 2>",
"<answer 2 for question 2>"
],
"correct_answer_document_keys": ["<list of document keys based on which correct answers were generated>"]
}
]
All benchmark questions and answers must be derived from your knowledge base documents, and every
correct_answer_document_keys entry must equal the name of one of the documents you loaded above.
import pandas as pd
benchmark_data = pd.read_json("<path to benchmark_data.json>")
Define and constrain search space
The search space defines all possible parameter combinations, where each combination creates a unique RAG Pattern. During the experiment, the engine will optimize the RAG Pattern for the selected metric over the given search space, using an objective function to evaluate each configuration.
from ai4rag.search_space.src.parameter import Parameter
from ai4rag.search_space.src.search_space import AI4RAGSearchSpace
from dev_utils.utils import build_maas_model
search_space = AI4RAGSearchSpace(
params=[
Parameter(
name="foundation_model",
param_type="C",
values=[build_maas_model(client, model_id="qwen3-8b-fp8-dynamic", model_type="llm")],
),
Parameter(
name="embedding_model",
param_type="C",
values=[
build_maas_model(
client,
model_id="bge-m3",
model_type="embedding",
embedding_params={"embedding_dimension": 1024, "context_length": 8192},
)
],
),
Parameter(
name="chunking_method",
param_type="C",
values=["recursive", "hybrid"],
),
Parameter(
name="chunk_size",
param_type="C",
values=[512, 1024, 2048],
),
Parameter(
name="chunk_overlap",
param_type="C",
values=[0, 128, 256],
),
]
)
[!tip]
chunking_methodcontrols the chunking strategy:"recursive"uses LangChain'sRecursiveCharacterTextSplitter, while"hybrid"uses docling's structure-awareHybridChunker(requireschunk_overlap=0). When omitted, both methods are included by default.
[!tip] To validate model IDs and build a search space from a MaaS deployment in one call, use
prepare_search_space_with_maas()fromai4rag.search_space.prepare, passing the MaaS client and the foundation/embedding model IDs per type.
Configure optimizer
You have full control over the optimization algorithm. Configure the GAMOptimizer by adjusting GAMOptSettings.
from ai4rag.core.hpo.gam_opt import GAMOptSettings
optimizer_settings = GAMOptSettings(
max_evals=10, n_random_nodes=4
)
Run the experiment
Using the information from the previous steps, create an experiment and run the ai4rag optimization engine.
[!note] Select the vector store by passing a
vector_store_configtoAI4RAGExperiment:ChromaConfig()for a zero-config in-memory store (vector-only search), orMilvusConfig.from_env()/PGVectorConfig.from_env()for a server-backed store with hybrid (dense + keyword) search.
from ai4rag.core.experiment.experiment import AI4RAGExperiment
from ai4rag.rag.vector_store import MilvusConfig
from ai4rag.utils.event_handler import LocalEventHandler
experiment = AI4RAGExperiment(
documents=documents,
benchmark_data=benchmark_data,
search_space=search_space,
vector_store_config=MilvusConfig.from_env(),
optimizer_settings=optimizer_settings,
event_handler=LocalEventHandler(output_path="<local-path-to-store-your-output-files>"),
)
experiment.search()
best_eval = experiment.results.get_best_evaluations(k=1)[0]
print(best_eval)
print(f"Best pattern: {best_eval.pattern_name} (score: {best_eval.final_score})")
[!note] Each trial closes its vector store once it finishes, so
EvaluationResultno longer exposes a reusablerag_pattern. Read the outcome from its fields (pattern_name,final_score,scores,rag_params); rebuild the pattern from those settings if you want to run inference.
[!tip] For production use, implement your own custom
EventHandlerto handle status changes and artifacts produced during the experiment. See theBaseEventHandlerimplementation for reference.
Contribution
Pull requests are very welcome! Make sure your patches are well tested. Ideally create a topic branch for every separate change you make.
Development setup
This project uses uv for dependency management.
# Clone the repository
git clone https://github.com/IBM/ai4rag.git
cd ai4rag
# Install all development dependencies
uv sync --extra dev
# Run tests
uv run pytest tests/unit/
# Check code style
uv run black --check ai4rag/
uv run pylint ai4rag/
# Build and serve documentation locally
uv run mkdocs serve
Pull request workflow
- Fork the repo
- Create your feature branch (
git checkout -b my-new-feature) - Commit your changes (
git commit -s -am 'Added some feature') - Push to the branch (
git push origin my-new-feature) - Create new Pull Request
See more details in contributing section.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai4rag-0.16.0.tar.gz.
File metadata
- Download URL: ai4rag-0.16.0.tar.gz
- Upload date:
- Size: 157.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
365d8622e7408d27de1bfd3e40e65d1c4ecd2ab4a12afa5e081e5d35c566c729
|
|
| MD5 |
85b4e51bf15dcad4181bf2e05373d46f
|
|
| BLAKE2b-256 |
eb658382f8195949355f25fc8008e5943c542e243e66cd05691b9b7ab6b1f40f
|
Provenance
The following attestation bundles were made for ai4rag-0.16.0.tar.gz:
Publisher:
publish-pypi.yml on IBM/ai4rag
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ai4rag-0.16.0.tar.gz -
Subject digest:
365d8622e7408d27de1bfd3e40e65d1c4ecd2ab4a12afa5e081e5d35c566c729 - Sigstore transparency entry: 2784114904
- Sigstore integration time:
-
Permalink:
IBM/ai4rag@04249e52e7864bbfd0dbd197a5ce21838925ca65 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/IBM
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@04249e52e7864bbfd0dbd197a5ce21838925ca65 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file ai4rag-0.16.0-py3-none-any.whl.
File metadata
- Download URL: ai4rag-0.16.0-py3-none-any.whl
- Upload date:
- Size: 194.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
04cd163a25bc37e5bdcb6c2627e273aec045e73714c482bfba27cec928900ebb
|
|
| MD5 |
115a1937e663971d51c98fc60e40b580
|
|
| BLAKE2b-256 |
c1e2ee9cd7dd5bf60e195fd79a1f50d3c6015735f6f06857881f5b78852907e2
|
Provenance
The following attestation bundles were made for ai4rag-0.16.0-py3-none-any.whl:
Publisher:
publish-pypi.yml on IBM/ai4rag
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ai4rag-0.16.0-py3-none-any.whl -
Subject digest:
04cd163a25bc37e5bdcb6c2627e273aec045e73714c482bfba27cec928900ebb - Sigstore transparency entry: 2784114966
- Sigstore integration time:
-
Permalink:
IBM/ai4rag@04249e52e7864bbfd0dbd197a5ce21838925ca65 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/IBM
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@04249e52e7864bbfd0dbd197a5ce21838925ca65 -
Trigger Event:
workflow_dispatch
-
Statement type: