Skip to main content

License mloda Python

rag-integration

RAG integration plugin for mloda. Composes modular FeatureGroups into text and image processing pipelines with PII redaction, chunking, deduplication, embedding, FAISS vector search, retrieval, evaluation, and LLM response generation.

See the demo notebook for an interactive walkthrough or the CLI for command-line usage.

Project Structure

rag_integration/
  feature_groups/
    rag_pipeline/       # Text: source, PII, chunk, dedup, embed, index, retrieve, LLM
    image_pipeline/     # Image: source, PII, preprocess, dedup, embed
    datasets/           # BEIR (SciFact) and image (Flickr30k) loaders
    evaluation/         # Retrieval metrics (precision, recall, NDCG, MAP)
cli/                    # Command-line demo tools (see cli/README.md)
tests/                  # Unit and integration tests
docs/                   # Detailed guides

Quick Start

1. Define your documents

Documents are loaded through a DataCreator-based FeatureGroup. The simplest option is DictDocumentSource, which loads from a Python list:

documents = [
    {"doc_id": "1", "text": "Contact support@example.com for help."},
    {"doc_id": "2", "text": "Our office is at 123 Main St."},
]

2. Build the pipeline

Each pipeline stage is expressed as a feature name. Stages chain with __:

docs                                    # raw documents
docs__pii_redacted                      # PII removed
docs__pii_redacted__chunked             # text split into chunks
docs__pii_redacted__chunked__deduped    # duplicates removed
docs__pii_redacted__chunked__deduped__embedded  # vector embeddings

3. Run with mlodaAPI

from mloda.user import mlodaAPI, PluginCollector, Feature, Options
from mloda_plugins.compute_framework.base_implementations.python_dict.python_dict_framework import (
    PythonDictFramework,
)

from rag_integration.feature_groups.rag_pipeline import (
    DictDocumentSource,
    RegexPIIRedactor,
    FixedSizeChunker,
    ExactHashDeduplicator,
    MockEmbedder,
)

providers = {
    DictDocumentSource,
    RegexPIIRedactor,
    FixedSizeChunker,
    ExactHashDeduplicator,
    MockEmbedder,
}

results = mlodaAPI.run_all(
    features=["docs__pii_redacted__chunked__deduped__embedded"],
    compute_frameworks={PythonDictFramework},
    plugin_collector=PluginCollector.enabled_feature_groups(providers),
)

4. Configure pipeline stages

Use Options to select specific implementations and tune parameters:

feature = Feature(
    "docs__pii_redacted__chunked__deduped__embedded",
    options=Options(context={
        "redaction_method": "regex",        # or "simple", "pattern", "presidio"
        "chunking_method": "sentence",      # or "fixed_size", "paragraph", "semantic"
        "deduplication_method": "exact_hash",  # or "normalized", "ngram"
        "embedding_method": "sentence_transformer",  # or "hash", "tfidf", "mock"
        "chunk_size": 512,
        "chunk_overlap": 128,
    }),
)

Available Components

Text Pipeline

Stage Implementations
Document Source DictDocumentSource, FileDocumentSource
PII Redaction RegexPIIRedactor, SimplePIIRedactor, PatternPIIRedactor, PresidioPIIRedactor
Chunking FixedSizeChunker, SentenceChunker, ParagraphChunker, SemanticChunker
Deduplication ExactHashDeduplicator, NormalizedDeduplicator, NGramDeduplicator
Embedding MockEmbedder, HashEmbedder, TfidfEmbedder, SentenceTransformerEmbedder
Vector Store FaissFlatIndexer, FaissIVFIndexer, FaissHNSWIndexer
Retrieval FaissRetriever
LLM Response ClaudeCliResponse

Image Pipeline

Stage Implementations
Image Source DictImageSource, FileImageSource
PII Redaction BlurPIIRedactor, PixelPIIRedactor, SolidFillPIIRedactor
Preprocessing ResizePreprocessor, NormalizePreprocessor, ThumbnailPreprocessor
Deduplication ExactHashImageDeduplicator, PerceptualHashImageDeduplicator, DifferenceHashImageDeduplicator
Embedding MockImageEmbedder, HashImageEmbedder, CLIPImageEmbedder

Connector families

Alongside the build-your-own stage pipeline, the connectors/ package wraps whole external open-source RAG tools under one mloda surface, organized into six families by query-contract shape (retrieve, rerank, generate, graph_rag, structured, orchestrator). You swap backends by changing options, not by rewriting a pipeline.

The two layers share one seam: the FAISS retrieval stage is the native dense path of the retrieve family (retrieve_backend="faiss"), and a stage and its connector counterpart emit the same passage / answer row shape under the same canonical feature name, so migrating between them is an option swap. See "Relationship to the stage pipeline" in docs/rag-connector-base-classes.md.

See feature_groups/connectors/README.md for the family map (per-family contract, backends, no-Docker concrete, and pedigree), runnable examples, and links to the contract suites. The design rationale is in docs/rag-connector-base-classes.md.

from mloda.user import mlodaAPI, Feature, Options, PluginCollector
from mloda_plugins.compute_framework.base_implementations.python_dict.python_dict_framework import (
    PythonDictFramework,
)
from rag_integration.feature_groups.connectors.retrieve import Bm25sRetriever

feature = Feature(
    "retrieved_passages",
    options=Options(context={
        "retrieve_backend": "bm25s",
        "query_text": "cat pet",
        "corpus": [
            {"doc_id": "d1", "text": "A cat is an independent and curious pet."},
            {"doc_id": "d2", "text": "Cars need regular engine oil and maintenance."},
        ],
        "top_k": 3,
    }),
)
results = mlodaAPI.run_all(
    [feature],
    compute_frameworks={PythonDictFramework},
    plugin_collector=PluginCollector.enabled_feature_groups({Bm25sRetriever}),
)

Install a family's backend with uv sync --extra connectors (or rerank / graph / structured / orchestrator).

Swapping one backend for another is an option change, not a pipeline rewrite. python -m cli.swap_demo runs that swap within a family (retrieve_backend="bm25s" -> "tfidf") and across families (retrieve vs orchestrator over the same inputs); the contract is written up under "Swapping backends" in docs/rag-connector-base-classes.md.

Installation

Clone the repository and install with uv:

git clone https://github.com/mloda-ai/rag_integration.git
cd rag_integration
uv venv
source .venv/bin/activate
uv sync --all-extras

To install only specific extras, use uv sync --extra <name>:

Extra What it adds
faiss FAISS vector indexing (faiss-cpu)
advanced Presidio, sentence-transformers, joblib, Pillow, FAISS
eval BEIR benchmark datasets, pandas, numpy
graph networkx graph-RAG backend (NetworkxGraphRag)
dev tox, pytest, ruff, mypy, bandit

CLI

A command-line interface is available for running pipelines interactively. See cli/README.md for full usage.

python3 -m cli.rag_demo run --input cli/docs/ --pii regex --chunking sentence --embedding tfidf -v

Development Setup

uv venv
source .venv/bin/activate
uv sync --all-extras

Run all checks (pytest, ruff, mypy, bandit):

tox

Run individual checks

pytest
ruff format --check --line-length 120 .
ruff check .
mypy --strict --ignore-missing-imports .
bandit -c pyproject.toml -r -q .

Related

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rag_integration-0.4.1.tar.gz (100.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

rag_integration-0.4.1-py3-none-any.whl (175.6 kB view details)

Uploaded Python 3

File details

Details for the file rag_integration-0.4.1.tar.gz.

File metadata

  • Download URL: rag_integration-0.4.1.tar.gz
  • Upload date:
  • Size: 100.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.20

File hashes

Hashes for rag_integration-0.4.1.tar.gz
Algorithm Hash digest
SHA256 859bdbc181b13bd7249f906c41e48e2286ca01ba0260effaa90deee5235036b5
MD5 2d568e98687a07ff38b21ef481737faa
BLAKE2b-256 3f53f7b8cb4c6bf026590fa87afdabfe9265c8f865841d0d1c6d9b7e94c2c57b

See more details on using hashes here.

File details

Details for the file rag_integration-0.4.1-py3-none-any.whl.

File metadata

File hashes

Hashes for rag_integration-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 5f18faccad2fc13104a4ff51bfb78cef8c00bbca2cd05028e9b1fe7a4aa30527
MD5 88ebb573ac5e2fb3043def0e6d6b16a1
BLAKE2b-256 ed776b93fa7cb8dba0e1094a1ae4650119d887d07917dd51557a23db625cb060

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page