Skip to main content

Verbatim RAG

ChiliGround Logo
Chill, I Ground! 🌶 ️

A minimalistic approach to Retrieval-Augmented Generation (RAG) that prevents hallucination by ensuring all generated content is explicitly derived from source documents.

PyPI License Open In Colab ACL 2025

Concept

Traditional RAG systems retrieve relevant documents and then allow an LLM to freely generate responses based on that context. This can lead to hallucinations where the model invents facts not present in the source material.

Verbatim RAG solves this by extracting verbatim text spans from documents and composing responses entirely from these exact passages, with direct citations linking back to sources.

For extraction, we can use LLM-based span extractors or fine-tuned encoder-based models like ModernBERT. We've trained our own ModernBERT model for this purpose, which is available on HuggingFace (we've trained it on the RAGBench dataset).

With this approach, the whole RAG pipeline can be run without any usage of LLMs, and with using SPLADE embeddings, the pipeline can be run entirely on CPU, making it lightweight and efficient.

Installation

# Install the package
pip install verbatim-rag

For local development:

pip install -e packages/core/
pip install -e .

Lightweight Core

If you only need the reusable verbatim core without the full RAG pipeline (no torch, transformers, or Milvus):

pip install verbatim-core
from verbatim_core import VerbatimTransform

vt = VerbatimTransform()
response = vt.transform(
    question="What is the main finding?",
    context=[
        {"content": "The study found that X leads to Y.", "title": "Paper A"},
        {"content": "Results show Z is significant.", "title": "Paper B"},
    ],
)
print(response.answer)

Dependencies: only openai, pydantic, rapidfuzz, and jinja2.

Quick Start

from verbatim_rag import VerbatimIndex, VerbatimRAG
from verbatim_rag.ingestion import DocumentProcessor
from verbatim_rag.vector_stores import LocalMilvusStore
from verbatim_rag.embedding_providers import SpladeProvider

# Process documents with intelligent chunking
processor = DocumentProcessor()

# Process PDFs from URLs
document = processor.process_url(
    url="https://aclanthology.org/2025.bionlp-share.8.pdf",
    title="KR Labs at ArchEHR-QA 2025: A Verbatim Approach for Evidence-Based Question Answering",
    metadata={"authors": ["Adam Kovacs", "Paul Schmitt", "Gabor Recski"]}
)

# Create embedding provider and vector store
sparse_provider = SpladeProvider(
    model_name="opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distill",
    device="cpu"
)
vector_store = LocalMilvusStore(
    db_path="./index.db",
    collection_name="verbatim_rag",
    enable_dense=False,
    enable_sparse=True,
)

# Create index with providers
index = VerbatimIndex(
    vector_store=vector_store,
    sparse_provider=sparse_provider
)
index.add_documents([document])

# Then query the index
rag = VerbatimRAG(index)

response = rag.query("What is the main contribution of the paper?")
print(response.answer)

Environment Setup

Set your OpenAI API key before using the system:

export OPENAI_API_KEY=your_api_key_here

How It Works

  1. Document Processing: Documents are processed using docling for format conversion and chonkie for chunking
  2. Document Indexing: Documents are indexed using vector embeddings (both dense and sparse)
  3. Template Management: Response templates are created and stored for common question types
  4. Query Processing:
    • Relevant documents are retrieved
    • Key passages are extracted verbatim using either LLM-based or fine-tuned span extractors
    • Responses are structured using templates
    • Citations link back to source documents

This ensures all responses are grounded in the source material, preventing hallucinations.

Architecture

Core Components

  • VerbatimRAG (verbatim_rag/core.py): Main orchestrator that coordinates document retrieval, span extraction, and response generation
  • VerbatimIndex (verbatim_rag/index.py): Vector-based document indexing and retrieval
  • SpanExtractor (verbatim_rag/extractors.py): Abstract interface for extracting relevant text spans from documents
    • LLMSpanExtractor: Uses OpenAI models to identify relevant spans
    • ModelSpanExtractor: Uses fine-tuned BERT-based models for span classification
  • DocumentProcessor (verbatim_rag/ingestion/): Docling + Chonkie integration for intelligent document processing
  • Document (verbatim_rag/document.py): Core document representation with metadata

Data Flow

  1. Documents are processed and chunked using docling and chonkie
  2. Documents are indexed using vector embeddings
  3. User queries retrieve relevant documents
  4. Span extractors identify verbatim passages that answer the question
  5. Response templates structure the final answer with citations
  6. All responses include exact text spans and document references

Web Interface

The package includes a full web interface with React frontend and FastAPI backend:

# Start API server
python api/app.py

# Start React frontend (in another terminal)
cd frontend/
npm install
npm start

ModernBERT Based Span Extractor

We've trained our own encoder model based on ModernBERT for sentence classification. This model is designed to classify text spans as relevant or not, providing a robust alternative to LLM-based extractors.

You can find our model on HuggingFace: KRLabsOrg/verbatim-rag-modern-bert-v1.

You can use it with the defined index as follows:

from verbatim_rag.core import VerbatimRAG
from verbatim_rag.index import VerbatimIndex
from verbatim_rag.extractors import ModelSpanExtractor
from verbatim_rag.vector_stores import LocalMilvusStore
from verbatim_rag.embedding_providers import SpladeProvider

# Load your trained extractor
extractor = ModelSpanExtractor("KRLabsOrg/verbatim-rag-modern-bert-v1")

# Create embedding provider and vector store
sparse_provider = SpladeProvider(
    model_name="opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distill",
    device="cpu"
)
vector_store = LocalMilvusStore(
    db_path="./index.db",
    collection_name="verbatim_rag",
    enable_dense=False,
    enable_sparse=True,
)

# Create index with providers
# (Assuming you have already populated the index)
index = VerbatimIndex(
    vector_store=vector_store,
    sparse_provider=sparse_provider
)

# Create VerbatimRAG system with custom extractor
rag_system = VerbatimRAG(
    index=index,
    extractor=extractor,
    k=5
)

# Query the system
response = rag_system.query("Main findings of the paper?")
print(response.answer)

Citation

If you use Verbatim RAG in your research, please cite our paper:

@inproceedings{kovacs-etal-2025-kr,
    title = "{KR} Labs at {A}rch{EHR}-{QA} 2025: A Verbatim Approach for Evidence-Based Question Answering",
    author = "Kovacs, Adam  and
      Schmitt, Paul  and
      Recski, Gabor",
    editor = "Soni, Sarvesh  and
      Demner-Fushman, Dina",
    booktitle = "Proceedings of the 24th Workshop on Biomedical Language Processing (Shared Tasks)",
    month = aug,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.bionlp-share.8/",
    pages = "69--74",
    ISBN = "979-8-89176-276-3",
    abstract = "We present a lightweight, domain{-}agnostic verbatim pipeline for evidence{-}grounded question answering. Our pipeline operates in two steps: first, a sentence-level extractor flags relevant note sentences using either zero-shot LLM prompts or supervised ModernBERT classifiers. Next, an LLM drafts a question-specific template, which is filled verbatim with sentences from the extraction step. This prevents hallucinations and ensures traceability. In the ArchEHR{-}QA 2025 shared task, our system scored 42.01{\%}, ranking top{-}10 in core metrics and outperforming the organiser{'}s 70B{-}parameter Llama{-}3.3 baseline. We publicly release our code and inference scripts under an MIT license."
}

Metadata

Release files for verbatim-rag 0.2.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verbatim-rag 0.2.8
File Size Uploaded
verbatim_rag-0.2.8.tar.gz 59.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verbatim-rag 0.2.8
File Interpreter ABI Platform
verbatim_rag-0.2.8-py3-none-any.whl Python 3 none any Details

Total release size: 130.8 kB

Release files / verbatim_rag-0.2.8.tar.gz

Download URL verbatim_rag-0.2.8.tar.gz
Size 59.5 kB
Tags Source
SHA-256 checksum
How to use checksums
310d23c9526ee261efcc9c66333a92bebde2681cfc0d2c9e559eb59709042d53
BLAKE2b-256 checksum
How to use checksums
a9779ae2a545de0999dea4a83d42fe402e1af9ebe018a6ce31faae8c4d549e04
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.13

Release files / verbatim_rag-0.2.8-py3-none-any.whl

Download URL verbatim_rag-0.2.8-py3-none-any.whl
Size 71.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bb2ca2753dae9ae87cbfcdcb77bb0191ab0da5b846324754423fe39fbf67601e
BLAKE2b-256 checksum
How to use checksums
5a44326d1083d8978d7166c520494bfbf0de64154a767c5cc86484ab33374d79
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.13

Release history Release notifications | RSS feed

This release

0.2.8 This release

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

1 release file

0.2.3

1 release file

0.2.2

1 release file

0.2.1

1 release file

0.2.0

1 release file

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page