RAG Framework
A modular, extensible Python framework for building Retrieval-Augmented Generation (RAG) pipelines. Plug in your own loaders, embedders, vector stores, and generators — or use the built-in implementations to get started in minutes.
Features
- Modular by design — every component (loader, chunker, embedder, retriever, generator) is an abstract base class you can swap out
- Works out of the box — built-in text/Markdown loaders, fixed-size chunker, in-memory cosine retriever, and placeholder implementations that need no API keys
- Extensible ecosystem — simple contracts mean integrating OpenAI, HuggingFace, ChromaDB, FAISS, or any other tool is just a subclass away
- Batteries-optional — core dependency is
numpyonly; add[pdf],[openai],[chromadb], … as you need them - Fully tested — pytest-based test suite with coverage reporting
- Contributor-friendly — clear abstractions, good first issues, and detailed contributing guide
Architecture
┌─────────────────────────────────────┐
│ RAGPipeline │
└─────────────┬───────────────────────┘
│
┌────────────────────────┼─────────────────────────┐
│ │ │
▼ ▼ ▼
┌───────────────┐ ┌──────────────┐ ┌──────────────────┐
│ DocumentLoader│──────▶│ TextChunker │──────┐ │ │
└───────────────┘ └──────────────┘ │ │ │
(TextFileLoader, (FixedSizeChunker, │ │ │
MarkdownLoader, SentenceChunker, │ │ │
PDFLoader*, …) SemanticChunker*) │ │ │
▼ │ │
┌──────────────┐ │
│ Embedder │ │
└──────┬───────┘ │
│ (RandomEmbedder, │
│ OpenAIEmbedder*, │
│ HFEmbedder*) │
▼ │
┌──────────────┐ │
│ Retriever │ │
└──────┬───────┘ │
│ (InMemoryRetriever,│
│ FAISSRetriever*, │
│ ChromaRetriever*) │
▼ │
┌──────────────┐ │
│ Generator │◀────────────┘
└──────────────┘
(EchoGenerator,
OpenAIGenerator*,
AnthropicGenerator*)
* = open contribution opportunity — see .github/GOOD_FIRST_ISSUES.md
Installation
Note: The package is not yet on PyPI. Install directly from GitHub:
# Core (numpy only)
pip install git+https://github.com/adaumsilva/RAG-framework.git
# With PDF support
pip install "ragframework[pdf] @ git+https://github.com/adaumsilva/RAG-framework.git"
# With OpenAI support
pip install "ragframework[openai] @ git+https://github.com/adaumsilva/RAG-framework.git"
# With HuggingFace embeddings
pip install "ragframework[huggingface] @ git+https://github.com/adaumsilva/RAG-framework.git"
# Everything
pip install "ragframework[all] @ git+https://github.com/adaumsilva/RAG-framework.git"
Once published to PyPI, installation will simplify to pip install ragframework.
Quick Start
from ragframework import RAGPipeline, RAGConfig
from ragframework.document import TextFileLoader, FixedSizeChunker
from ragframework.embeddings import RandomEmbedder # swap for OpenAIEmbedder
from ragframework.retriever import InMemoryRetriever # swap for FAISSRetriever
from ragframework.generator import EchoGenerator # swap for OpenAIGenerator
pipeline = RAGPipeline(
loader=TextFileLoader(),
chunker=FixedSizeChunker(chunk_size=512, chunk_overlap=64),
embedder=RandomEmbedder(dim=384),
retriever=InMemoryRetriever(),
generator=EchoGenerator(),
config=RAGConfig(top_k=5),
)
# Ingest a document
n_chunks = pipeline.ingest("my_document.txt")
print(f"Indexed {n_chunks} chunks")
# Query
response = pipeline.query("What is this document about?")
print(response.answer)
for chunk in response.source_chunks:
print(f" Source: {chunk.metadata.get('source')} — {chunk.content[:80]}…")
Implementing your own component
from ragframework.base import Embedder
class MyEmbedder(Embedder):
def embed(self, texts: list[str]) -> list[list[float]]:
# call your embedding API / model here
...
That's it — plug MyEmbedder() into RAGPipeline and everything else stays the same.
Roadmap
Community contributions are the engine that drives this roadmap. Pick up a Good First Issue and open a PR!
| Priority | Item | Status |
|---|---|---|
| High | PDF document loader | Open |
| High | DOCX document loader | Open |
| High | OpenAI embeddings integration | Open |
| High | HuggingFace Sentence Transformers | Open |
| High | OpenAI / Anthropic generator | Open |
| Medium | FAISS vector store retriever | Open |
| Medium | ChromaDB retriever integration | Open |
| Medium | Semantic / recursive chunker | Open |
| Medium | Async pipeline support | Open |
| Low | Jupyter notebook examples | Open |
Contributing
Contributions are what make open source great. Please read CONTRIBUTING.md before opening a PR.
- Fork the repo and create a branch:
git checkout -b feat/my-feature - Install dev dependencies:
pip install -e ".[dev]" - Write your code and tests
- Run the suite:
pytest tests/ -v - Open a pull request
License
Distributed under the MIT License. See LICENSE for more information.
Acknowledgements
Architecture inspired by RAG-Anything by HKUDS.
Release files for ragframework 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ragframework-0.2.0.tar.gz | 21.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ragframework-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 47.4 kB
Release files / ragframework-0.2.0.tar.gz
| Download URL | ragframework-0.2.0.tar.gz |
|---|---|
| Size | 21.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0e715959f1b71a73a6a60166a2d7937ed312b76df2ff8ca4874a2dced07e6411
|
|
BLAKE2b-256 checksum How to use checksums |
91c136f22c7dc3a9cd57e5ba0aedf6a365ea15c0c610d820c94c3b9e58c95f0d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency logRelease files / ragframework-0.2.0-py3-none-any.whl
| Download URL | ragframework-0.2.0-py3-none-any.whl |
|---|---|
| Size | 25.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
244b53e2c9316d4a96204fb4bcf41a2e61ccde08a0dcb10fb60d2f27ef3e01fe
|
|
BLAKE2b-256 checksum How to use checksums |
3b53960f59e8283c9c045ef3a9b341cc72f2d538b66ad307d48b578f882b19f3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency log