Skip to main content

🚀 RocketRAG

Fast, efficient, minimal, extendible and elegant RAG system

RocketRAG is a high-performance Retrieval-Augmented Generation (RAG) system designed with a focus on speed, simplicity, and extensibility. Built on top of state-of-the-art libraries, it provides both CLI and web server capabilities for seamless integration into any workflow.

https://github.com/user-attachments/assets/1bd7cc50-9eac-4402-80bd-42933ac35ca3

🎯 Mission

RocketRAG aims to be the fastest and most efficient RAG library while maintaining:

  • Minimal footprint - Clean, lightweight codebase
  • Maximum extensibility - Pluggable architecture for all components
  • Peak performance - Leveraging the best-in-class libraries
  • Ease of use - Simple CLI and API interfaces

⚡ Performance-First Architecture

RocketRAG is built on top of cutting-edge, performance-optimized libraries:

🚀 Quick Start

Installation

Using pip

pip install rocketrag

Using uvx (recommended for CLI usage)

# Run directly without installation
uvx rocketrag --help

# Or install globally
uvx install rocketrag

Basic Usage

from rocketrag import RocketRAG

rag = RocketRAG("./data") # Path do your data (supports PDF, TXT, MD, etc.)
rag.prepare() # Construct vector database

# Ask questions
answer, sources = rag.ask("What is the main topic of the documents?")
print(answer)

CLI Usage

# Prepare documents from a directory
rocketrag prepare --data-dir ./documents

# Ask questions via CLI
rocketrag ask "What are the key findings?"

# Start web server
rocketrag server --port 8000

Using uvx (no installation required)

# Same commands work with uvx
uvx rocketrag prepare --data-dir ./documents
uvx rocketrag ask "What are the key findings?"
uvx rocketrag server --port 8000

# Run as module
uvx --from rocketrag python -m rocketrag --help

🏗️ Architecture

RocketRAG follows a modular, plugin-based architecture:

┌─────────────────┐    ┌─────────────────┐    ┌─────────────────┐
│   Document      │    │    Chunking     │    │   Vectorization │
│   Loaders       │───▶│   (Chonkie)     │───▶│ (SentenceTransf)│
│  (Kreuzberg)    │    │                 │    │                 │
└─────────────────┘    └─────────────────┘    └─────────────────┘
                                                        │
┌─────────────────┐    ┌─────────────────┐             │
│      LLM        │    │   Vector DB     │◀────────────┘
│ (llama-cpp-py)  │◀───│ (Milvus Lite)   │
│                 │    │                 │
└─────────────────┘    └─────────────────┘

Core Components

  • BaseLoader: Pluggable document loading (PDF, TXT, MD, etc.)
  • BaseChunker: Configurable chunking strategies (semantic, recursive, etc.)
  • BaseVectorizer: Flexible embedding models
  • BaseLLM: Swappable language models
  • MilvusLiteDB: High-performance vector storage and retrieval

🔧 Configuration

Custom Components

from rocketrag import RocketRAG
from rocketrag.vectors import SentenceTransformersVectorizer
from rocketrag.chonk import ChonkieChunker
from rocketrag.llm import LLamaLLM
from rocketrag.loaders import KreuzbergLoader

# Configure high-performance components
vectorizer = SentenceTransformersVectorizer(
    model_name="minishlab/potion-multilingual-128M"  # Fast multilingual model
)

chunker = ChonkieChunker(
    method="semantic",  # Semantic chunking for better context
    embedding_model="minishlab/potion-multilingual-128M",
    chunk_size=512
)

llm = LLamaLLM(
    repo_id="unsloth/gemma-3n-E2B-it-GGUF",
    filename="*Q8_0.gguf"  # Quantized for speed
)

loader = KreuzbergLoader()  # Ultra-fast document processing

rag = RocketRAG(
    vectorizer=vectorizer,
    chunker=chunker,
    llm=llm,
    loader=loader
)

CLI Configuration

# Custom chunking strategy
rocketrag prepare \
  --chonker chonkie \
  --chonker-args '{"method": "semantic", "chunk_size": 512}' \
  --vectorizer-args '{"model_name": "all-MiniLM-L6-v2"}'

# Custom LLM for inference
rocketrag ask "Your question" \
  --repo-id "microsoft/DialoGPT-medium" \
  --filename "*.gguf"

🌐 Web Server

RocketRAG includes a FastAPI-based web server with OpenAI-compatible endpoints:

# Start server
rocketrag server --port 8000 --host 0.0.0.0

API Endpoints

  • GET / - Interactive web interface
  • POST /ask - Question answering
  • POST /ask/stream - Streaming responses
  • GET /chat - Chat interface
  • GET /browse - Document browser
  • GET /visualize - Vector visualization
  • GET /health - Health check

Example API Usage

import requests

response = requests.post(
    "http://localhost:8000/ask",
    json={"question": "What are the main findings?"}
)

result = response.json()
print(result["answer"])
print(result["sources"])

🎨 Features

Core Features

  • ⚡ Ultra-fast document processing with Kreuzberg
  • 🧠 Semantic chunking with Chonkie and model2vec
  • 🔍 High-performance vector search with Milvus Lite
  • 🤖 Optimized LLM inference with llama-cpp-python
  • 📊 Rich CLI interface with progress bars and formatting
  • 🌐 Web server with interactive UI
  • 🔌 Pluggable architecture for easy customization

Advanced Features

  • 📈 Vector visualization for debugging and analysis
  • 📚 Document browsing interface
  • 💬 Streaming responses for real-time interaction
  • 🔄 Batch processing for large document sets
  • 📝 Metadata preservation throughout the pipeline
  • 🎯 Context-aware chunking for better retrieval

🛠️ Development

Installation for Development

git clone https://github.com/yourusername/rocketrag.git
cd rocketrag
pip install -e ".[dev]"

Running Tests

pytest tests/

Code Quality

ruff check .
ruff format .

📊 Performance

RocketRAG is designed for speed:

  • Document Loading: 10x faster with Kreuzberg's optimized parsers
  • Chunking: Semantic chunking with model2vec for superior context preservation
  • Vectorization: Optimized batch processing with sentence-transformers
  • Retrieval: Sub-millisecond vector search with Milvus Lite
  • Generation: GGUF quantization for 4x faster inference

🤝 Contributing

We welcome contributions! RocketRAG's modular architecture makes it easy to:

  • Add new document loaders
  • Implement custom chunking strategies
  • Integrate different embedding models
  • Support additional LLM backends
  • Enhance the web interface

🙏 Acknowledgments

RocketRAG builds upon the excellent work of:

Metadata

Release files for rocketrag 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rocketrag 0.1.5
File Size Uploaded
rocketrag-0.1.5.tar.gz 47.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rocketrag 0.1.5
File Interpreter ABI Platform
rocketrag-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 88.3 kB

Release files / rocketrag-0.1.5.tar.gz

Download URL rocketrag-0.1.5.tar.gz
Size 47.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f8234680515192e877c23bbb7a46909811e3dd075c210f59087f92d5cc2e4fab
BLAKE2b-256 checksum
How to use checksums
647ed9a98fd700d1b53410142678e9da58da3dfce8046e3d02a628d62584364e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.21

Release files / rocketrag-0.1.5-py3-none-any.whl

Download URL rocketrag-0.1.5-py3-none-any.whl
Size 40.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d89b9798a0b8d0d1654b6e76a1d737ce97be77523935b2a6496aadd25a6df965
BLAKE2b-256 checksum
How to use checksums
16347190748148eecc54d6f1feb6fa98a22e338165e7368f8774ccfec958242e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.21

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page