Skip to main content

Rago

CI Python Versions Package Version License Discord

Rago is a lightweight framework for RAG.

Features

  • Vector Database support
    • FAISS
  • Retrieval features
    • Support PDF extraction via Langchain
  • Augmentation (Embedding + Vector Database Search)
    • Support for Sentence Transformer (Hugging Face)
    • Support for Open AI
    • Support for SpaCy
  • Generation (LLM)
    • Support for Hugging Face
    • Support for Llama (Hugging Face)
    • Support for OpenAI
    • Support for Gemini

Roadmap

1. Add new Backends

As noted in several GitHub issues, our initial goal is to support as many backends as possible. This approach will provide valuable insights into user needs and inform the structure for the next phase.

2. Declarative API for Rago

Objective

To simplify and streamline the user experience in configuring RAG by introducing a declarative, composable API—similar to how Plotnine or Altair allows users to build visualizations.

Overview

The current procedural approach in Rago requires users to instantiate and connect individual components (retrieval, augmentation, generation, etc.) manually. This can become cumbersome as support for multiple backends grows. We propose a new declarative interface that lets users define their entire RAG steps in a single, fluent expression using operator overloading.

Proposed Syntax Example

from pathlib import Path

from rago import Rago, Retrieval, Augmented, Generation, DB, Cache

datasource = ...

rag = (
    Rago()
    | DB(backend="faiss")
    | Cache(backend="file", target_dir=Path(".rago-cache"))
    | Retrieval(backend="string")
    | Augmented(
        backend="openai",
        model_name="text-embedding-3-small",
        top_k=5,
    )
    | Generation(
        backend="openai",
        model_name="gpt-4o-mini",
        prompt_template="Question: {query}\nContext: {context}\nAnswer:"
    )
)

result = rag.run(query="What is the capital of France?", source=datasource)
print(result.result)

Key Benefits

  • Intuitive Composition: Users can build complex pipelines by simply adding layers together.
  • Modularity: Each component is encapsulated, making it easy to swap or extend backends without altering the overall architecture.
  • Reduced Boilerplate: The declarative syntax minimizes the need for repetitive setup code, focusing on the "what" rather than the "how."
  • Enhanced Readability: The pipeline’s structure becomes immediately clear, promoting easier maintenance and collaboration.

Implementation Plan

  1. Define Base Classes: Develop abstract base classes for each component (DB, Cache, Retrieval, Augmented, Generation) to standardize interfaces and facilitate future extensions.
  2. Operator Overloading: Implement the __or__ method in the main Rago class to allow chaining of components, effectively building the pipeline through a fluent interface.
  3. Configuration and Defaults: Integrate sensible defaults and validation (using tools like Pydantic) so that users can override only when necessary.
  4. Documentation and Examples: Provide comprehensive documentation and examples to illustrate the new declarative syntax and usage scenarios.

Installation

If you want to install it for cpu only, you can run:

$ pip install rago[cpu]

But, if you want to install it for gpu (cuda), you can run:

$ pip install rago[gpu]

Setup

Llama 3

In order to use a Llama model, visit its page on Hugging Face and request access via its form, for example: https://huggingface.co/meta-llama/Llama-3.2-1B.

After you are granted access to the desired model, you will be able to use it with Rago.

You will also need to provide a Hugging Face token in order to download the models locally, for example:

from rago import Augmented, Generation, Rago, Retrieval

# For Gated LLMs
HF_TOKEN = 'YOUR_HUGGING_FACE_TOKEN'

animals_data = [
    "The Blue Whale is the largest animal ever known to have existed, even "
    "bigger than the largest dinosaurs.",
    "The Peregrine Falcon is renowned as the fastest animal on the planet, "
    "capable of reaching speeds over 240 miles per hour.",
    "The Giant Panda is a bear species endemic to China, easily recognized by "
    "its distinctive black-and-white coat.",
    "The Cheetah is the world's fastest land animal, capable of sprinting at "
    "speeds up to 70 miles per hour in short bursts covering distances up to "
    "500 meters.",
    "The Komodo Dragon is the largest living species of lizard, found on "
    "several Indonesian islands, including its namesake, Komodo.",
]

rag = (
    Rago()
    | Retrieval(backend='string')
    | Augmented(
        backend='sentence_transformers',
        model_name='paraphrase-MiniLM-L12-v2',
        top_k=2,
    )
    | Generation(
        backend='llama',
        model_name='meta-llama/Llama-3.2-1B',
        api_key=HF_TOKEN,
    )
)

rag.prompt('What is the fastest animal on Earth?', source=animals_data)

Ollama

For testing the generation with Ollama, run first the following commands:

$ ollama pull llama3.2:1b
$ ollama serve

Metadata

Release files for rago 0.14.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rago 0.14.5
File Size Uploaded
rago-0.14.5.tar.gz 38.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rago 0.14.5
File Interpreter ABI Platform
rago-0.14.5-py3-none-any.whl Python 3 none any Details

Total release size: 87.0 kB

Release files / rago-0.14.5.tar.gz

Download URL rago-0.14.5.tar.gz
Size 38.7 kB
Tags Source
SHA-256 checksum
How to use checksums
44563353e7d7769bedbb6624083067185d8e1b815feb78a267647fcab7dbcea8
BLAKE2b-256 checksum
How to use checksums
6ab1daca2b1499f236de289550e4443281d8374fd956a40c4e53e91ec9aabec5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.15

Release files / rago-0.14.5-py3-none-any.whl

Download URL rago-0.14.5-py3-none-any.whl
Size 48.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d3cb9e55ee7c79f72f03fa8fcb7506070107b661ce15a5af5ea7506dae5b6b30
BLAKE2b-256 checksum
How to use checksums
4dc337b73fe409e51b7e42d7d5f7ca36cbc31a96b0dfb4ebe4409224b3718fac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.15

Release history Release notifications | RSS feed

This release

0.14.5 This release

2 release files

0.14.3

2 release files

0.14.2

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.2

2 release files

0.11.1

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page