Skip to main content

Ragora

PyPI version Python versions License GitHub stars

Build smarter, grounded, and transparent AI with Ragora.

Ragora is an open-source framework for building Retrieval-Augmented Generation (RAG) systems that connect your language models to real, reliable knowledge. It provides a clean, composable interface for managing knowledge bases, document retrieval, and grounding pipelines, so your AI can reason with context instead of guesswork.

The name Ragora blends RAG with the ancient Greek Agora, the public square where ideas were exchanged, debated, and refined. In the same spirit, Ragora is the meeting place of data and dialogue, where your information and your AI come together to think.

✨ Key Features

  • 📄 Specialized Document Processing: Native support for LaTeX parsing and email handling with more formats coming
  • 🏗️ Clean Architecture: Three-layer design (DatabaseManager → VectorStore → Retriever) for maintainability
  • 🔍 Flexible Search: Vector, keyword, and hybrid search modes for optimal retrieval
  • 🧩 Composable Components: Use high-level APIs or build custom pipelines with low-level components
  • ⚡ Performance Optimized: Batch processing, GPU acceleration, and efficient vector search with Weaviate
  • 🔒 Privacy-First: Run completely local with sentence-transformers and Weaviate

🚀 Installation

pip install ragora

Prerequisites

You need a Weaviate instance running. Download the pre-configured Ragora database server:

# Download from GitHub releases
wget https://github.com/Vahidlari/aiApps/releases/download/v<x.y.z>/database_server-<x.y.z>.tar.gz

# Extract and start
tar -xzf database_server-<x.y.z>.tar.gz
cd database-server
./database-manager.sh start

Update <x.y.z> with the actual package version- For example use 1.0.0 for version v1.0.0. The database server is a zero-dependency solution (only requires Docker) that works on Windows, macOS, and Linux.

Document Processing

Process LaTeX documents with specialized handling:

from ragora.core import DocumentPreprocessor, DataChunker

# Parse LaTeX with citations
preprocessor = DocumentPreprocessor()
document = preprocessor.parse_latex(
    "paper.tex",
    bibliography_path="references.bib"
)

# Chunk with configurable size and overlap using new API
from ragora import DataChunker, ChunkingContextBuilder

chunker = DataChunker()
context = ChunkingContextBuilder().for_document().build()
chunks = chunker.chunk(document.content, context)

🔍 Search Modes

Ragora supports three search strategies:

from ragora import SearchStrategy

# Semantic search (best for conceptual queries)
results = kbm.search("explain machine learning", strategy=SearchStrategy.SIMILAR)

# Keyword search (best for exact terms)
results = kbm.search("Schrödinger equation", strategy=SearchStrategy.KEYWORD)

# Hybrid search (recommended - combines both)
results = kbm.search("neural networks", strategy=SearchStrategy.HYBRID, alpha=0.7)

🎯 Use Cases

  • 📖 Academic Research: Build knowledge bases from scientific papers and LaTeX documents
  • 📝 Documentation Search: Create searchable knowledge bases from technical documentation
  • 🤖 AI Assistants: Ground LLM responses in your specific domain knowledge
  • 💬 Question Answering: Build Q&A systems over your document collections
  • 🔬 Literature Review: Efficiently search and synthesize information from research papers

📖 Documentation & Examples

  • Tool Documentation: Overal tool documentation, including instructions to get started
  • API Reference: Complete API documentation
  • Examples Directory: Working code examples
    • basic_usage.py: Basic usage examples and getting started
    • advanced_usage.py: Advanced features and custom pipelines
    • email_usage_examples.py: Email integration examples

📊 Requirements

  • Python: 3.11 or higher
  • Weaviate: 1.22.0 or higher (for vector storage)
  • Dependencies: See requirements.txt

🤝 Contributing

We welcome contributions! Please see our Contributing Guidelines for:

  • Setting up your development environment
  • Code style and standards
  • Writing tests
  • Submitting pull requests

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

📮 Contact

For questions, feedback, or collaboration opportunities:

  • Open an issue on GitHub
  • Start a discussion in GitHub Discussions
  • Contact the maintainers directly

Build smarter, grounded, and transparent AI with Ragora.

Metadata

Release files for ragora 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ragora 1.3.0
File Size Uploaded
ragora-1.3.0.tar.gz 182.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ragora 1.3.0
File Interpreter ABI Platform
ragora-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 277.4 kB

Release files / ragora-1.3.0.tar.gz

Download URL ragora-1.3.0.tar.gz
Size 182.6 kB
Tags Source
SHA-256 checksum
How to use checksums
8527fdd5a8cb5ea2a526354448e53363a6e9bfc489cf9fd85af4db6f19998f83
BLAKE2b-256 checksum
How to use checksums
e9f930eef323dbbef04c68de7974e19097ba9cff12a767a104241bfad32a298b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release files / ragora-1.3.0-py3-none-any.whl

Download URL ragora-1.3.0-py3-none-any.whl
Size 94.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c1cabd31091d5ae409bafa2fd5c85ef789336318e57cc82bb341e07b0ca78cfe
BLAKE2b-256 checksum
How to use checksums
6a5c0899f3bb62e22873f73b0f5ee2b166ef504def4b3991c4d9e599fade1d71
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page