Skip to main content

RAG PDF Chatbot

RAG PDF Chatbot Logo

A Professional, Enterprise-Grade Retrieval-Augmented Generation System for PDF Documents

License: MIT Python 3.8+ Code Style: Black

🎯 Problem Statement

Organizations struggle with extracting actionable insights from large collections of PDF documents. Traditional search methods fail to provide contextual, accurate answers to complex questions. RAG PDF Chatbot solves this by combining:

  • Document Retrieval: Find relevant information from PDF collections
  • Contextual Understanding: Use LLM to understand and synthesize information
  • Natural Language Interface: Ask questions in plain English and get precise answers

🏗️ Architecture

graph TD
    A[PDF Documents] --> B[Document Processor]
    B --> C[Vector Store]
    C --> D[Retriever]
    D --> E[RAG Chain]
    E --> F[LLM]
    F --> G[Answer]
    G --> H[User]
    H -->|Question| E

Key Components

  1. Document Processor: Loads and chunks PDF documents
  2. Vector Store: Stores document embeddings for efficient retrieval
  3. Retriever: Finds relevant documents for a given question
  4. RAG Chain: Combines retrieved context with LLM for answer generation
  5. LLM Interface: Uses Ollama to run local language models

🛠️ Tech Stack

  • Core: Python 3.8+
  • Document Processing: LangChain, PyMuPDF
  • Embeddings: Ollama (nomic-embed-text)
  • Vector Store: FAISS
  • LLM: Ollama (llama3.2:3b)
  • Configuration: Python dataclasses + environment variables
  • Testing: pytest

🚀 Quick Start

Prerequisites

  • Python 3.8+
  • Ollama running locally with required models
  • PDF documents in the rag-dataset/ directory

Installation

# Clone the repository
git clone https://github.com/your-org/rag-pdf-chatbot.git
cd rag-pdf-chatbot

# Create virtual environment (recommended)
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

# Set up environment variables
cp .env.example .env
# Edit .env with your configuration

Running the Application

# Basic usage
python -m src.main --help

# Ask a specific question
python -m src.main --question "What are the benefits of BCAA supplements?"

# Interactive mode
python -m src.main --interactive

# Rebuild vector store
python -m src.main --rebuild --interactive

📂 Project Structure

rag-pdf-chatbot/
├── src/                  # Core application code
│   ├── __init__.py       # Package initialization
│   ├── config.py         # Configuration management
│   ├── document_processor.py  # Document loading and processing
│   ├── vector_store.py   # Vector storage and retrieval
│   ├── rag_chain.py      # RAG pipeline implementation
│   └── main.py           # Main application entry point
├── tests/                # Unit and integration tests
├── docs/                 # Architecture and design documentation
├── config/               # Configuration files
├── scripts/              # Automation and utility scripts
├── .env.example          # Environment variable template
├── .gitignore            # Git ignore patterns
├── README.md             # This file
└── requirements.txt      # Python dependencies

🔧 Configuration

The application uses environment variables for configuration. See .env.example for all available options:

# Ollama Configuration
OLLAMA_BASE_URL=http://localhost:11434
EMBEDDING_MODEL=nomic-embed-text
LLM_MODEL=llama3.2:3b

# Document Processing
DATASET_PATH=rag-dataset
CHUNK_SIZE=1000
CHUNK_OVERLAP=100

# Vector Store
VECTOR_STORE_PATH=health_supplements
SAVE_VECTOR_STORE=true

# Retrieval
RETRIEVAL_TYPE=mmr
RETRIEVAL_K=3
RETRIEVAL_FETCH_K=100
RETRIEVAL_LAMBDA=1.0

🧪 Testing

# Run all tests
pytest tests/

# Run specific test
pytest tests/test_document_processor.py

# Run with coverage
pytest --cov=src tests/

📖 Documentation

🤝 Contributing

We welcome contributions! Please see CONTRIBUTING.md for guidelines.

📜 License

This project is licensed under the MIT License - see the LICENSE file for details.

🎯 Value Proposition

For Developers:

  • Clean, modular architecture following SOLID principles
  • Easy to extend and customize
  • Comprehensive documentation and examples

For Organizations:

  • Extract insights from PDF documents efficiently
  • Reduce manual document review time
  • Improve knowledge discovery and decision making

For Recruiters:

  • Professional, enterprise-grade codebase
  • Follows best practices for security and maintainability
  • Demonstrates advanced Python and AI/ML skills

🔒 Security

This project follows GitGuardian security standards:

  • No hardcoded secrets
  • Environment variable configuration
  • Secure dependency management
  • Regular security audits

📞 Support

For issues, questions, or feature requests, please open an issue on GitHub.


Built with ❤️ for developers, by developers.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rag_pdf_chatbot-1.0.0.tar.gz (14.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

rag_pdf_chatbot-1.0.0-py3-none-any.whl (12.0 kB view details)

Uploaded Python 3

File details

Details for the file rag_pdf_chatbot-1.0.0.tar.gz.

File metadata

  • Download URL: rag_pdf_chatbot-1.0.0.tar.gz
  • Upload date:
  • Size: 14.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for rag_pdf_chatbot-1.0.0.tar.gz
Algorithm Hash digest
SHA256 ec21cdcfc628e4ad3d7dd9d92acd454d4742f347bf043ba84ed0afce5701c7de
MD5 1fdc79499c242fe67c7b17a8482b6da8
BLAKE2b-256 645077544a49ef7ab074ee0d0d170167bbc57c786122fb6317ea3ad5360accad

See more details on using hashes here.

File details

Details for the file rag_pdf_chatbot-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for rag_pdf_chatbot-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f369157fa4f95aacd5ffee17d377543a8405f134e27cd1a0ce8d120e72634918
MD5 49253a45b98f8df77587181678d8dd85
BLAKE2b-256 001b3866a15ea453fc84d72b8a05ced3370aa12b4a1d4d19db01d745ba13899d

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page