Skip to main content

A local citation-aware research and teaching assistant for technical PDFs

Project description

Multimodal AI Research & Teaching Assistant

Upload a research paper PDF, index it locally, and get answers grounded in the document with exact page citations — no cloud API required. The system runs entirely on your machine using Ollama for language models and FAISS for vector search. This repository is also a 10-notebook tutorial series covering every component end-to-end, from PDF ingestion to evaluation.

Features

  • Upload a PDF and build a searchable FAISS index
  • Ask questions and receive answers with page citations
  • Five teaching modes: beginner, graduate, interview prep, quiz generation, and figure explanation
  • Optional figure captioning with a vision-language model
  • Source-scoped retrieval — Explain figure mode constrains search to the selected document
  • Duplicate upload detection — re-uploading the same PDF returns a cached response instantly
  • OpenTelemetry tracing — per-request spans with retrieval scores, token counts, and latency
  • Fully local: Ollama + Hugging Face, no API keys required
  • Production-style architecture with typed modules, API/UI separation, Docker, testing, CI, evaluation, and observability

Prerequisites

Required:

Text model:

ollama pull llama3.2:latest

Optional — enables figure and image captioning (~6 GB):

ollama pull qwen2.5vl:latest

The vision model is not required for text-only PDF question answering. When it is not installed, the Explain figure mode falls back to text-based retrieval and shows an in-app prompt with the install command.

Quick start

cp .env.example .env
ollama pull llama3.2:3b
docker compose up --build

Optional — enable figure captioning:

ollama pull qwen2.5vl:7b

Open:

Demo workflow:

  1. In the sidebar, upload data/sample/attention_is_all_you_need.pdf
  2. Click Index document
  3. Ask: "What problem does self-attention solve?"

Python package

The core library is distributed as mrta-rag on PyPI:

# Core only (config, schemas, LLM client, prompts)
pip install mrta-rag

# Add PDF ingestion
pip install "mrta-rag[pdf]"

# Add chunking, embeddings, and FAISS vector search
pip install "mrta-rag[retrieval]"

# Full install (matches the Docker environment)
pip install "mrta-rag[all]"
import mrta

print(mrta.__version__)   # 0.1.0

# Core API available after pip install mrta-rag:
from mrta import rag_query, LLMClient, Settings, load_prompt

# Requires mrta-rag[pdf]:
from mrta import load_pdf, chunk_pdf

# Requires mrta-rag[retrieval]:
from mrta import Embedder, VectorStore

Development

Local setup (without Docker)

python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
pip install -e ".[all]"

Run the backend and frontend in separate terminals:

uvicorn apps.api.main:app --reload --port 8000
streamlit run apps/streamlit/app.py

Environment switching

Config is loaded from configs/{MRTA_ENV}.yaml, with env vars and .env taking priority:

MRTA_ENV=test pytest   # lighter models, fast CI
MRTA_ENV=dev pytest    # full dev config

Tests

pytest
pytest tests/unit/        # unit tests only
pytest tests/evaluation/  # retrieval gate tests

Observability

Tracing is controlled by three .env variables:

ENABLE_TRACING=true           # activate the OTEL SDK
OTEL_CONSOLE_EXPORTER=true    # print spans to stdout (local dev)
OTEL_SERVICE_NAME=mrta
OTEL_EXPORTER_OTLP_ENDPOINT=  # set to export to Jaeger / Tempo

With console export enabled, each /ask call prints a span to the API logs showing retrieval scores, cited sources, token counts, and end-to-end latency.

Linting and type checking

ruff check src/ tests/ apps/
black --check src/ tests/ apps/
.venv311/bin/mypy src/ apps/ --ignore-missing-imports

Note: Use a Python 3.11 virtual environment for mypy. The default .venv uses Python 3.14, whose NumPy stubs use syntax that mypy rejects when python_version = "3.11" is set. CI uses Python 3.11 and passes.

Tutorial notebooks

jupyter lab notebooks/

Two parallel versions of the 10-part series:

  • notebooks/production/ — imports from src/mrta/; the reference implementation
  • notebooks/tutorials/ — every function defined inline; use for learning
# Phase Topic
0 Setup Repo scaffold, Ollama, Hugging Face
1 Ingestion PyMuPDF text and image extraction
2 Chunking Fixed, recursive, and semantic strategies
3 Embeddings sentence-transformers + FAISS index
4 RAG End-to-end pipeline with citations
5 Backend FastAPI endpoints and Pydantic schemas
6 Frontend Streamlit upload, ask, cite
7 Multimodal Figure extraction, CLIP, VLM captioning
8 Teaching modes Prompt templates for different audiences
9 Evaluation DeepEval metrics, structured logs, Docker

Architecture and design decisions

Limitations

  • Math is rendered as text; LaTeX-aware parsing would improve recall on equation-heavy papers.
  • Table extraction is basic; ColPali or unstructured would help for table-heavy domains.
  • Reranking is a stub; adding a cross-encoder (bge-reranker-base) is a one-day improvement.
  • No multi-document graph reasoning yet — a clear next step toward an "agentic" research assistant.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mrta_rag-0.1.0.tar.gz (1.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mrta_rag-0.1.0-py3-none-any.whl (31.5 kB view details)

Uploaded Python 3

File details

Details for the file mrta_rag-0.1.0.tar.gz.

File metadata

  • Download URL: mrta_rag-0.1.0.tar.gz
  • Upload date:
  • Size: 1.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for mrta_rag-0.1.0.tar.gz
Algorithm Hash digest
SHA256 2746a5b96b71323ffff1155735f7e2f5ba9de31109ce06c24f5d87c3566aba08
MD5 6c0755fb06cd024e58d9ce67a2eafa91
BLAKE2b-256 08eb256952795783ae79ec0c13d854c0b3a72e98976b0b674497b79da17b2d36

See more details on using hashes here.

File details

Details for the file mrta_rag-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: mrta_rag-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 31.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for mrta_rag-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a7909658c78cc29d792779c0fa3a6f05533453741e812289a0e6daee1b5e6191
MD5 eac7d24a88bf332e19f5bc85dee2cbd8
BLAKE2b-256 6830c3e71dbdd61ea71088a8f79ea480ad7fe65af2cec7c9fed153dfd2c7d407

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page