Multimodal AI Research & Teaching Assistant
Upload a research paper PDF, index it locally, and get answers grounded in the document with exact page citations — no cloud API required. The system runs entirely on your machine using Ollama for language models and FAISS for vector search. This repository is also a 10-notebook tutorial series covering every component end-to-end, from PDF ingestion to evaluation.
Features
- Upload a PDF and build a searchable FAISS index
- Ask questions and receive answers with page citations
- Five teaching modes: beginner, graduate, interview prep, quiz generation, and figure explanation
- Optional figure captioning with a vision-language model
- Source-scoped retrieval — Explain figure mode constrains search to the selected document
- Duplicate upload detection — re-uploading the same PDF returns a cached response instantly
- OpenTelemetry tracing — per-request spans with retrieval scores, token counts, and latency
- Fully local: Ollama + Hugging Face, no API keys required
- Production-style architecture with typed modules, API/UI separation, Docker, testing, CI, evaluation, and observability
Prerequisites
Required:
- Docker Desktop
- Ollama
- Git
Text model:
ollama pull llama3.2:latest
Optional — enables figure and image captioning (~6 GB):
ollama pull qwen2.5vl:latest
The vision model is not required for text-only PDF question answering. When it is not installed, the Explain figure mode falls back to text-based retrieval and shows an in-app prompt with the install command.
Quick start
cp .env.example .env
ollama pull llama3.2:3b
docker compose up --build
Optional — enable figure captioning:
ollama pull qwen2.5vl:7b
Open:
- UI: http://localhost:8501
- API docs: http://localhost:8000/docs
Demo workflow:
- In the sidebar, upload
data/sample/attention_is_all_you_need.pdf - Click Index document
- Ask: "What problem does self-attention solve?"
Python package
The core library is distributed as mrta-rag on PyPI:
# Core only (config, schemas, LLM client, prompts)
pip install mrta-rag
# Add PDF ingestion
pip install "mrta-rag[pdf]"
# Add chunking, embeddings, and FAISS vector search
pip install "mrta-rag[retrieval]"
# Full install (matches the Docker environment)
pip install "mrta-rag[all]"
import mrta
print(mrta.__version__) # 0.1.0
# Core API available after pip install mrta-rag:
from mrta import rag_query, LLMClient, Settings, load_prompt
# Requires mrta-rag[pdf]:
from mrta import load_pdf, chunk_pdf
# Requires mrta-rag[retrieval]:
from mrta import Embedder, VectorStore
Development
Local setup (without Docker)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[all]"
Run the backend and frontend in separate terminals:
uvicorn apps.api.main:app --reload --port 8000
streamlit run apps/streamlit/app.py
Environment switching
Config is loaded from configs/{MRTA_ENV}.yaml, with env vars and .env
taking priority:
MRTA_ENV=test pytest # lighter models, fast CI
MRTA_ENV=dev pytest # full dev config
Tests
pytest
pytest tests/unit/ # unit tests only
pytest tests/evaluation/ # retrieval gate tests
Observability
Tracing is controlled by three .env variables:
ENABLE_TRACING=true # activate the OTEL SDK
OTEL_CONSOLE_EXPORTER=true # print spans to stdout (local dev)
OTEL_SERVICE_NAME=mrta
OTEL_EXPORTER_OTLP_ENDPOINT= # set to export to Jaeger / Tempo
With console export enabled, each /ask call prints a span to the API logs showing
retrieval scores, cited sources, token counts, and end-to-end latency.
Linting and type checking
ruff check src/ tests/ apps/
black --check src/ tests/ apps/
.venv311/bin/mypy src/ apps/ --ignore-missing-imports
Note: Use a Python 3.11 virtual environment for
mypy. The default.venvuses Python 3.14, whose NumPy stubs use syntax that mypy rejects whenpython_version = "3.11"is set. CI uses Python 3.11 and passes.
Tutorial notebooks
jupyter lab notebooks/
Two parallel versions of the 10-part series:
notebooks/production/— imports fromsrc/mrta/; the reference implementationnotebooks/tutorials/— every function defined inline; use for learning
| # | Phase | Topic |
|---|---|---|
| 0 | Setup | Repo scaffold, Ollama, Hugging Face |
| 1 | Ingestion | PyMuPDF text and image extraction |
| 2 | Chunking | Fixed, recursive, and semantic strategies |
| 3 | Embeddings | sentence-transformers + FAISS index |
| 4 | RAG | End-to-end pipeline with citations |
| 5 | Backend | FastAPI endpoints and Pydantic schemas |
| 6 | Frontend | Streamlit upload, ask, cite |
| 7 | Multimodal | Figure extraction, CLIP, VLM captioning |
| 8 | Teaching modes | Prompt templates for different audiences |
| 9 | Evaluation | DeepEval metrics, structured logs, Docker |
Architecture and design decisions
- Tech stack, system diagram, and repo layout:
docs/architecture/overview.md - Key design decisions (FAISS vs Qdrant, Ollama vs API, etc.):
docs/adr/
Limitations
- Math is rendered as text; LaTeX-aware parsing would improve recall on equation-heavy papers.
- Table extraction is basic; ColPali or
unstructuredwould help for table-heavy domains. - Reranking is a stub; adding a cross-encoder (
bge-reranker-base) is a one-day improvement. - No multi-document graph reasoning yet — a clear next step toward an "agentic" research assistant.
License
This project is licensed under the MIT License. See the LICENSE file for details.
Metadata
Release files for mrta-rag 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mrta_rag-0.1.0.tar.gz | 1.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mrta_rag-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.4 MB
Release files / mrta_rag-0.1.0.tar.gz
| Download URL | mrta_rag-0.1.0.tar.gz |
|---|---|
| Size | 1.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2746a5b96b71323ffff1155735f7e2f5ba9de31109ce06c24f5d87c3566aba08
|
|
BLAKE2b-256 checksum How to use checksums |
08eb256952795783ae79ec0c13d854c0b3a72e98976b0b674497b79da17b2d36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.15
|
Release files / mrta_rag-0.1.0-py3-none-any.whl
| Download URL | mrta_rag-0.1.0-py3-none-any.whl |
|---|---|
| Size | 31.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a7909658c78cc29d792779c0fa3a6f05533453741e812289a0e6daee1b5e6191
|
|
BLAKE2b-256 checksum How to use checksums |
6830c3e71dbdd61ea71088a8f79ea480ad7fe65af2cec7c9fed153dfd2c7d407
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.15
|