Skip to main content

Production-grade self-hosted document ingestion and retrieval service

Project description

RAG Agent

Production-grade, self-hosted document ingestion and retrieval service.

Features

  • Multi-format parsing: PDF, DOCX, TXT, Markdown with smart layout preservation
  • Chunking strategies: Recursive character, Markdown headers
  • Image extraction and LLM description: Makes visual content searchable
  • OCR fallback: LLM vision for scanned pages
  • Vector storage: Milvus with cosine similarity search
  • Hybrid retrieval: BM25 keyword search + vector fusion (RRF)
  • Cross-encoder reranking: Sentence Transformers for relevance scoring
  • ARQ task queue: Background processing with retry and backoff
  • SSE status streaming: Real-time ingestion progress
  • Pluggable connectors: Local filesystem (ready), S3/Google Drive (pluggable)
  • Deduplication: Content hash + source path matching
  • Batched embedding: Configurable batch size with retry
  • Web dashboard: Responsive dark-themed UI for all operations
  • uv Python manager: Fast dependency installation and management
  • Makefile: Simplified task management

Architecture

Upload → Validate → Store → Track (DB) → Queue (ARQ)
  ┌─── Worker ───────────────────────────────────────┐
  │ Parse → Describe images → Chunk → Dedup → Embed → Store (Milvus)
  └──────────────────────────────────────────────────┘
                ↓
          SSE Status Events
                ↓
           Query → Search
                ↓
          Web Dashboard  ←── You are here

Web UI

The project includes a responsive dark-themed web dashboard built with Vue 3 (served as static files from the FastAPI application).

Pages:

Page Description
Dashboard System health, collection stats, recent documents
Documents Upload (drag & drop), list/filter, delete, retry, download
Collections Create, browse, delete vector collections
Search Semantic search with reranker, multi-collection mode, score visualization

Access the UI at http://localhost:8100/ (redirects to /ui/).

The frontend is served directly by the API server — no separate build step or dev server needed. Source lives in the frontend/ directory.

Quick Start

Two workflows are available:

🐳 Full Container Stack (production-like)

# Install uv (https://github.com/astral-sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install dependencies
make setup

# Create .env from example
cp .env.example .env
# Edit .env with your settings

# Start everything (app + data infra in containers)
make up

# Wait for services to be ready
make wait

# Upload a document
curl -X POST http://localhost:8100/api/v1/documents/upload \
  -F "file=@example.pdf"

# Search
curl -X POST http://localhost:8100/api/v1/search \
  -H "Content-Type: application/json" \
  -d '{"query": "What is the revenue?", "limit": 5}'

⚡ Dev-Fast (app on host, hot-reload)

Run the Python code directly on your machine for instant feedback — only the data stores (Postgres, Valkey, Milvus) run in containers.

# Install uv (https://github.com/astral-sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install dependencies
make setup

# Start ONLY data infrastructure containers
make dev-up

# Create media directory + run database migrations
make dev-setup

# (Terminal 1) Start API with hot-reload
make dev-fast

# (Terminal 2) Start background worker
make dev-fast-worker

# Open http://localhost:8100/ in your browser

# When done, stop data containers
make dev-down

Changes to Python files are picked up instantly — no Docker rebuilds needed.

Manual uv Commands

# Install dependencies
uv sync --all-extras

# Run tests
uv run pytest tests/ -v

# Start API server (uses .env)
uv run uvicorn rag_agent.app:create_app --reload --factory

# Start ARQ worker
uv run arq rag_agent.worker.settings.WorkerSettings

API Endpoints

Health

  • GET /health — Liveness
  • GET /ready — Readiness with dependency checks
  • GET /live — Minimal liveness

Collections

  • GET /api/v1/collections — List collections
  • POST /api/v1/collections?name=... — Create collection
  • GET /api/v1/collections/{name} — Collection stats
  • DELETE /api/v1/collections/{name} — Drop collection

Documents

  • POST /api/v1/documents/upload — Upload file (multipart)
  • GET /api/v1/documents — List tracked documents
  • GET /api/v1/documents/{id} — Document detail
  • DELETE /api/v1/documents/{id} — Delete (cascade)
  • POST /api/v1/documents/{id}/retry — Re-queue failed ingestion
  • GET /api/v1/documents/{id}/download — Download original

Search

  • POST /api/v1/search — Vector search
  • POST /api/v1/search/multi — Multi-collection search
  • GET /api/v1/collections/{name}/documents/{id} — Search within a document

Sync & Connectors

  • POST /api/v1/sync — Trigger directory sync
  • GET /api/v1/sync/logs — Sync history
  • GET /api/v1/connectors — Available connectors
  • GET /api/v1/status — SSE stream for progress events

Configuration

See .env.example for all environment variables.

Key settings:

  • EMBEDDING_BASE_URL — OpenAI-compatible embedding endpoint
  • EMBEDDING_MODEL — Model name (e.g., all-MiniLM-L6-v2)
  • MILVUS_URI — Milvus connection
  • CHUNK_SIZE, CHUNK_OVERLAP — Text chunking
  • ENABLE_HYBRID_SEARCH — BM25 + vector fusion
  • ENABLE_IMAGE_DESCRIPTION — LLM vision for images

Development Tasks

The project includes a comprehensive Makefile for common tasks:

# Show available tasks
make

# ── Setup & Quality ──────────────────────────────────────────
make setup          # Install Python dependencies
make test           # Run test suite
make lint           # Lint code (ruff)
make format         # Format code (ruff format)
make typecheck      # Type checking (mypy)
make clean          # Clean build artifacts

# ── Container Stack (app + data in containers) ──────────────
make up             # Start full stack
make down           # Stop stack
make logs           # Show service logs
make ps             # Show running containers

# ── Dev-Fast (app on host, hot-reload) ──────────────────────
make dev-up          # Start data infra only (Postgres, Valkey, Milvus)
make dev-down        # Stop data infra containers
make dev-logs        # Show data infra logs
make dev-setup       # Create media dir + run migrations
make dev-fast        # Run API with hot-reload (no container)
make dev-fast-worker # Run ARQ worker directly (no container)
make dev-migrate     # Run Alembic migrations
make dev-create-tables # Create tables directly (no Alembic)

Integration with pydantic-deepagents

from rag_agent.client import RAGAgentClient

client = RAGAgentClient(base_url="http://localhost:8100")

# Upload
result = await client.upload_document("report.pdf")

# Search
results = await client.search("quarterly earnings")

Requirements

  • uv (https://github.com/astral-sh/uv) — Modern Python package installer and resolver
  • Python 3.12+ — Runtime environment
  • Docker and Docker Compose — For running data infrastructure (Postgres, Valkey, Milvus). Required by both the full container stack and the dev-fast workflow. If you only run unit tests, Docker is optional (tests use SQLite).

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

verity_rag-0.1.2.tar.gz (250.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

verity_rag-0.1.2-py3-none-any.whl (63.4 kB view details)

Uploaded Python 3

File details

Details for the file verity_rag-0.1.2.tar.gz.

File metadata

  • Download URL: verity_rag-0.1.2.tar.gz
  • Upload date:
  • Size: 250.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for verity_rag-0.1.2.tar.gz
Algorithm Hash digest
SHA256 44737240b47f9d59431f2bc7078506ded76e37f02b64ec65ea7f2d34bc0c4f64
MD5 ecd040bb194175938f778642d3a1363f
BLAKE2b-256 b8f93679d771069e42a14a2e7d391eb8db5621d9f0f52b906478db35a86dc0d7

See more details on using hashes here.

File details

Details for the file verity_rag-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: verity_rag-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 63.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for verity_rag-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 00339617db8c4c5e882063e7be2e4f9dd341b3d14db216a35cbdd9b3c8326894
MD5 b67f7389fdd6ebd76a91f4e5a48fee29
BLAKE2b-256 675ca639bf537b94ff676cdd0705293001bcf7b47e7c4599842be1e515b27a4f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page