Skip to main content

Production-grade self-hosted document ingestion and retrieval service

Project description

RAG Agent

Production-grade, self-hosted document ingestion and retrieval service.

Features

  • Multi-format parsing: PDF, DOCX, TXT, Markdown with smart layout preservation
  • Chunking strategies: Recursive character, Markdown headers
  • Image extraction and LLM description: Makes visual content searchable
  • OCR fallback: LLM vision for scanned pages
  • Vector storage: Milvus with cosine similarity search
  • Hybrid retrieval: BM25 keyword search + vector fusion (RRF)
  • Cross-encoder reranking: Sentence Transformers for relevance scoring
  • ARQ task queue: Background processing with retry and backoff
  • SSE status streaming: Real-time ingestion progress
  • Pluggable connectors: Local filesystem (ready), S3/Google Drive (pluggable)
  • Deduplication: Content hash + source path matching
  • Batched embedding: Configurable batch size with retry
  • Web dashboard: Responsive dark-themed UI for all operations
  • uv Python manager: Fast dependency installation and management
  • Makefile: Simplified task management

Architecture

Upload → Validate → Store → Track (DB) → Queue (ARQ)
  ┌─── Worker ───────────────────────────────────────┐
  │ Parse → Describe images → Chunk → Dedup → Embed → Store (Milvus)
  └──────────────────────────────────────────────────┘
                ↓
          SSE Status Events
                ↓
           Query → Search
                ↓
          Web Dashboard  ←── You are here

Web UI

The project includes a responsive dark-themed web dashboard built with Vue 3 (served as static files from the FastAPI application).

Pages:

Page Description
Dashboard System health, collection stats, recent documents
Documents Upload (drag & drop), list/filter, delete, retry, download
Collections Create, browse, delete vector collections
Search Semantic search with reranker, multi-collection mode, score visualization

Access the UI at http://localhost:8100/ (redirects to /ui/).

The frontend is served directly by the API server — no separate build step or dev server needed. Source lives in the frontend/ directory.

Quick Start

Two workflows are available:

🐳 Full Container Stack (production-like)

# Install uv (https://github.com/astral-sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install dependencies
make setup

# Create .env from example
cp .env.example .env
# Edit .env with your settings

# Start everything (app + data infra in containers)
make up

# Wait for services to be ready
make wait

# Upload a document
curl -X POST http://localhost:8100/api/v1/documents/upload \
  -F "file=@example.pdf"

# Search
curl -X POST http://localhost:8100/api/v1/search \
  -H "Content-Type: application/json" \
  -d '{"query": "What is the revenue?", "limit": 5}'

⚡ Dev-Fast (app on host, hot-reload)

Run the Python code directly on your machine for instant feedback — only the data stores (Postgres, Valkey, Milvus) run in containers.

# Install uv (https://github.com/astral-sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install dependencies
make setup

# Start ONLY data infrastructure containers
make dev-up

# Create media directory + run database migrations
make dev-setup

# (Terminal 1) Start API with hot-reload
make dev-fast

# (Terminal 2) Start background worker
make dev-fast-worker

# Open http://localhost:8100/ in your browser

# When done, stop data containers
make dev-down

Changes to Python files are picked up instantly — no Docker rebuilds needed.

Manual uv Commands

# Install dependencies
uv sync --all-extras

# Run tests
uv run pytest tests/ -v

# Start API server (uses .env)
uv run uvicorn rag_agent.app:create_app --reload --factory

# Start ARQ worker
uv run arq rag_agent.worker.settings.WorkerSettings

API Endpoints

Health

  • GET /health — Liveness
  • GET /ready — Readiness with dependency checks
  • GET /live — Minimal liveness

Collections

  • GET /api/v1/collections — List collections
  • POST /api/v1/collections?name=... — Create collection
  • GET /api/v1/collections/{name} — Collection stats
  • DELETE /api/v1/collections/{name} — Drop collection

Documents

  • POST /api/v1/documents/upload — Upload file (multipart)
  • GET /api/v1/documents — List tracked documents
  • GET /api/v1/documents/{id} — Document detail
  • DELETE /api/v1/documents/{id} — Delete (cascade)
  • POST /api/v1/documents/{id}/retry — Re-queue failed ingestion
  • GET /api/v1/documents/{id}/download — Download original

Search

  • POST /api/v1/search — Vector search
  • POST /api/v1/search/multi — Multi-collection search
  • GET /api/v1/collections/{name}/documents/{id} — Search within a document

Sync & Connectors

  • POST /api/v1/sync — Trigger directory sync
  • GET /api/v1/sync/logs — Sync history
  • GET /api/v1/connectors — Available connectors
  • GET /api/v1/status — SSE stream for progress events

Configuration

See .env.example for all environment variables.

Key settings:

  • EMBEDDING_BASE_URL — OpenAI-compatible embedding endpoint
  • EMBEDDING_MODEL — Model name (e.g., all-MiniLM-L6-v2)
  • MILVUS_URI — Milvus connection
  • CHUNK_SIZE, CHUNK_OVERLAP — Text chunking
  • ENABLE_HYBRID_SEARCH — BM25 + vector fusion
  • ENABLE_IMAGE_DESCRIPTION — LLM vision for images

Development Tasks

The project includes a comprehensive Makefile for common tasks:

# Show available tasks
make

# ── Setup & Quality ──────────────────────────────────────────
make setup          # Install Python dependencies
make test           # Run test suite
make lint           # Lint code (ruff)
make format         # Format code (ruff format)
make typecheck      # Type checking (mypy)
make clean          # Clean build artifacts

# ── Container Stack (app + data in containers) ──────────────
make up             # Start full stack
make down           # Stop stack
make logs           # Show service logs
make ps             # Show running containers

# ── Dev-Fast (app on host, hot-reload) ──────────────────────
make dev-up          # Start data infra only (Postgres, Valkey, Milvus)
make dev-down        # Stop data infra containers
make dev-logs        # Show data infra logs
make dev-setup       # Create media dir + run migrations
make dev-fast        # Run API with hot-reload (no container)
make dev-fast-worker # Run ARQ worker directly (no container)
make dev-migrate     # Run Alembic migrations
make dev-create-tables # Create tables directly (no Alembic)

Integration with pydantic-deepagents

from rag_agent.client import RAGAgentClient

client = RAGAgentClient(base_url="http://localhost:8100")

# Upload
result = await client.upload_document("report.pdf")

# Search
results = await client.search("quarterly earnings")

Requirements

  • uv (https://github.com/astral-sh/uv) — Modern Python package installer and resolver
  • Python 3.12+ — Runtime environment
  • Docker and Docker Compose — For running data infrastructure (Postgres, Valkey, Milvus). Required by both the full container stack and the dev-fast workflow. If you only run unit tests, Docker is optional (tests use SQLite).

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

verity_rag-0.1.1.tar.gz (248.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

verity_rag-0.1.1-py3-none-any.whl (62.4 kB view details)

Uploaded Python 3

File details

Details for the file verity_rag-0.1.1.tar.gz.

File metadata

  • Download URL: verity_rag-0.1.1.tar.gz
  • Upload date:
  • Size: 248.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for verity_rag-0.1.1.tar.gz
Algorithm Hash digest
SHA256 cbea9fc0f876b378f6021ce94b5f91e1baa6d25f47d7d47171fc911b40d8b0a0
MD5 d1a39d3adf4e1ce39180b2ccd67a1cf1
BLAKE2b-256 8e8be500e8d5bba3f52ef5c055a08d9b1121abdd67a122ffe0cac614419dfb60

See more details on using hashes here.

File details

Details for the file verity_rag-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: verity_rag-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 62.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for verity_rag-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 abab486570ca1284ca8bb8568b069c34c2fd4a5a89210406aa2777f11272c317
MD5 5a7494b0815216c1cbf09716e3d095e3
BLAKE2b-256 e9af786294d850da86d1fff7bcf066389af8080646207a91c81dcac175d89866

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page