Skip to main content

Sci RAG Kit

CI Documentation License: BSD-3-Clause Python 3.11 | 3.12

A template repository for retrieval-augmented generation, built around your scientific domain. It implements hybrid GraphRAG retrieval, grounded answer generation with citations, an evaluation harness, and REST API plus MCP endpoints.

Read the documentation site for the guided path, or go directly to the methodology for the design specification.

To start a project of your own, run pipx install sci-rag-kit and then sci-rag new; the wizard asks about your field and writes a configured, git-initialized project. To evaluate the kit first, run the quickstart below against the bundled demo corpus.

Components

  • Ingestion: PDF, HTML, Markdown, and plain-text parsing (Docling when installed, pypdf fallback; HTML through the standard library, with page chrome stripped), structure-aware chunking that preserves section hierarchy and keeps tables intact, content-hash deduplication, per-document license metadata.
  • Retrieval: five candidate generators run in parallel and fuse by weighted reciprocal rank: dense vectors (pgvector + HNSW), Postgres full-text search, knowledge-graph traversal, community summaries, and HyDE. Per-stage timeouts, traces, and graceful degradation.
  • Knowledge graph: LLM extraction of entities and typed relationships constrained to a user-defined ontology (domain/domain.yaml), with evidence provenance per edge; deterministic community detection with LLM-written, embedded summaries. Stored as ordinary Postgres rows; no separate graph database (see ADR 0001).
  • Access control on content: each document carries a redistribution class (public, open_commercial, open_noncommercial, restricted, unknown). Retrieval scopes are applied inside every layer's SQL before ranking; an empty allowlist returns nothing.
  • Answering: numbered inline citations tied to retrieved sources; when retrieval finds nothing in scope, the system states that instead of answering from model priors.
  • Model providers: Gemini, Claude, and any OpenAI-compatible endpoint, chosen per role with a provider:model setting. On Google Cloud that reaches the Vertex Model Garden partner models (Claude, Grok, Llama, Mistral) with no credentials beyond the project you already have.
  • Evaluation: expert-authored seed questions, retrieval metrics (hit@k, MRR) with per-layer ablation configs, and a two-pass LLM judge: the grounding pass never sees the reference answer, and correctness is graded separately against it. Reports are stamped with a corpus fingerprint and git commit.
  • Serving: a FastAPI service (/v1, OpenAPI at /docs) and an MCP server (eight tools, mounted at /mcp and runnable over stdio) backed by the same service instance. Static API keys with scopes and rate limits, an interface seam for OAuth, and per-request LLM key override.

Everything runs in a single PostgreSQL database: text, vectors, full-text indexes, and the graph.

Quickstart

Requirements: uv, a supported PostgreSQL backend (Docker by default, local PostgreSQL, or the opt-in Cloud SQL development backend), and optionally a Google AI Studio API key or Vertex AI credentials for real embeddings and generation.

git clone https://github.com/sustainability-software-lab/sci-rag-kit.git
cd sci-rag-kit

cp .env.example .env
chmod 600 .env          # owner only: it is about to hold a credential
# In .env, set one of:
#   SCI_RAG_GOOGLE_API_KEY=...              AI Studio key
#   SCI_RAG_GCP_PROJECT=...                 Vertex AI (after gcloud auth application-default login)
#   SCI_RAG_EMBEDDING_PROVIDER=local-hash   offline mode: no credentials, lexical-only retrieval, no generation
# To generate with Claude or Grok instead of Gemini, see docs/extend.md.

make setup     # sync dependencies, start the selected backend, create the schema
make demo      # ingest the demo corpus, run a traced retrieval, score it

make setup starts the selected database backend and applies every migration. Docker is the template default; generated projects may select conda-forge, system PostgreSQL, or the optional Cloud SQL development helper. See Run Postgres your way for the supported combinations.

With credentials configured, generation and the graph work too:

uv run sci-rag answer "What conversion route suits rice straw given its ash content?"
make demo-cloud   # graph extraction + communities + a deep answer + ablation report
uv run sci-rag serve   # REST at /docs, MCP at /mcp

Example answer from the demo corpus (five numbered sources across three documents, all claims cited):

Given its ash content, anaerobic digestion is a suitable conversion route for rice straw [2][4]. Rice straw has an ash content near 18 percent, which includes high silica [2]. This high ash and silica limit direct combustion [2] and prevent its use in gasifiers due to accelerated clinker formation [5]. ... a mild alkali soak raises the biogas yield to 320 cubic meters per dry ton [1][3].

The demo corpus is five synthetic documents about agricultural residues (realistic form, fictional numbers, CC0), included so the pipeline can be exercised end to end before you commit your own documents.

Customizing to your domain

You do not have to write the domain files cold. sci-rag draft creates corpus-grounded first passes; the LLM-assisted setup guide also shows a copy-paste workflow that needs no model credentials.

  1. Put documents in data/raw/ and describe them in a JSONL corpus manifest (title, authors, license class, source).
  2. Edit domain/domain.yaml: entity types, relationship types, and HyDE query classes for your field.
  3. Adjust the wording of the prompts in domain/prompts/ where needed.
  4. Write 10 to 20 ground-truth questions in domain/eval_seed_questions.jsonl.
  5. Run sci-rag ingest, sci-rag graph extract, sci-rag graph communities, then sci-rag eval retrieval --ablation.

The step-by-step version with worked examples is Bring your own domain. Inside a checkout, uv run sci-rag init lets you choose Quick or Advanced setup, while uv run sci-rag init --advanced asks every applicable question; uv run python scripts/init_domain.py is the narrow path when all you want is the project name, description, and a seed-question reset.

Repository layout

domain/            Ontology, prompts, seed questions (everything specific to your field)
src/sci_rag/       ingest, embed, graph, retrieve, answer, evals, server, cli
data/demo/         Demo corpus (synthetic, CC0; optional)
migrations/        Alembic schema (pgvector + HNSW + FTS indexes)
tests/             Offline-first suite; database tests use a disposable selected backend
infra/terraform/   Optional production deployment plus a separate dev database module
docs/              Methodology, tutorials, API reference, ADRs

CLI

Command Purpose
sci-rag new Create a configured project with Quick or Advanced setup
sci-rag init Configure the checkout in the current directory
sci-rag db upgrade Create or upgrade the database schema
sci-rag ingest <folder> / --manifest file.jsonl Parse, chunk, embed, store
sci-rag corpus enrich --mailto you@example.org Add Crossref journal, citation-count, and retraction metadata (--dry-run first)
sci-rag campaign discover --topic ... | --doi-file ... Build a deduplicated, resumable DOI list through OpenAlex or Crossref
sci-rag campaign build --topic ... | --doi-file ... --dry-run Map explicit license signals, download verified direct OA PDFs, and write an ingest manifest
sci-rag campaign screen --name ... --criteria-file ... Screen discovered abstracts and route uncertain or invalid model results to human review
sci-rag campaign review --name ... Walk the pending review queue and append explicit human decisions
sci-rag graph extract Extract entities and relationships from chunks
sci-rag graph resolve-entities --dry-run Preview alias, fuzzy, and optional LLM duplicate-entity merges
sci-rag graph citations --dry-run Reconcile cached Crossref references into resolved and unresolved DOI pointers
sci-rag graph communities Cluster the graph and write summaries
sci-rag retrieve "question" Ranked results with per-layer traces (filter with --year-min/--year-max/--author/--journal/--exclude-doi/--license/--source)
sci-rag answer "question" Grounded answer with citations; known retracted papers excluded by default
sci-rag eval retrieval [--ablation] Retrieval metrics against seed questions
sci-rag eval answers Generate and judge answers
sci-rag serve REST + MCP server
sci-rag mcp MCP over stdio (for local agents)
sci-rag stats Corpus contents and relationship-confidence summary
sci-rag doctor Check config, database, corpus, and credentials in one pass

Register the MCP server with a local agent:

claude mcp add my-corpus -- uv run --directory /path/to/your/repo sci-rag mcp

Documentation

The complete, searchable site is published at sustainability-software-lab.github.io/sci-rag-kit.

Quickstart Setup, first run, troubleshooting
FAQ Short answers, and the reasoning behind each design decision
Bring your own domain Configure the kit for your field
Run Postgres your way Choose Docker, conda-forge, system PostgreSQL, or the optional Cloud helper
Run a corpus campaign Polite, resumable discovery from topics or DOI seeds
Methodology Design rationale for every component
Architecture Code layout, data model, extension points
Evaluate your pipeline Seed questions, ablations, the judge
Benchmarks Measured demo-corpus results, reproducible via make benchmark
Choosing Sci RAG Kit Honest comparison vs GraphRAG, LightRAG, PaperQA2, LlamaIndex
Roadmap Waves 2-3, collaboration seams, launch-gated decisions
REST, MCP, and Python API REST endpoints, MCP tools, auth, error codes
Deploy on Google Cloud Cloud SQL + Cloud Run via Terraform
Decision records Postgres-native graph, embedding dimensions, Docling, template format
Versioning + Governance What 0.x promises; how decisions get made
Adopters Who runs a knowledge base built from the kit

Defaults and requirements

Python 3.11+; PostgreSQL 16 to 18 with pgvector from the selected backend. Docker is the template default and matches the PostgreSQL 16 CI service. Default models: gemini-embedding-001 at 1536 dimensions (within pgvector's HNSW index limit; see ADR 0002) and gemini-2.5-flash for generation, via AI Studio key or Vertex AI. A deterministic offline embedder covers tests and credential-free runs. Docling is an optional extra (uv sync --extra docling) because of its install size; without it, PDF parsing falls back to pypdf at reduced table fidelity.

License

BSD 3-Clause. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sci_rag_kit-0.4.1.tar.gz (6.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sci_rag_kit-0.4.1-py3-none-any.whl (309.9 kB view details)

Uploaded Python 3

File details

Details for the file sci_rag_kit-0.4.1.tar.gz.

File metadata

  • Download URL: sci_rag_kit-0.4.1.tar.gz
  • Upload date:
  • Size: 6.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for sci_rag_kit-0.4.1.tar.gz
Algorithm Hash digest
SHA256 df00f37eeac23439815fc873ba27646ea269bd56f17bb031a6cc77a80f605a35
MD5 3dad862f1fbc0337348f256b88effdfc
BLAKE2b-256 f5c0f45c01d0c09f0de6bd8d9b2547d265384f5659c17a9ce65e8063e138afe6

See more details on using hashes here.

Provenance

The following attestation bundles were made for sci_rag_kit-0.4.1.tar.gz:

Publisher: release.yml on sustainability-software-lab/sci-rag-kit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file sci_rag_kit-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: sci_rag_kit-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 309.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for sci_rag_kit-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 8f2d858ad345ba23a3f8d9da8cc4900e0ae21bcee3fdb7264af8d17fd7ca314c
MD5 ccf9f193fa816d28cd4135dfc3026a91
BLAKE2b-256 411154bf04ef3350516138f3e1ef0d3861049bc0bbcd7d44b268f1887b311b14

See more details on using hashes here.

Provenance

The following attestation bundles were made for sci_rag_kit-0.4.1-py3-none-any.whl:

Publisher: release.yml on sustainability-software-lab/sci-rag-kit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.5.0

2 files

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page