Skip to main content

Darwin RAG

A local-first RAG engine that ingests documents, indexes them with BM25 + dense embeddings, and exposes search via an MCP server for AI agent integration.

  • Ingestion — PDF, Markdown, HTML, images (OCR), CSV, Excel, ODS, URLs
  • Indexing — BM25 keyword + dense embedding hybrid index with configurable chunking strategies
  • Search — Hybrid, semantic, or keyword retrieval with reranking, diversity rerank, and structural penalties
  • Generation — LLM-backed answer synthesis via LiteLLM (OpenAI, Anthropic, Gemini, etc.)
  • Observability — Structured logging with per-session history, queryable via MCP
  • Isolation — Multiple independent stores for tenant/project separation
  • Deployment — stdio (AI agent subprocess), SSE, or Streamable HTTP; Docker-ready
  • Local-first — Everything runs locally, fully offline-capable after setup

Built by BrightDotDev.
License: MIT with Attribution


Quick Start

1. Install

pip install darwin-rag

2. Set up models

# Interactive — detects hardware, pick your models
darwin-admin setup interactive

# Or one-shot (embedding-only, no prompts)
darwin-admin setup --preset required

3. Start the MCP server & connect

# stdio mode — for AI agent subprocess (Claude Desktop, Cursor, etc.)
darwin mcp

# Or HTTP mode — for remote clients
darwin mcp --http --port 8765

Configure your MCP client:

{
  "mcpServers": {
    "darwin": {
      "command": "darwin",
      "args": ["mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-..."  // At least one LLM provider key
      }
    }
  }
}

Or generate config automatically:

darwin config claude          # Claude Desktop config
darwin config cursor          # Cursor config
darwin config all --copy      # All clients + copy to clipboard

MCP Tools

Tool Description
search_darwin Query the knowledge base with hybrid/semantic/keyword search
get_search_results List saved search results
get_search_result_by_id Load a saved search result by filename
run_pipeline Ingest + index documents from a path or URL
purge_artifacts Delete pipeline artifacts for specific files
create_store Create a new isolated data store
list_files List all tracked files with pipeline status
file_status Detailed status for a single file across all stages
get_logs Query session logs (oldest first, INFO excluded)

Full documentation: docs/mcp.md


Remote / HTTP Mode

Start the server on a network-accessible endpoint:

# SSE transport (legacy)
darwin mcp --sse --host 0.0.0.0 --port 8765

# Streamable HTTP transport (recommended for production)
darwin mcp --http --host 0.0.0.0 --port 8765

Configure your MCP client with the URL:

{
  "mcpServers": {
    "darwin": {
      "url": "http://your-host:8765/mcp"  // or /sse for SSE mode
    }
  }
}

Environment Variables

Variable Required Description
OPENAI_API_KEY No* OpenAI provider key
ANTHROPIC_API_KEY No* Anthropic provider key
GEMINI_API_KEY No* Google Gemini provider key
MISTRAL_API_KEY No* Mistral AI provider key
GROQ_API_KEY No* Groq provider key
COHERE_API_KEY No* Cohere provider key
TOGETHER_API_KEY No* Together AI provider key
OPENROUTER_API_KEY No* OpenRouter provider key
DEEPSEEK_API_KEY No* DeepSeek provider key
DARWIN_BASE_DIR No Override the base data directory
NO_COLOR No Set to any value to disable ANSI color output

* At least one LLM provider key is required for answer generation. Search/indexing works without any.


Setup Details

Command What it does
darwin-admin setup interactive Guided setup — detect hardware, choose models
darwin-admin setup --preset required Download embedding model only (fastest)
darwin-admin setup --preset recommended Embedding + reranker + OCR models
darwin-admin setup logging Reconfigure logging only
darwin-admin setup validate Validate current setup

See docs/setup.md for the full walkthrough including Docker, from-source install, and API key configuration.


CLI Reference

darwin — User CLI

Command Description
darwin mcp Start MCP server (stdio, --sse or --http for network)
darwin config [client] Generate MCP client config

darwin-admin — Power-user CLI

Command Description
darwin-admin setup Setup models, logging, and configuration
darwin-admin status System status overview
darwin-admin models Model registry: list, install, switch, keys
darwin-admin store Data store: status, files, audit, health, repair
darwin-admin pipeline Ingestion pipeline: run, ingest, index, purge
darwin-admin search Interactive search
darwin-admin logs Structured log viewer and management
darwin-admin system System information
darwin-admin uninstall Remove Darwin data and configuration

See docs/admin.md for the full command reference.


Documentation

Doc What
setup.md Full setup walkthrough
mcp.md MCP server, tools, resources, transports
admin.md Admin CLI reference
architecture.md For developers and contributors
pipeline.md Ingestion & indexing
retrieval.md Search engine
storage.md DarwinStore
models.md Model registry & inference
logger.md Structured logging
orchestrators.md High-level business logic

Contributing

Found a bug? Want to add something? You're welcome here.

  • Issues — open one at github.com/BrightDotDev/DARWIN/issues
  • Code — fork, branch, PR. Keep it minimal.
  • AI-generated code is fine — but you own what you ship. Test it before submitting.

Read CONTRIBUTING.md for the full guidelines.


License

MIT with Attribution — see LICENSE.

Core architecture and implementation by BrightDotDev.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

darwin_rag-0.2.16.tar.gz (311.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

darwin_rag-0.2.16-py3-none-any.whl (365.2 kB view details)

Uploaded Python 3

File details

Details for the file darwin_rag-0.2.16.tar.gz.

File metadata

  • Download URL: darwin_rag-0.2.16.tar.gz
  • Upload date:
  • Size: 311.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.8

File hashes

Hashes for darwin_rag-0.2.16.tar.gz
Algorithm Hash digest
SHA256 111a2f0e458dec0f6bfced9cbd26abae2cc0054d1aca2b11410a434ea8285ba0
MD5 90cde483b2e6ee8ab7b8ac83df71892a
BLAKE2b-256 b4e70758f4565ba062b22ca20d8694315bd866fa1af8bb8a9330f7f2e9ce907e

See more details on using hashes here.

File details

Details for the file darwin_rag-0.2.16-py3-none-any.whl.

File metadata

  • Download URL: darwin_rag-0.2.16-py3-none-any.whl
  • Upload date:
  • Size: 365.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.8

File hashes

Hashes for darwin_rag-0.2.16-py3-none-any.whl
Algorithm Hash digest
SHA256 3c75f42a8e03e1023cb2be6001bf31396d9e92862efb745c0e4f8520fd3f5fda
MD5 63763ab716000ce7015f91534a0eef55
BLAKE2b-256 0ba5f4548b61672f9e1e39f526f325b205c9149ece98604b98c46977a3698494

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page