Skip to main content

Darwin RAG

A local-first RAG engine that ingests documents, indexes them with BM25 + dense embeddings, and exposes search via an MCP server for AI agent integration.

  • Ingestion — PDF, Markdown, HTML, images (OCR), CSV, Excel, ODS, URLs
  • Indexing — BM25 keyword + dense embedding hybrid index with configurable chunking strategies
  • Search — Hybrid, semantic, or keyword retrieval with reranking, diversity rerank, and structural penalties
  • Generation — LLM-backed answer synthesis via LiteLLM (OpenAI, Anthropic, Gemini, etc.)
  • Observability — Structured logging with per-session history, queryable via MCP
  • Isolation — Multiple independent stores for tenant/project separation
  • Deployment — stdio (AI agent subprocess), SSE, or Streamable HTTP; Docker-ready
  • Local-first — Everything runs locally, fully offline-capable after setup

Built by BrightDotDev.
License: MIT with Attribution


Quick Start

1. Install

pip install darwin-rag

2. Set up models

# Interactive — detects hardware, pick your models
darwin-admin setup interactive

# Or one-shot (embedding-only, no prompts)
darwin-admin setup --preset required

3. Start the MCP server & connect

# stdio mode — for AI agent subprocess (Claude Desktop, Cursor, etc.)
darwin mcp

# Or HTTP mode — for remote clients
darwin mcp --http --port 8765

Configure your MCP client:

{
  "mcpServers": {
    "darwin": {
      "command": "darwin",
      "args": ["mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-..."  // At least one LLM provider key
      }
    }
  }
}

Or generate config automatically:

darwin config claude          # Claude Desktop config
darwin config cursor          # Cursor config
darwin config all --copy      # All clients + copy to clipboard

MCP Tools

Tool Description
search_darwin Query the knowledge base with hybrid/semantic/keyword search
get_search_results List saved search results
get_search_result_by_id Load a saved search result by filename
run_pipeline Ingest + index documents from a path or URL
purge_artifacts Delete pipeline artifacts for specific files
create_store Create a new isolated data store
list_files List all tracked files with pipeline status
file_status Detailed status for a single file across all stages
get_logs Query session logs (oldest first, INFO excluded)

Full documentation: docs/mcp.md


Remote / HTTP Mode

Start the server on a network-accessible endpoint:

# SSE transport (legacy)
darwin mcp --sse --host 0.0.0.0 --port 8765

# Streamable HTTP transport (recommended for production)
darwin mcp --http --host 0.0.0.0 --port 8765

Configure your MCP client with the URL:

{
  "mcpServers": {
    "darwin": {
      "url": "http://your-host:8765/mcp"  // or /sse for SSE mode
    }
  }
}

Environment Variables

Variable Required Description
OPENAI_API_KEY No* OpenAI provider key
ANTHROPIC_API_KEY No* Anthropic provider key
GEMINI_API_KEY No* Google Gemini provider key
MISTRAL_API_KEY No* Mistral AI provider key
GROQ_API_KEY No* Groq provider key
COHERE_API_KEY No* Cohere provider key
TOGETHER_API_KEY No* Together AI provider key
OPENROUTER_API_KEY No* OpenRouter provider key
DEEPSEEK_API_KEY No* DeepSeek provider key
DARWIN_BASE_DIR No Override the base data directory
NO_COLOR No Set to any value to disable ANSI color output

* At least one LLM provider key is required for answer generation. Search/indexing works without any.


Setup Details

Command What it does
darwin-admin setup interactive Guided setup — detect hardware, choose models
darwin-admin setup --preset required Download embedding model only (fastest)
darwin-admin setup --preset recommended Embedding + reranker + OCR models
darwin-admin setup logging Reconfigure logging only
darwin-admin setup validate Validate current setup

See docs/setup.md for the full walkthrough including Docker, from-source install, and API key configuration.


CLI Reference

darwin — User CLI

Command Description
darwin mcp Start MCP server (stdio, --sse or --http for network)
darwin config [client] Generate MCP client config

darwin-admin — Power-user CLI

Command Description
darwin-admin setup Setup models, logging, and configuration
darwin-admin status System status overview
darwin-admin models Model registry: list, install, switch, keys
darwin-admin store Data store: status, files, audit, health, repair
darwin-admin pipeline Ingestion pipeline: run, ingest, index, purge
darwin-admin search Interactive search
darwin-admin logs Structured log viewer and management
darwin-admin system System information
darwin-admin uninstall Remove Darwin data and configuration

See docs/admin.md for the full command reference.


Documentation

Doc What
setup.md Full setup walkthrough
mcp.md MCP server, tools, resources, transports
admin.md Admin CLI reference
architecture.md For developers and contributors
pipeline.md Ingestion & indexing
retrieval.md Search engine
storage.md DarwinStore
models.md Model registry & inference
logger.md Structured logging
orchestrators.md High-level business logic

Contributing

Found a bug? Want to add something? You're welcome here.

  • Issues — open one at github.com/BrightDotDev/DARWIN/issues
  • Code — fork, branch, PR. Keep it minimal.
  • AI-generated code is fine — but you own what you ship. Test it before submitting.

Read CONTRIBUTING.md for the full guidelines.


License

MIT with Attribution — see LICENSE.

Core architecture and implementation by BrightDotDev.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

darwin_rag-0.2.14.tar.gz (310.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

darwin_rag-0.2.14-py3-none-any.whl (364.3 kB view details)

Uploaded Python 3

File details

Details for the file darwin_rag-0.2.14.tar.gz.

File metadata

  • Download URL: darwin_rag-0.2.14.tar.gz
  • Upload date:
  • Size: 310.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.8

File hashes

Hashes for darwin_rag-0.2.14.tar.gz
Algorithm Hash digest
SHA256 f9d2ae8883db225aaeb13c82152f23194264a31c22cad29aa3f62485bbe32d7f
MD5 b7b73d22a448fe8f079eefaae204b950
BLAKE2b-256 fa913904726a9d72229c772c90dcd567e432a2491e7f6d2a078e4ad6570848f4

See more details on using hashes here.

File details

Details for the file darwin_rag-0.2.14-py3-none-any.whl.

File metadata

  • Download URL: darwin_rag-0.2.14-py3-none-any.whl
  • Upload date:
  • Size: 364.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.8

File hashes

Hashes for darwin_rag-0.2.14-py3-none-any.whl
Algorithm Hash digest
SHA256 9b8255802fc794a97518f2b49337b81b69a6142e944dcd1fd14f795e12164a66
MD5 b95eba4b08182cef07ca8adaf6250e93
BLAKE2b-256 a1fd54815754f355235d9040846fabf431d31dd74f970c0ecc272f2ee10558af

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page