Skip to main content

Local knowledge base for AI agents with SQLite FTS5 search, URI routing, and conflict detection

Project description

DocsHaven

License: MIT PyPI version GitHub stars GitHub last commit CI Python

Your AI agent keeps forgetting what it learned last session. DocsHaven fixes that.

Add your repos, and your agent always has them at hand. DocsHaven is a local knowledge base that lets you index any GitHub repository, search it instantly, and keep your agent informed across sessions. Zero dependencies, works with any MCP-compatible agent — Claude, Cursor, Gemini, Codex.

Why DocsHaven?

  • 🧠 Add repos once, search forever — your agent always has the knowledge it needs
  • Instant search — SQLite FTS5 finds relevant docs in <100ms
  • 🔌 Works with any AI agent — Claude, Cursor, Gemini, Codex via MCP
  • 🏠 100% local — no cloud, no API keys, no data leaves your machine
  • 📦 Minimal dependencies — only mcp[cli], no Docker or external services

Features

  • SQLite FTS5 search — BM25 ranking with LIKE fallback
  • URI routing — organize knowledge by domain: core://, ref://, guide://
  • Git sync — compressed chunks for multi-machine sync (no merge conflicts)
  • Conflict detection — flag contradictions when adding documents
  • MCP server — 16 tools for any MCP-compatible agent
  • Document chunking — split long documents for better search precision

Installation

pip install docs-haven

Or from source:

git clone https://github.com/Cipher208/docs-haven.git
cd docs-haven
pip install -e .
With test dependencies
pip install -e ".[test]"

How to Use

1. Add Repositories to Your Knowledge Base
from storage import Storage
from pathlib import Path

storage = Storage(Path.home() / ".docshaven")

# Add a GitHub repo (clones and indexes markdown files)
result = storage.add_repo(
    url="https://github.com/fastapi/fastapi",
    description="FastAPI web framework",
)
print(f"Indexed {result['files_indexed']} files in {result['chunks']} chunks")

# Add with custom file mask (index Python files)
result = storage.add_repo(
    url="https://github.com/pallets/flask",
    mask="**/*.py",
)
2. Search Your Knowledge Base
# Basic search
results = storage.search("dependency injection")
for r in results:
    print(f"{r['score']:.2f} [{r['collection']}] {r['title']}")
    print(f"  {r['content'][:100]}...")
    print()

# Search with filters
results = storage.search(
    "async middleware",
    collections=["fastapi"],
    limit=5,
    min_score=0.3,
)

# Use auto strategy (fts for short queries, hybrid for long)
results = storage.search("how to use Depends()", strategy="auto")
3. Organize with URI Routing
from uri import URI, URIRouter

# Parse URIs
uri = URI.parse("core://fastapi/dependencies")
print(uri.domain)      # "core"
print(uri.path)        # "fastapi/dependencies"
print(uri.to_collection())  # "core__fastapi"

# Search within a URI scope
router = URIRouter(storage)
results = router.search_by_uri("core://fastapi", limit=5)

# List all domains
domains = router.list_all_domains()
# {'core': {'count': 3, 'doc': 'Core documentation'}, 'ref': {'count': 2, ...}}

Available domains: core, ref, guide, lib, src, test, note

4. Detect Conflicts
from conflicts import ConflictDetector

detector = ConflictDetector(storage)
result = detector.detect(
    title="FastAPI dependency injection",
    content="How to use Depends()...",
)

if result.has_conflicts:
    print(f"Found {len(result.candidates)} similar documents:")
    for c in result.candidates:
        print(f"  - {c['title']} (score: {c['score']})")
        print(f"    {c['snippet'][:80]}...")

    # Record judgment
    detector.judge("new_doc_id", "existing_doc_id", "supersedes")
5. Sync Between Machines
from sync import Syncer
from pathlib import Path

syncer = Syncer(Path.home() / ".docshaven-sync")

# On machine A: export
result = syncer.export(
    {"fastapi": docs, "sqlalchemy": docs},
    created_by="alice",
)
print(f"Exported chunk {result['chunk_id']}")

# On machine B: import
result = syncer.import_chunks()
print(f"Imported {result['chunks_imported']} chunks")

# Check status
status = syncer.status()
print(f"Chunks: {status['local_chunks']}")
6. Use as MCP Server

Add to your MCP client config (Claude Desktop, Cursor, etc.):

{
  "mcpServers": {
    "docs-haven": {
      "command": "python",
      "args": ["/path/to/docs-haven/server.py"]
    }
  }
}

Then ask your agent:

"Search for FastAPI middleware examples" "What documentation do we have about SQLAlchemy?" "Check if this new doc conflicts with existing ones"

MCP Tools (16)

Tool Description
kb_search Search with BM25 ranking + highlighted excerpts
kb_add_repo Clone and index a GitHub repo
kb_get Get document content
kb_update Update document content
kb_delete Delete document from knowledge base
kb_list_collections List all collections
kb_stats Database statistics
kb_uri_resolve URI to collection mapping
kb_uri_search Search within URI scope (supports wildcards)
kb_uri_list List URIs in domain
kb_uri_domains All domains with counts
kb_sync_export Export compressed chunk
kb_sync_import Import chunks
kb_sync_status Sync status
kb_conflict_check Detect conflicts
kb_conflict_judge Record judgment

Comparison

Feature DocsHaven QMD Elasticsearch Context7
Dependencies 1 (mcp) 1 (npm) JVM + plugins External service
Setup time 10 seconds 5 minutes 30+ minutes API key needed
MCP server Built-in No No Yes
URI routing Yes No No No
Conflict detection Yes No No No
Git sync Compressed chunks No No No
Cost Free Free Free (self-hosted) Paid tiers

FAQ

What is DocsHaven?

DocsHaven is a local knowledge base designed for AI agents. It provides full-text search via SQLite FTS5, organizes knowledge by URI domains (core://, ref://, guide://), and detects contradictions when adding new documents. It runs as an MCP server with 16 tools.

How is this different from just using SQLite?

DocsHaven adds a complete knowledge management layer on top of SQLite: automatic document chunking, BM25 ranking with LIKE fallback, URI-based organization, conflict detection, and compressed multi-machine sync — all exposed via MCP tools.

Can I use this with Claude Desktop / Cursor / other AI agents?

Yes. DocsHaven runs as an MCP server. Add it to your MCP client config and all 14 tools become available to your agent.

How fast is search?

SQLite FTS5 with BM25 ranking handles 1,000+ documents in under 100ms on modern hardware. No network latency since everything is local.

Is my data sent anywhere?

No. DocsHaven is fully local. The only network operation is cloning GitHub repositories (which you initiate). All search and storage happens on your machine.

Architecture

docs-haven/
├── server.py        # MCP server (16 tools)
├── storage.py       # SQLite FTS5 backend + Result types
├── uri.py           # URI routing
├── sync.py          # Git sync (compressed chunks)
├── conflicts.py     # Conflict detection
├── result.py        # Ok/Err Result type
├── cli.py           # CLI interface
├── benchmark.py     # Performance benchmarks
├── tests/           # pytest test suite (78 tests)
├── docs/            # Documentation + ADRs
└── pyproject.toml   # Package config

Performance

Benchmarked on Linux (Python 3.14, SQLite FTS5):

Operation Time
Index 1,000 docs 0.076s (13,219 docs/sec)
Search (avg) 3.3ms
Search (P95) 4.3ms
Throughput 299 queries/sec

Run benchmark: python benchmark.py

Integrations

Client Config Status
Claude Desktop claude_desktop_config.json
Cursor .cursor/mcp.json
Gemini CLI gemini mcp add
VS Code (Copilot) .vscode/mcp.json
Codex .codex/config.toml

Development

# Install with test dependencies
pip install -e ".[test]"

# Run tests
pytest tests/ -v

# Run linting
ruff check .

# Run formatting
ruff format .

# Run type checking
mypy . --ignore-missing-imports

Security

  • SQLite FTS5 with parameterized queries (no SQL injection)
  • Input validation on all MCP tool parameters
  • No external network dependencies (local-only operation)
  • Secret scanning via GitHub Actions (gitleaks)

See SECURITY.md for vulnerability reporting.

Author

Built with ❤️ by Cipher208

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docs_haven-0.4.0.tar.gz (55.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docs_haven-0.4.0-py3-none-any.whl (56.1 kB view details)

Uploaded Python 3

File details

Details for the file docs_haven-0.4.0.tar.gz.

File metadata

  • Download URL: docs_haven-0.4.0.tar.gz
  • Upload date:
  • Size: 55.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for docs_haven-0.4.0.tar.gz
Algorithm Hash digest
SHA256 2ee35a9c391c9a581ece7b0f204a7598c34066110de29b1a787482a2a9a51c5e
MD5 d9e98667248582c2d98fbf924993569e
BLAKE2b-256 a149a9d40b0e1785b0f7eb4ae50d501e09cae1fbd86f1eddc4e44247e39b5729

See more details on using hashes here.

File details

Details for the file docs_haven-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: docs_haven-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 56.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.14

File hashes

Hashes for docs_haven-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 55f75abb128945eb38250573e6c6704538ba82a092bc38ccaac9a9af0007ccd3
MD5 34f304f97d1d855ad9dd6ec159c607d7
BLAKE2b-256 88e610ebfc96ced49ecb073facc0e258301270b183799b2386c408ac351c4378

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page