Skip to main content

supermem

Persistent AI memory without RAG — four-tier retrieval that uses an LLM agent only as a last resort, backed by SQLite FTS5, an embedded graph database, and your local markdown vault.

Python 3.11 License: Apache 2.0 MCP

An MCP (Model Context Protocol) server that gives AI assistants — Claude Desktop, LM Studio, ChatGPT — persistent, structured memory backed by SQLite + an optional graph database. The LLM agent is tier 4, not the default path — most queries resolve in milliseconds via full-text search.


Quick Start (Personal, No GPU)

pip install supermem

# Point supermem at a directory of markdown files
export SUPERMEM_VAULT_PATH=~/notes
export SUPERMEM_LLM_PROVIDER=openrouter
export OPENROUTER_API_KEY=your_key_here

# Start the MCP server (add to Claude Desktop's mcp.json)
supermem serve

Add to Claude Desktop mcp.json:

{
  "mcpServers": {
    "supermem": {
      "command": "supermem",
      "args": ["serve"]
    }
  }
}

Quick Start (Production with Docker)

# Clone and configure
git clone https://github.com/lamenting-hawthorn/supermem
cp .env.example .env
# Edit .env: set SUPERMEM_VAULT_PATH, SUPERMEM_LLM_PROVIDER, API keys

# MCP server only (stdio, for Claude Desktop)
docker compose up supermem-mcp

# MCP server + HTTP dashboard
docker compose --profile worker up

# Dashboard at http://localhost:37777

Architecture: Four-Tier Retrieval

Every query goes through tiers in order, short-circuiting when enough results are found. Tiers 1–3 never call an LLM.

Query
  │
  ├─ Tier 1: SQLite FTS5 full-text search          ~1ms    always available
  │          porter tokenizer, WAL mode
  │
  ├─ Tier 2: Kuzu embedded graph expansion         ~5ms    optional (install kuzu)
  │          BFS traversal via [[wikilink]] edges
  │
  ├─ Tier 3: ChromaDB vector similarity            ~50ms   optional (SUPERMEM_VECTOR=true)
  │          sentence-transformer embeddings
  │
  └─ Tier 4: LLM agent fallback                   ~5-30s  always available
             navigates vault via Python sandbox

Short-circuit rule: if tier 1 returns ≥ min_results (default 3), tiers 2–4 are skipped entirely. Unavailable tiers are skipped with a WARNING log — no errors raised.


MCP Tool Reference

Tool Parameters Returns Notes
use_memory_agent query: str Formatted answer Backward-compatible. Routes through all 4 tiers; falls back to full agent only if tiers 1–3 insufficient
supermem_hybrid query: str, tier_limit: int = 4 JSON with obs_ids, source_tier, latency_ms Preferred for programmatic use. Token-efficient — returns IDs first
get_observations ids: list[int] JSON array of observation dicts Fetch full content for specific IDs
get_timeline obs_id: int, window: int = 5 JSON array of chronological observations Context around a specific observation

Progressive Disclosure Pattern

# 1. Search — cheap, returns IDs only
result = await supermem_hybrid("Alice's project status", tier_limit=2)
# {"obs_ids": [42, 17, 88], "source_tier": 1, "latency_ms": 2.1}

# 2. Fetch — only for IDs you actually need
obs = await get_observations([42, 17])
# [{"id": 42, "content": "...", "tier_used": 1}, ...]

# 3. Timeline — context around interesting observations  
ctx = await get_timeline(42, window=3)

Environment Variables

Variable Default Description
SUPERMEM_LLM_PROVIDER openrouter openrouter | ollama | vllm | claude | lmstudio
SUPERMEM_LLM_MODEL provider default Model string (e.g. openai/gpt-4o-mini, llama3)
SUPERMEM_DB_PATH ~/.supermem/supermem.db SQLite database path
SUPERMEM_VAULT_PATH .memory_path file Markdown vault directory
SUPERMEM_VECTOR false Set true to enable ChromaDB tier
SUPERMEM_API_KEY (none) Bearer token for HTTP API auth (disabled if unset)
SUPERMEM_RATE_LIMIT 60 Requests/minute limit
SUPERMEM_WORKER_PORT 37777 HTTP dashboard port
SUPERMEM_COMPRESS_EVERY 50 Observations written before LLM compression
OPENROUTER_API_KEY (required for openrouter) OpenRouter API key
ANTHROPIC_API_KEY (required for claude) Anthropic API key
OLLAMA_HOST http://localhost:11434 Ollama server URL
VLLM_HOST / VLLM_PORT localhost / 8000 vLLM server address
LMSTUDIO_HOST http://localhost:1234 LM Studio server URL

Connector Guide

Import external data into your vault with one command:

# ChatGPT export (Settings → Data controls → Export data → .zip)
supermem connect chatgpt ~/Downloads/chatgpt_export.zip

# Notion workspace export (.zip)
supermem connect notion ~/Downloads/notion_export.zip

# Nuclino workspace export (.zip)
supermem connect nuclino ~/Downloads/nuclino_export.zip

# GitHub repositories (live via API)
supermem connect github owner/repo1,owner/repo2 --token ghp_xxx

# Google Docs (OAuth, opens browser)
supermem connect google_docs "My Doc Name"

All connectors write markdown to your vault, then automatically index the files into SQLite + graph. Private content wrapped in <private>...</private> tags is stripped before indexing.


CLI Reference

supermem serve            # Start MCP server (stdio transport, for Claude Desktop)
supermem serve --worker   # Start MCP server + HTTP dashboard on :37777
supermem chat             # Interactive terminal REPL (no client required)
supermem backup           # Create timestamped .tar.gz (vault + SQLite)
supermem backup --output /path/to/archive.tar.gz
supermem restore <archive.tar.gz>
supermem connect <type> <source> [--token TOKEN] [--max-items N]

HTTP Dashboard (Optional)

Start with supermem serve --worker or docker compose --profile worker up.

Endpoint Method Description
/health GET {"status":"ok","db":true,"graph":false,"vector":false}
/sessions GET Paginated session list with summaries
/observations GET Filter by session/date/type
/search POST {"query": "...", "tier_limit": 4}
/index/rebuild POST Reindex entire vault
/backup GET Streams vault + DB as .tar.gz
/stats GET {obs_count, entity_count, session_count, db_size_mb}

Auth: Authorization: Bearer <SUPERMEM_API_KEY>. Disabled when env var is unset.


Privacy

Wrap sensitive content in <private>...</private> tags. It is stripped before writing to any storage layer (SQLite, Kuzu, ChromaDB). The content passes through to the agent sandbox only — it never persists.

# Meeting Notes

Alice discussed the roadmap.
<private>Budget: $2.4M approved for Q3</private>
Next steps: ship v2 by June.

Running Tests

uv run pytest tests/ -v                          # all tests
uv run pytest tests/unit/ -v                     # unit only (fast, no network)
uv run pytest tests/integration/ -v              # integration (real storage)
uv run pytest tests/ --cov=supermem --cov-report=term-missing  # with coverage

Coverage gate: 60% (CI enforced). Kuzu and Anthropic tests are auto-skipped if packages are not installed.


License

Apache 2.0 — see LICENSE.

Release files for supermem 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for supermem 0.3.1
File Size Uploaded
supermem-0.3.1.tar.gz 498.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for supermem 0.3.1
File Interpreter ABI Platform
supermem-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 656.8 kB

Release files / supermem-0.3.1.tar.gz

Download URL supermem-0.3.1.tar.gz
Size 498.9 kB
Tags Source
SHA-256 checksum
How to use checksums
c0e48e49870b2ebec9b9941281cb46480e85ad79df19fae18229fd30da380164
BLAKE2b-256 checksum
How to use checksums
610d07971febf05dc8e2fa85489330a99b38e5dcd4872a03c61d9cae03b0fa2d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 10, 2026.

Transparency log

Release files / supermem-0.3.1-py3-none-any.whl

Download URL supermem-0.3.1-py3-none-any.whl
Size 157.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
24ae796a67ed65213564a17a6d5f6811dadeae77726b922263d82b143c213516
BLAKE2b-256 checksum
How to use checksums
c8ea4fa6c7b1332d9c161653a1fc323e33c1a50568b7c38d0b6f2e038171125b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page