Skip to main content

Production-grade memory reliability system for AI coding agents. Lifecycle-managed claims with citations, conflict detection, steward governance, and MCP integration.

Project description

MemoryMaster

Production-grade memory reliability system for AI coding agents.

Lifecycle-managed claims with citations, conflict detection, steward governance, hybrid retrieval, and MCP integration. Give your AI agents persistent, trustworthy memory.

License: MIT Python 3.10+ Tests MCP Tools CLI Commands PyPI

MemoryMaster prevents the #1 problem with agent memory: drift, stale assumptions, and unsafe disclosure. It gives Claude Code, Codex, and any MCP-compatible agent persistent, verifiable memory with a full claim lifecycle, citation tracking, conflict detection, and human-in-the-loop governance.

How it's different

Most agent-memory systems (mem0, Letta/MemGPT, Zep) optimize for storing and recalling more — embeddings, summaries, a fast vector store. MemoryMaster optimizes for trusting what you recall. The differentiator is governance: every memory is a lifecycle-managed claim, not an opaque embedding.

mem0 / Letta / Zep MemoryMaster
Unit of memory text chunk / summary claim with status, tier, citations, bitemporal validity
Stale / wrong facts linger until overwritten decay → stale, conflict detection, supersession
Contradictions silently coexist surfaced as conflicted, auto-resolved (5-tier) or queued for review
Provenance usually none citation per claim + per-agent provenance
Secret leakage your problem sensitivity filter at ingest (JWT/AWS/Bearer/SSH redaction)
Operator control API only steward governance + a dashboard you can see

If you want an agent that recalls more, any vector store works. If you want an agent that recalls correctly — and can prove where a fact came from and retire it when it goes stale — that's the gap MemoryMaster fills.


Architecture

MemoryMaster is layered around MCP/CLI entry points, the MemoryService facade, SQLite/Postgres storage, optional Qdrant vector search, scheduled jobs, and the Obsidian wiki/vault layer. The canonical ingest path is:

MCP/CLI -> sensitivity filter -> MemoryService.ingest -> store write -> FTS5 index

The query path is:

query_memory -> MemoryService.query -> storage reads + optional Qdrant candidates -> ranked context

See docs/architecture.md for the current module map, data-flow details, recent PR status, and sensitivity-filter invariants.

Key features

  • 6-state lifecycle: candidateconfirmedstalesupersededconflictedarchived
  • Citation tracking with provenance for every claim
  • Hybrid retrieval: vector (sentence-transformers / Gemini) + FTS5 + freshness + confidence
  • Context optimizer: query_for_context(budget=4000) returns auto-curated memory that fits your token budget
  • Entity graph with typed relationships and alias resolution
  • Rule-shaped claims (new in v3.21.0): prescriptive when <trigger>, do <action> because <rationale> claims (ingest_rule / query_rules) — the shape an agent needs to actually change behaviour next time, not just recall a fact
  • Correction mining (new in v3.21.0): mine-rules scans the verbatim transcript archive for user corrections and distills them into rule claims; the Stop hook also mines each session's latest correction automatically
  • Versioned schema migrations (new in v3.20.0): migrate applies SQLite/Postgres migrations with sha256 drift detection; incremental export-delta ships small claim deltas for cheap cross-machine sync
  • Retrieval quality (new in v3.22.0): floor-ratio boost gate (MEMORYMASTER_BOOST_FLOOR_RATIO) stops fresh-but-wrong claims outranking the true match; query --explain shows per-stage score attribution; an opt-in correctness-safe query cache (MEMORYMASTER_QUERY_CACHE) with a generation gate
  • Semantic contradiction probe (new in v3.22.0, wired as a steward phase in v3.23.0): detect-contradictions finds claims that genuinely contradict each other (beyond the deterministic same-subject conflict check) via an LLM judge with a Wilson-CI rate and verdict cache; in v3.23 the same probe runs inside run-steward and emits paste-ready conflicted proposals
  • Verbatim archive cleanup (new in v3.23.0): verbatim-cleanup dedups the raw-transcript table and optionally purges pre-#128 junk rows, with a dry-run default and FTS5 mirror sync
  • Steward governance: multi-probe validators (filesystem, format, citation, semantic, tool) with proposal review
  • Conflict resolution: 5-tier auto (confidence > freshness > citations > LLM > manual)
  • Auto-redaction at ingest: JWT, GitHub tokens, Bearer, AWS keys, SSH keys, custom patterns
  • LLM Wiki: compiled-truth + append-only timeline articles with progressive-disclosure frontmatter, explored: true|false operator-review marker, and inline > [!contradiction] Obsidian callouts
  • Atlas Inbox V1 (new in v3.13.0): WhatsApp ingestion → source/evidence/action proposal lifecycle → Super-Productivity export. Versioned API/CLI contract for downstream consumers (LifeAgent, etc.) — see docs/atlas-api-contract-v1.md. Real provider adapters (OpenAIWhisperTranscriptionProvider, TesseractOcrProvider) behind Protocols; mock providers stay default.
  • Dual backend: SQLite (zero-config) and Postgres (full feature parity with pgvector)
  • Dream Bridge for bidirectional sync with Claude Code's Auto Dream
  • 7-hook stack: recall, classify, validate-wiki, session-start, auto-ingest, precompact, steward-cron

Full feature index lives in docs/handbook.md.

Benchmarks

LongMemEval-S (N=500, retrieval-only) — v3.15.0 now leads the publicly-reported numbers from agentmemory on R@5 and MRR, after wiring sentence-transformers/all-MiniLM-L6-v2 into the bench harness (the v3.14 baseline was unintentionally BM25-only).

LongMemEval-S benchmark

Metric v3.14.0 v3.15.0 agentmemory Δ vs agentmemory
Recall@5 0.894 0.966 0.952 +0.014
Recall@10 0.942 0.984 0.986 -0.002
MRR 0.799 0.902 0.882 +0.020

Reproduce: python tests/bench_longmemeval.py --retrieval-only. Full methodology, experiment-by-experiment deltas (1 KEEP, 2 REVERT, 3 NULL), and the architectural findings that surfaced along the way live in docs/archive/longmemeval-results.md and docs/archive/v315-experiments/. QA-accuracy pass (with judge) is deferred until provider quotas allow.

Prerequisites

Required (the package won't function without these)

  • Python 3.10+ with pip
  • Claude Code, Codex, or any MCP-compatible agent
  • An LLM provider — pick one: Claude Code OAuth (free if you're a subscriber, set MEMORYMASTER_LLM_PROVIDER=claude_cli), a free Gemini API key from aistudio.google.com, OpenAI, Anthropic API, or local Ollama. The steward, auto-ingest, and wiki-absorb cycles all need an LLM — without one, claims pile up as candidate and never get validated, deduped, or compiled into the wiki.

Strongly recommended (you'll lose ~80% of the value without these)

  • Node.js 18+ for graphify and GitNexus — these are the cached intelligence layers that make MemoryMaster cheap to query. Without them, every "what does this codebase do?" question burns tokens cold-exploring files the graph already mapped. The intelligence-first workflow in CLAUDE.md assumes both are installed.
  • Obsidian 1.6+ with the Bases core plugin — the wiki engine writes plain Markdown so any editor works, but Obsidian's backlinks, graph view, and Bases dashboards are how you actually navigate wiki-absorb output. Without Obsidian, the wiki is just a folder of files.

Optional (nice to have)

  • Docker for Qdrant — vector retrieval. SQLite FTS5 is the default and works out of the box; add Qdrant when you want semantic recall on top of keyword search.

15-minute quickstart

From zero to a recalled claim and a live dashboard. No Qdrant, no Postgres, no LLM key required for these steps (SQLite + FTS5 is the default).

1. Install (2 min)

pip install "memorymaster[mcp]"
memorymaster --db memorymaster.db init-db

2. Configure a provider (3 min, optional for this walkthrough)

Recall and ingest below work with zero config. An LLM provider is only needed for the steward/wiki cycles — pick one when you're ready (see Pick your LLM provider). For a Claude Code subscriber, the cheapest path is:

export MEMORYMASTER_LLM_PROVIDER=claude_cli   # reuses your Claude Code OAuth, no API key

3. Ingest a claim via CLI (1 min)

memorymaster --db memorymaster.db ingest \
  --text "Server uses PostgreSQL 16" \
  --source "session://chat|turn-3|user confirmed"

4. Recall it (1 min)

A freshly-ingested claim starts life as a candidate (unvalidated). The CLI query/context paths exclude candidates by default — that's the governance model: unvalidated facts don't silently leak into recall until the steward promotes them. To see your brand-new claim before a validation cycle, pass --include-candidates:

# Hybrid retrieval (lexical + freshness + confidence)
memorymaster --db memorymaster.db query "database version" \
  --retrieval-mode hybrid --include-candidates

# Token-budgeted context block — the killer feature for agents
memorymaster --db memorymaster.db context "database" \
  --budget 4000 --format xml --include-candidates

You should see the PostgreSQL 16 claim come back, ranked, with its citation. (Drop --include-candidates and you'll get zero results until step 7's run-cycle promotes it to confirmed — that's working as designed, not a bug.)

5. Open the dashboard (2 min)

memorymaster --db memorymaster.db run-dashboard   # serves on http://127.0.0.1:8765

Open the URL: you'll see your claim in Claims, plus governance panels — Conflicts, Review Queue, Recall Analysis (why each claim ranked where it did), Audit Log, Provenance by Agent, and Reliability.

6. Wire it into your agent (3 min)

memorymaster-setup     # interactive: hooks, MCP, steward cron, CLAUDE.md / AGENTS.md

That installs the MCP server and the auto-ingest Stop hook so your agent recalls and stores memory automatically. See MCP server for the config block.

7. Run a validation cycle (1 min, needs a provider)

memorymaster --db memorymaster.db run-cycle   # extract, validate, decay, compact

For the one-prompt agent install (paste into any agent with shell access), see docs/handbook.md#one-prompt-agent-install.

Pick your LLM provider

Provider Env vars Default model Cost
Claude Code OAuth (recommended for subscribers) MEMORYMASTER_LLM_PROVIDER=claude_cli (requires claude CLI on PATH) claude-haiku-4-5-20251001 included in Claude Code plan
Google Gemini (default) MEMORYMASTER_LLM_PROVIDER=google + GEMINI_API_KEY=... gemini-3.1-flash-lite-preview ~free
OpenAI MEMORYMASTER_LLM_PROVIDER=openai + OPENAI_API_KEY=... gpt-4o-mini ~$0.001/call
Anthropic API MEMORYMASTER_LLM_PROVIDER=anthropic + ANTHROPIC_API_KEY=... claude-haiku-4-5-20251001 ~$0.001/call
Ollama (local) MEMORYMASTER_LLM_PROVIDER=ollama + OLLAMA_URL=http://localhost:11434 llama3.2:3b free

The claude_cli provider shells out to your local claude --print binary, so it inherits the OAuth session you're already logged into in Claude Code — no API key, no rotator, no quota juggling. Caveat: cold-start adds 3-15s per call (subprocess spawn), so it's ideal for batched/cron paths (steward, wiki-absorb) and not for latency-sensitive recall. Override with MEMORYMASTER_CLAUDE_CLI_BIN and MEMORYMASTER_CLAUDE_CLI_TIMEOUT. On VM installs the OAuth token expires ~24h, so pair with MEMORYMASTER_LLM_FALLBACK_PROVIDER=ollama; desktop tokens don't expire.

For zero-cost offline use, install Ollama, ollama pull llama3.2:3b, and set MEMORYMASTER_LLM_PROVIDER=ollama.

MCP server

{
  "mcpServers": {
    "memorymaster": {
      "command": "memorymaster-mcp",
      "env": {
        "MEMORYMASTER_DEFAULT_DB": "/path/to/memorymaster.db",
        "MEMORYMASTER_WORKSPACE": "/path/to/your/project"
      }
    }
  }
}

30 MCP tools spanning setup/lifecycle, ingest, query/retrieval, listing, knowledge graph, and governance: init_db, ingest_claim, ingest_rule, query_rules, rules_export, run_cycle, run_steward, classify_query, query_memory, query_for_context, query_for_task, query_claim_paths, query_meta_decisions, federated_query, recall_analysis, read_active_tasks, list_claims, redact_claim_payload, pin_claim, compact_memory, list_events, search_verbatim, open_dashboard, list_steward_proposals, resolve_steward_proposal, extract_entities, entity_stats, find_related_claims, quality_scores, recompute_tiers.

See docs/MCP-TOOLS.md for the grouped reference (one line per tool), and .mcp.json.example for the full config template.

Backends

Backend Install Use case
SQLite Built-in Local development, single-agent, zero-config
Postgres pip install "memorymaster[postgres]" Team deployment, multi-agent, pgvector search

Docker Compose

Run the full stack (MemoryMaster + Qdrant + Ollama) with one command:

docker compose up -d

See INSTALLATION.md for Kubernetes / Helm.

Development

# Install with dev dependencies
pip install -e ".[dev,mcp,security,embeddings,qdrant]"

# Run tests
pytest tests/ -q

# Lint and format
ruff check memorymaster/ && ruff format memorymaster/

# Performance benchmarks
python benchmarks/perf_smoke.py

See CONTRIBUTING.md for the full workflow.

Documentation

Document Description
docs/README.md Documentation index — where to find each living doc
docs/handbook.md Full operator handbook — hooks, dashboard, steward, dream bridge, troubleshooting, one-prompt install
docs/MCP-TOOLS.md Reference for all 30 MCP tools, grouped by purpose
docs/INTEGRATING.md Integration guide for embedding MemoryMaster in your agent
INSTALLATION.md Setup guide: pip, Docker, Helm, MCP config
CONTRIBUTING.md Dev setup, testing, PR workflow
ARCHITECTURE.md System design and subsystem details
USER_GUIDE.md Usage, MCP integration, troubleshooting
CHANGELOG.md Version history and release notes
ROADMAP.md Release plan and future tracks
docs/enabling-v2-systems.md v3 statistical classifier + cadence policy opt-in

License

MIT — Built by wolverin0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

memorymaster-4.0.0.tar.gz (1.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

memorymaster-4.0.0-py3-none-any.whl (673.5 kB view details)

Uploaded Python 3

File details

Details for the file memorymaster-4.0.0.tar.gz.

File metadata

  • Download URL: memorymaster-4.0.0.tar.gz
  • Upload date:
  • Size: 1.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for memorymaster-4.0.0.tar.gz
Algorithm Hash digest
SHA256 6b416ba5623f2aeb9a3175f6a452caa0a4e444a552d99094fa1bea06f71d1654
MD5 914a56b22e90a00c472a76a1b5db745f
BLAKE2b-256 fcd9693f3e73c080802466d92331752d882269b0a34dc3f7d3d3b86eaa7ced92

See more details on using hashes here.

Provenance

The following attestation bundles were made for memorymaster-4.0.0.tar.gz:

Publisher: publish.yml on wolverin0/memorymaster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file memorymaster-4.0.0-py3-none-any.whl.

File metadata

  • Download URL: memorymaster-4.0.0-py3-none-any.whl
  • Upload date:
  • Size: 673.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for memorymaster-4.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 91c31c1c44478d053ac51fbeca24988b1b34883d454ee5f33879ca55c80ac09c
MD5 0e4413fe0b6a895c5874fb102ab602b6
BLAKE2b-256 202a56fcb35421116a357df6d129fdff310c2d7e51bc846846a0a3ac30e94d80

See more details on using hashes here.

Provenance

The following attestation bundles were made for memorymaster-4.0.0-py3-none-any.whl:

Publisher: publish.yml on wolverin0/memorymaster

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page