Skip to main content

Developer memory tool — mine codebases and conversations into a LanceDB-backed searchable palace. No API key required.

Project description

mempalace-code

mempalace-code

Your AI's long-term memory. Local. Instant. Private.

Index your codebase once. Your AI can recall architecture decisions, debugging sessions, and API patterns across sessions and projects without re-reading the repo.

No cloud service, no API keys, no subscription. After the one-time embedding model download, indexing and search stay on your machine.


Get Started in 30 seconds · How It Works · All Features · Benchmarks


Language-Aware Mining
AST, regex, and adaptive chunking
matched to each file type
29 MCP Tools
Any MCP-capable agent
search, store, traverse · static profiles
Temporal Knowledge Graph
Facts that change over time
with validity windows
595x Token Savings
measured peak · median 80x
scales with project size
Cross-Project Tunnels
Search auth in one project
find it everywhere
2200+ Tests · $0 Cost
Every feature acceptance-gated
offline after model setup

Quick Start

uv tool install mempalace-code        # recommended (fast, Rust-based)
# or
pipx install mempalace-code           # alternative
# or
pip install mempalace-code            # into current environment
# or
uvx --from mempalace-code mempalace-code --help  # try without installing

mempalace-code is the default command name so this fork can coexist with upstream/vanilla mempalace on the same machine. If mempalace is unused on your PATH and you want the shorter alias, run mempalace-code install-alias. Packaged installs use the Python import package mempalace_code, so they can coexist with vanilla MemPalace in the same Python environment. Source checkouts keep a small mempalace.mcp_server shim only so older repo-local MCP configs that run with PYTHONPATH=/path/to/mempalace-code continue to start.

Then ask your AI to read docs/AGENT_INSTALL.md — it will handle setup, MCP wiring, prompt injection, and verification automatically.

Or do it manually
mempalace-code init ~/projects/myapp       # detect rooms, cache embedding model (~80 MB)
mempalace-code init ~/projects/myapp --detect-entities  # optional people/project detection for notes/convos
mempalace-code mine ~/projects/myapp       # index your codebase
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server  # connect Claude Code
codex mcp add mempalace-code -- python -m mempalace_code.mcp_server   # connect Codex CLI

Optional: auto-sync on commit (requires [watch] extra — see Auto-Watch):

mempalace-code watch ~/projects/           # re-mines on every commit, zero noise

This makes all 29 tools available to your AI. To expose a reduced subset, add a --profile flag (e.g. -- python -m mempalace_code.mcp_server --profile=minimal). For proactive search and storage (without you asking), you'll also need to add usage rules to your CLAUDE.md (or equivalent agent-instruction file) — copy from docs/LLM_USAGE_RULES.md, or let docs/AGENT_INSTALL.md Section 7 inject them for you.

Do you need it? Ask your LLM:

Read:
- https://github.com/rergards/mempalace-code/blob/main/README.md
- https://github.com/rergards/mempalace-code/blob/main/docs/AGENT_INSTALL.md
- https://github.com/rergards/mempalace-code/blob/main/docs/LLM_USAGE_RULES.md

Evaluate mempalace-code for this repo specifically, not generic memory tooling.

1. Current context burn: what do I re-explain, and how many tokens/week are spent re-reading files?
2. Highest-ROI surfaces: rank code, docs/prose, design notes, conversation/log exports, manual decisions, KG facts, architecture relationships.
3. Stack support: list supported languages/file types actually indexed; name unsupported extensions that are skipped by default. Distinguish tree-sitter AST, regex structural, Markdown/prose, Kubernetes YAML, and adaptive chunking.
4. Agent workflow fit: which MCP tools should you use proactively (`mempalace_search`, `mempalace_code_search`, `mempalace_file_context`, KG, tunnels, diary, manual drawers)?
5. Ops fit: recommend global MCP, project-scoped MCP, or no MCP; cover 29-tool prompt surface and profile subsetting options (`--profile`, `--tools`, `--exclude`), `mine` vs `mine-all`, `watch` cadence, backup/repair, and offline/model constraints.
6. Decision: install now, try scoped for a week, wait for a named feature, or skip. Give the first 3 commands you would run.
7. If waiting: ask me whether to draft a GitHub feature request.

For repos that already have a hand-curated memory file, evaluate mempalace-code as a complement, not a replacement. A good trial order is:

  1. KG first — move volatile current facts (active branch, phase, owner, deadline, status) into temporal triples; keep curated prose for reasoning and narrative.
  2. Drawers + docs second — mine design docs, specs, long decision notes, and conversation/log exports where semantic retrieval beats manual grep.
  3. Code mining last — start with one high-value subproject, then expand if agents actually use the results.

Cost caveat: the MCP server defaults to all 29 tools. Use --profile=minimal or --tools=search,add_drawer at startup to reduce the prompt/tool-surface cost. Proactive use also depends on adding the usage-rules block to agent instructions. Prefer project-scoped MCP for trials, and keep it only if searches, KG lookups, or drawer writes show up in real sessions.

Supported MCP Clients

mempalace-code works with any MCP-compatible client:

  • Claude Code (CLI, desktop, web) — claude mcp add mempalace-code -- python -m mempalace_code.mcp_server
  • Codex CLIcodex mcp add mempalace-code -- python -m mempalace_code.mcp_server
  • Claude Desktop — add to claude_desktop_config.json
  • Cursor — add as MCP server in settings
  • Windsurf — add as MCP server in settings
  • Any MCP client — point it at python -m mempalace_code.mcp_server (stdio transport)

For local models without MCP support (Llama, Mistral, etc.), use mempalace-code wake-up to pipe context into the system prompt — see Memory Layers.


How It Actually Works

You write code. You make decisions. You debug things. Between sessions, all that context vanishes.

mempalace-code indexes it once into a local vector store, then your AI finds it in milliseconds — using 595x fewer tokens than grep + read at measured peak (median 80x on a 19k-chunk project, and it keeps scaling). Think of it as git log for everything that isn't in the code: the why, the discussions, the dead ends, the decisions.

What gets indexed or stored:

  • Code files — structural chunks for Python, TypeScript/JS/TSX/JSX, Go, Rust, Java, Kotlin, C#, F#, VB.NET, XAML, Swift, PHP, Scala, Dart, Lua, Ruby, Terraform/HCL, Markdown, Kubernetes manifests, Helm charts/templates, and Ansible playbooks/roles/inventory; adaptive chunks for C/C++, shell, SQL, HTML/CSS, JSON/YAML/TOML, CSV, Dockerfile, Make, templates, and config files
  • .NET solutions — .sln/.csproj project graphs, cross-project symbol relationships, interface implementations
  • Architecture facts — pattern, layer, namespace, and project membership facts for .NET and Python projects
  • Conversation/log exports — Claude Code JSONL, OpenAI Codex CLI JSONL, Gemini CLI JSONL, Claude.ai JSON, ChatGPT conversations.json, Slack JSON, plain text transcripts
  • MCP-saved context — manual drawers are vector-indexed; diary entries and temporal KG facts are stored in their own retrieval surfaces
  • Architecture notes, decisions, anything else you store manually

Generated helper files such as entities.json are skipped during project mining by default, because they are created by init/entity detection and should not become source-code drawers unless explicitly force-included.

How you use it: After setup, your AI calls mempalace tools automatically. You don't type search commands.


Features

Language-Aware Code Mining

mempalace-code mine walks your source tree and chooses the best chunker for each file type: AST boundaries where optional tree-sitter grammars are available, regex structural boundaries for supported languages, YAML-aware Kubernetes/Helm/Ansible resource splits, Markdown/prose sections, or adaptive line-count chunks for formats without reliable declarations. The shared catalog currently exposes 45 searchable language labels to code_search(language=...). Leading comments and docstrings stay attached to declarations where structural chunking is active; Markdown drawers keep heading path, section type, and Mermaid/code/table flags in search metadata.

Language Strategy AST Support
Python Functions, classes, methods, decorators Optional tree-sitter
TypeScript / JavaScript / TSX / JSX Functions, classes, exports, imports Optional tree-sitter
Go Functions, types, methods, interfaces Optional tree-sitter
Rust Functions, structs, enums, traits, impls Optional tree-sitter
Java Classes, interfaces, methods, annotations Regex
Kotlin Classes, objects, functions, extensions Regex
Scala Classes, case classes, objects, traits, enums, functions, implicits, type aliases, generics Regex
Swift Classes, structs, enums, protocols, functions, properties, extensions, actors, async/await Regex
Dart Classes, mixins, extensions, enums, functions, named/factory constructors, async/await Regex
PHP Classes, interfaces, traits, enums (8.1+), functions, methods, namespaces (Laravel/WP/Symfony aware) Regex
C# Classes, interfaces, records, methods, properties Regex
F# / VB.NET Modules, types, functions Regex
XAML Controls, resources, code-behind linking Regex
Terraform / HCL Terraform/HCL top-level blocks (resource, module, variable, moved, import, check, etc.) Regex
Kubernetes manifests Deployments, Services, ConfigMaps, Secrets, Ingresses, CRDs (indexed by kind/name) YAML-aware
Helm charts Chart.yaml, values*.yaml, raw templates with kind/name metadata; no template rendering YAML/Go-template aware
Ansible Playbooks, role tasks/handlers/defaults/vars, inventories; no Jinja evaluation or inventory semantics YAML/Jinja tolerant
Markdown / plain text Heading sections (#-######), heading paths, section metadata, paragraphs
Lua Functions, local functions, methods (dot/colon), module/table declarations Regex
C / C++ Indexed and searchable with best-effort symbol metadata; chunked adaptively today
Ruby Static classes, modules, methods, singleton methods, attrs, and constants; Rails DSL/metaprogramming not interpreted Regex
shell / SQL Indexed and searchable; chunked adaptively today
HTML / CSS / CSV Indexed and searchable; chunked adaptively today
YAML / JSON / TOML Adaptive line-count; Kubernetes YAML auto-detected separately
Dockerfile / Make / templates / config Dockerfile, Containerfile, Makefile, GNUmakefile, Vagrantfile, Go templates, Jinja2, .conf, .cfg, .ini

The mempalace_code_search language filter is generated from the same language catalog as the miner. If a file type is mined with a language label, the MCP schema and unsupported-language hints stay aligned with that catalog.

Tree-sitter is optional (pip install "mempalace-code[treesitter]"). When a grammar is missing, Python, TypeScript/JavaScript/TSX/JSX, Go, and Rust fall back to regex structural chunking. Other recognized formats use their regex, YAML-aware, prose, or adaptive chunker as listed above.

Extensions outside the miner catalog are skipped by normal project scans unless you explicitly force-include an exact path with --include-ignored path/to/file.

mempalace-code mine ~/projects/myapp                  # all supported file types
mempalace-code mine ~/projects/myapp --wing myapp     # tag with a specific wing
mempalace-code mine ~/chats/ --mode convos            # mine conversation exports
mempalace-code mine-all ~/projects/                   # sync all projects incrementally (one wing per project)
mempalace-code mine-all ~/projects/ --new-only        # skip projects whose wing already exists (first-run only)

Mining is incremental by default — content-hash based, only changed files are re-chunked. Use --full to force a rebuild.

Multi-project wing namingmine-all assigns one wing per project using this priority:

  1. wing: in the project's mempalace.yaml (explicit override)
  2. Git origin repo name (e.g. my-repo.gitmy_repo)
  3. Normalized folder name

If two projects resolve to the same wing name, mine-all exits with an error before mining anything. Fix this by adding a unique wing: value to each project's mempalace.yaml. Use --new-only to skip projects already present in the palace (useful for first-run batch ingestion).

Optional Entity Detection

mempalace-code init <dir> is config-first by default: it detects rooms from the directory structure and does not scan file contents for names. Add --detect-entities only when the directory contains prose where people or project names matter, such as meeting notes, client notes, personal notes, or conversation exports:

mempalace-code init ~/notes --detect-entities        # prompts to confirm detected people/projects
mempalace-code init ~/notes --detect-entities --yes  # auto-accept entity confirmation (no room prompts)

The detector is a lightweight bootstrap step, not the main miner. It samples up to 10 readable files, prefers prose files (.md, .txt, .rst, .csv), reads the first 5 KB of each sampled file, and looks for heuristic signals such as Alice said, thanks Bob, Apollo repo, deploy Apollo, or import Apollo. Confirmed results are written to <dir>/entities.json:

{
  "people": ["Alice", "Bob"],
  "projects": ["Apollo"]
}

Use it for human/project context. Leave it off for normal code repos unless their docs contain the entities you want captured. Full-repo scanning would be slower and noisier: class names, packages, examples, and variables often look like people or products to a heuristic pass. Code structure, symbols, languages, and architecture relationships are handled by mempalace-code mine, not by entity detection.

Auto-Watch

Keep your palace in sync automatically. By default, watches .git/refs/heads/ and re-mines only on commit — no noise from work-in-progress saves. Handles multiple branches and worktrees.

Requires the watch extra:

uv tool install "mempalace-code[watch]"   # or: pipx install "mempalace-code[watch]"

Already installed without it? Add watchfiles:

uv tool inject mempalace-code watchfiles  # or: pipx inject mempalace-code watchfiles
mempalace-code watch ~/projects/my-app                # watch an initialized project (on commit)
mempalace-code watch ~/projects/                      # watch all initialized projects in a parent directory
mempalace-code watch ~/projects/ --on-save            # watch all file saves instead (noisier)
mempalace-code watch ~/projects/ schedule             # print launchd/cron snippet for daemon

watch accepts either an initialized project directory (has mempalace.yaml) or a parent directory containing immediate initialized project subdirectories. Pointing it at a project root that has project files but no mempalace.yaml exits with the correct mempalace-code init <dir> command.

Install as persistent daemon (macOS):

mempalace-code watch ~/projects/ schedule > ~/Library/LaunchAgents/com.mempalace.watch.plist
launchctl load ~/Library/LaunchAgents/com.mempalace.watch.plist

Starts at login, restarts if crashed. Logs to /tmp/mempalace-watch.log.

Disk-budget guard: the daemon automatically skips mine/optimize cycles when free disk space falls below the configured floor (default 1 GiB). Use watch status to check the current state:

mempalace-code watch ~/projects/ status      # print disk-budget summary + launchd state

To pause or stop the daemon when disk is low:

launchctl unload ~/Library/LaunchAgents/com.mempalace.watch.plist   # stop until next login
launchctl load   ~/Library/LaunchAgents/com.mempalace.watch.plist   # re-enable after freeing space

Diagnosing and stopping a crash-looping job:

If the daemon was pointed at an uninitialized directory it will crash-loop under KeepAlive. Confirm with:

mempalace-code watch ~/projects/ status   # shows state, runs count, and last exit code

Stop and optionally remove it:

# macOS 10.11+ preferred — unregisters the job immediately:
launchctl bootout gui/$(id -u)/com.mempalace.watch

# Older macOS / alternative:
launchctl unload ~/Library/LaunchAgents/com.mempalace.watch.plist

# Remove permanently (re-install after fixing the watch root):
rm ~/Library/LaunchAgents/com.mempalace.watch.plist

Daemon health check:

The daemon emits WATCH_RUN lines at each startup transition so the appended log at /tmp/mempalace-watch.log can be searched to confirm the latest startup reached the watch loop. Use this sequence to diagnose daemon state:

# 1. Process state — is the daemon running?
launchctl print gui/$(id -u)/com.mempalace.watch
# Or use the watch status shortcut:
mempalace-code watch ~/projects/ status

# 2. Palace storage health
mempalace-code --palace ~/.mempalace/palace health

# 3. Find the latest startup that reached watch-ready
grep -a 'state=watch-ready' /tmp/mempalace-watch.log | tail -1
# Example output: WATCH_RUN run_id=20260616T120102Z-p12345 state=watch-ready

# 4. Filter that run's context (replace the run_id from step 3):
grep -a 'run_id=20260616T120102Z-p12345' /tmp/mempalace-watch.log

A log file may contain WATCH_RUN lines from older runs that exited with disk-budget or backup failures. The run_id on the latest state=watch-ready line identifies the current healthy startup — lines from prior runs with different run_id values are stale and can be ignored. If no watch-ready line appears, check the most recent run-started line and the state that followed it (for example state=pre-watch-backup-failed or state=initial-mine-skipped reason=disk-budget).

Configure the threshold via environment variable or ~/.mempalace/config.json:

# Environment variable (bytes or human suffix)
export MEMPALACE_WATCH_DISK_MIN_FREE_BYTES=2GiB    # watcher-specific floor
export MEMPALACE_DISK_MIN_FREE_BYTES=1GiB          # global floor (watcher + backup)

# ~/.mempalace/config.json
{
  "disk_min_free_bytes": 1073741824,         // 1 GiB global default
  "watch_disk_min_free_bytes": 2147483648,   // 2 GiB for watcher specifically
  "backup_disk_min_free_bytes": 2147483648   // 2 GiB for backups specifically
}

MCP read/search behavior is not affected by a paused watcher — agents can still search the palace while the daemon waits for disk space.


The Palace

mempalace-code organizes memories into a navigable structure — the same mental model ancient Greek orators used to memorize speeches.

  ┌─────────────────────────────────────────────────────────────┐
  │  WING: myapp                                               │
  │    ┌──────────┐  ──hall──  ┌──────────┐                    │
  │    │  backend │            │  frontend│                    │
  │    └────┬─────┘            └──────────┘                    │
  │         ▼                                                  │
  │    ┌──────────┐      ┌──────────┐                          │
  │    │  Closet  │ ───▶ │  Drawer  │  (verbatim content)     │
  │    └──────────┘      └──────────┘                          │
  └─────────┼──────────────────────────────────────────────────┘
            │ tunnel (auto-created when room names match)
  ┌─────────┼──────────────────────────────────────────────────┐
  │  WING: otherapp                                            │
  │    ┌────┴─────┐  ──hall──  ┌──────────┐                    │
  │    │  backend │            │  infra   │                    │
  │    └──────────┘            └──────────┘                    │
  └─────────────────────────────────────────────────────────────┘
Concept What it is
Wing A project, person, or domain. As many as you need.
Room A topic within a wing: backend, auth, deploy, decisions.
Drawer Verbatim content. Never summarized, never rewritten.
Hall Connection between rooms in the same wing.
Tunnel Auto-connection between wings when the same room name appears.

MCP Server — 29 Tools {#mcp-tool-profiles}

claude mcp add mempalace-code -- python -m mempalace_code.mcp_server

The MCP server registration name defaults to mempalace-code. The MCP tool identifiers remain mempalace_* for compatibility with existing agents and usage rules.

By default all 29 tools are exposed. Use startup flags to reduce the tool surface (GitHub issue #6 — static profiles lower prompt cost while preserving stable named-tool trigger patterns in usage rules):

# Named profiles — select a pre-defined subset at server startup
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=minimal
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=kg
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=code
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=notes

# Explicit tool list (replaces profile base set)
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --tools=search,add_drawer,diary_*

# Add or remove tools from a profile
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=minimal --include=kg_query
claude mcp add mempalace-code -- python -m mempalace_code.mcp_server --profile=full --exclude=delete_wing,delete_drawer
Profile Tools Best for
full (default) all 29 Full capability; no surface reduction
minimal 4 Search + store only
kg 8 Minimal + temporal knowledge graph
code 10 Code archaeology; no drawer-write/diary tools (mine included)
notes 12 Knowledge management + diary; no code-search

Selector rules for --tools, --include, --exclude:

  • Accept full names (mempalace_search), short names (search), or wildcards (diary_*).
  • --tools replaces the profile base set; cannot be combined with --include.
  • --include adds to the profile base set; --exclude removes last (wins over everything).
  • Invalid profile name, unknown selector, or empty result → process exits with nonzero status and a stderr message.
Palace — Read
Tool What
mempalace_status Palace overview — total drawers, wings, rooms
mempalace_list_wings All wings with drawer counts
mempalace_list_rooms Rooms within a wing
mempalace_get_taxonomy Full wing → room → count tree
mempalace_search Semantic search with optional wing/room filters; Markdown hits include heading path and section metadata
mempalace_code_search Filter by language, symbol name/type, file glob; optional rerank="hybrid"
mempalace_file_context All indexed chunks for a source file, ordered by chunk_index
mempalace_check_duplicate Similarity check before filing (0.9 threshold)
Palace — Write
Tool What
mempalace_add_drawer File verbatim content into a wing/room
mempalace_delete_drawer Remove a drawer by ID
mempalace_delete_wing Delete all drawers in a wing
mempalace_mine Trigger re-mining of a project directory (incremental or full)
Knowledge Graph
Tool What
mempalace_kg_query Entity relationships with time filtering
mempalace_kg_add Add a fact with optional validity window
mempalace_kg_invalidate Mark a fact as no longer true
mempalace_kg_timeline Chronological story of an entity
mempalace_kg_stats Graph overview
Architecture Retrieval
Tool What
mempalace_find_implementations Find all types implementing a given interface
mempalace_find_references Find all usages of a type (implementors, subclasses, deps)
mempalace_show_project_graph Project-level dependency graph, optionally filtered by solution
mempalace_show_type_dependencies Inheritance/implementation chain (ancestors + descendants)
mempalace_explain_subsystem Explain how a subsystem works: semantic search + KG expansion
mempalace_extract_reusable Classify deps as core/platform/glue; identify extraction boundary
mempalace_kg_query (entity="Service", direction="incoming") Show all services in the project
mempalace_kg_query (entity="Data", direction="incoming") Show all types in the data layer
Navigation & Diary
Tool What
mempalace_traverse Walk the graph from a room across wings
mempalace_find_tunnels Find rooms bridging two wings
mempalace_graph_stats Graph connectivity overview
mempalace_diary_write Write a session journal entry
mempalace_diary_read Read recent diary entries

MCP tools are discoverable by any MCP-capable client automatically. To teach the AI when and how to use them, paste the usage rules from docs/LLM_USAGE_RULES.md into your agent's instructions (CLAUDE.md, AGENTS.md, .cursorrules, etc.) — otherwise the tools are available but the assistant will not know the protocol. If you use a named profile, see the matching profile block in docs/LLM_USAGE_RULES.md for profile-scoped routing guidance.


Knowledge Graph

Temporal entity-relationship triples — local SQLite, no Neo4j, no cloud.

kg = KnowledgeGraph()
kg.add_triple("myapp", "uses", "Postgres", valid_from="2025-11-03")
kg.add_triple("myapp", "uses", "Redis",    valid_from="2026-01-15")

kg.query_entity("myapp")                    # → Postgres (current), Redis (current)
kg.query_entity("myapp", as_of="2025-12-01")  # → Postgres only

kg.invalidate("myapp", "uses", "Postgres", ended="2026-03-01")  # fact expired

Good candidates: version numbers, team assignments, tech stack choices, deployment states, deadlines.

Architecture extractionmempalace-code mine automatically emits higher-level KG facts for .NET and Python projects after each mine:

Predicate Example Query
is_pattern UserService → is_pattern → Service kg_query(entity="Service", direction="incoming")
is_layer UserRepository → is_layer → Data kg_query(entity="Data", direction="incoming")
in_namespace UserService → in_namespace → Company.App kg_query(entity="UserService")
in_project UserService → in_project → myapp kg_query(entity="myapp", direction="incoming")

Default patterns: Service, Repository, Controller, ViewModel, Factory. Default layers: UI (*.UI, *.Web, *.Presentation), Business (*.Application, *.Domain), Data (*.Data, *.Persistence), Infrastructure (*.Infrastructure).

Re-mining a project refreshes architecture facts for that project's wing only, so a multi-project palace can update one repo without expiring facts from another.

Override or extend via the architecture: block in mempalace.yaml:

architecture:
  enabled: true
  patterns:
    - name: Service
      suffixes: [Service]
      type_names: [AuditHandler]   # explicit names bypass suffix matching
  layers:
    - name: Business
      namespace_globs: ["*.Application", "*.Domain", "*.Audit"]
      type_suffixes: [Service]
      priority: 1

Set enabled: false to disable the pass entirely.


Memory Layers

Layer What When
L0 Identity — project, persona Always loaded (~50 tokens)
L1 Critical facts — team, decisions Always loaded (~120 tokens)
L2 Room recall — current topic On demand
L3 Deep search — full semantic query On demand
mempalace-code wake-up --wing myapp    # emit L0 + L1 context (~170 tokens)

For local models (Llama, Mistral) that don't speak MCP, pipe wake-up into the system prompt.


Backup & Restore

mempalace-code backup create                           # create backup archive (default: <palace_parent>/backups/)
mempalace-code backup create --out ~/safe/my.tar.gz   # custom path
mempalace-code backup create --kind scheduled          # create with 'scheduled' kind prefix
mempalace-code backup                                  # back-compat: same as 'backup create'
mempalace-code backup --out ~/safe/my.tar.gz           # back-compat: same as 'backup create --out ...'
mempalace-code backup list                             # list existing backups (with stale/oversized flags)
mempalace-code backup list --dir ~/old_backups/        # include extra directory in discovery
mempalace-code restore palace_backup_2026-04-14.tar.gz # restore
mempalace-code restore backup.tar.gz --force           # overwrite existing

Backups are written to <palace_parent>/backups/ by default. For a palace at ~/.mempalace/palace, that is ~/.mempalace/backups/.

Backup kinds: Each archive has a kind that controls its filename prefix and per-kind retention:

Kind Prefix Created by
manual mempalace_backup_ backup create (default)
scheduled scheduled_ backup create --kind scheduled / cron
pre_optimize pre_optimize_ Auto-backup before optimize

Scheduled backups:

# Print a scheduler snippet (does NOT install — owner action required)
mempalace-code backup schedule --freq daily    # daily at 03:00
mempalace-code backup schedule --freq weekly   # weekly on Sunday at 03:00
mempalace-code backup schedule --freq hourly   # every hour

# macOS: save and load the launchd plist
mempalace-code backup schedule --freq daily > ~/Library/LaunchAgents/com.mempalace.backup.plist
launchctl load ~/Library/LaunchAgents/com.mempalace.backup.plist

# Linux: paste the printed cron line into crontab -e
mempalace-code backup schedule --freq daily
# → 0 3 * * * /usr/local/bin/mempalace-code backup create --kind scheduled --palace /path/to/palace

Retention (automatic pruning):

pre_optimize archives are bounded by default to the newest 5 (implicit safe default for watcher daemons). scheduled archives are bounded by default to the newest 14 (implicit safe default for cron and launchd jobs). manual archives are unbounded by default (retain all, no pruning).

export MEMPALACE_BACKUP_RETAIN_COUNT=5   # explicit limit for all kinds; overrides implicit pre_optimize and scheduled bounds
# Deliberate keep-all opt-out (including pre_optimize and scheduled):
export MEMPALACE_BACKUP_RETAIN_COUNT=0   # 0 disables pruning for every kind

Or in ~/.mempalace/config.json: {"backup_retain_count": 5}. Retention only affects archives in the managed backups/ directory; explicit --out archives are never pruned.

backup list shows [stale] for archives that would be pruned at the current retain count, and [oversized] for archives larger than MEMPALACE_BACKUP_WARN_SIZE_BYTES.

After a successful optimize and readability check, MemPalace also runs best-effort verified Lance cleanup (cleanup_stale_fragments with unsafe_now=false) so future backups do not keep archiving stale table versions. Optimize and cleanup verification re-opens the Lance table, so it checks the same fresh-handle path the next CLI, MCP server, or watcher process will use. Use the manual cleanup command for older installations that already accumulated stale versions or for emergency recovery.

Disk-budget guard (1 GiB default):

export MEMPALACE_BACKUP_DISK_MIN_FREE_BYTES=2GiB    # require 2 GiB projected free after backup
# Legacy alias still accepted:
export MEMPALACE_BACKUP_MIN_FREE_BYTES=2GiB

When projected free space after the archive would fall below the configured floor, backup raises a disk budget error before writing the archive and optimize is skipped (fail-closed). The backup floor falls back to backup_disk_min_free_bytesdisk_min_free_bytes1 GiB default. Set the backup-specific floor to 0 to disable the backup guard.

Auto-backup before optimize (on by default):

backup_before_optimize is true by default. A backup is created under <palace_parent>/backups/pre_optimize_*.tar.gz before every optimize() call (runs after mining).

To opt out, add to ~/.mempalace/config.json:

{
  "auto_backup_before_optimize": false
}

Or set env var: MEMPALACE_AUTO_BACKUP_BEFORE_OPTIMIZE=0 (preferred) or MEMPALACE_BACKUP_BEFORE_OPTIMIZE=0.

Disable auto-optimize (paranoid mode):

{
  "optimize_after_mine": false
}

Skips compaction entirely. Storage will grow with more fragments but avoids any compaction-related corruption risk.

Why backup matters: Manual drawer additions (via mempalace_add_drawer) are not recoverable from source code. If LanceDB storage gets corrupted, only backups preserve this data. Code-mined drawers can be restored by re-running mempalace-code mine.

Also available: mempalace-code export --only-manual for JSONL export of manually-stored drawers.

Remote mirror risk — backups vs file mirroring:

Managed backups and Lance cleanup protect local palace state. They do not protect against delete-mode rsync between independent hosts. rsync --delete syncing a whole MemPalace state directory removes remote-owned drawers, diary entries, and KG triples that were never synced back to the source — even when local backups are healthy.

Use these recommended excludes for any delete-mode state-directory mirror:

rsync -a --delete \
  --exclude=palace/ \
  --exclude=knowledge_graph.sqlite3 \
  --exclude=config.json \
  --exclude=backups/ \
  ~/.mempalace/ user@host:.mempalace/

Run a preflight check before installing a mirror job (the command is never executed):

mempalace-code preflight mirror --command \
  "rsync -a --delete --exclude=palace/ --exclude=knowledge_graph.sqlite3 \
   --exclude=config.json --exclude=backups/ ~/.mempalace/ user@host:.mempalace/"
# OK

See docs/BACKUP_RESTORE.md for the full mirror-risk guidance and the safer export/import alternative for cross-host transfer of non-regenerable content.


Scan Excludes

By default mempalace-code mine already skips common generated directories (node_modules, __pycache__, .git, etc.). For project-specific noise — generated LSP state, build artifacts, IDE files — configure app-level excludes in ~/.mempalace/config.json:

{
  "scan_skip_dirs":  [".kotlin-lsp"],
  "scan_skip_files": ["workspace.json"],
  "scan_skip_globs": ["generated/**/*.js", "build/**"]
}
Key Match rule Default
scan_skip_dirs directory basename — prunes the whole subtree [".kotlin-lsp"]
scan_skip_files file basename — skips matching files anywhere []
scan_skip_globs project-relative POSIX glob — skips matching file paths []

workspace.json as opt-in example: a root workspace.json can be a legitimate monorepo config file, so it is not excluded by default. Add it to scan_skip_files only if your LSP generates it as noise inside generated directories.

These rules apply to both mempalace-code mine and the auto-watcher (mempalace-code mine --watch and mempalace-code watch). Force-include paths (--include-ignored) always win over app-level excludes.

Watcher loops reload these app-level rules between scan cycles, so edits to ~/.mempalace/config.json apply to subsequent re-mines without restarting mempalace-code watch.

Removing previously indexed noise: scan excludes prevent future scans from indexing the excluded paths. To remove content that was indexed before adding the exclusion, run a full re-mine:

mempalace-code mine <dir> --full

--full forces a clean rebuild and sweeps drawers from files that are no longer discovered by the scanner — including previously indexed files that now fall under an exclusion rule.


Health & Repair

mempalace-code health              # probe palace for fragment corruption
mempalace-code health --json       # machine-readable report
mempalace-code cleanup --older-than-days 7  # reclaim stale Lance versions
mempalace-code cleanup --unsafe-now         # emergency only; stop MemPalace processes first

mempalace-code repair --dry-run    # show what would be recovered
mempalace-code repair --rollback   # roll back to last working version

What health checks:

  1. Manifest read (count_rows)
  2. Data fragment read (head)
  3. Metadata scan (count_by_pair) - catches the silent-failure surface

What repair --rollback does:

  1. Walks LanceDB version history from newest to oldest
  2. Finds the most recent version where all probes pass
  3. Restores to that version (loses data added after corruption)

Use --dry-run first to see how many rows would be lost.

Normal optimize runs already prune verified stale Lance versions after a successful compaction and fresh-handle verification. Manual cleanup is still useful after older installs, large historical accumulations, or emergency disk recovery after stopping watchers, miners, maintenance commands, and MCP servers.


Version Check (opt-in)

mempalace-code can notify you when a newer release is available on PyPI. This feature is strictly opt-in — no network calls are made by default.

# Check your current status (local-only, no network)
mempalace-code version-check

# Opt in to periodic checks (contacts PyPI for package metadata only)
mempalace-code version-check --enable

# Opt out (suppresses future first-run prompts)
mempalace-code version-check --disable

# Check right now regardless of the interval setting
mempalace-code version-check --check-now

How it works:

  • On the first interactive command after a fresh install, the CLI prompts once: "Enable periodic new-version checks?" — answering n records the opt-out permanently. Non-interactive (piped, CI, non-TTY) invocations never prompt.
  • When opted in, a background check runs at most once per interval (default: 168 hours / 1 week). Any update hint appears on stderr only — stdout remains machine-parseable.
  • Explicit --check-now ignores the interval, contacts PyPI, and prints current/latest/error to stdout.
  • Only https://pypi.org/pypi/mempalace-code/json is contacted. No telemetry, no user IDs, no installed-package inventory.

Environment overrides:

Variable Effect
MEMPALACE_VERSION_CHECK=1 Force-enable (overrides config and state)
MEMPALACE_VERSION_CHECK=0 Force-disable (overrides config and state)
MEMPALACE_VERSION_CHECK_INTERVAL_HOURS=N Override interval (default: 168)

Setting MEMPALACE_VERSION_CHECK=0 in a CI pipeline guarantees no network calls regardless of any saved preference.


This Fork vs Upstream

This is a code-first fork of milla-jovovich/mempalace. We inherited the good parts — the palace metaphor, the MCP integration, the LongMemEval harness — and rebuilt what was broken. Every claim here is backed by code, tests, and documented benchmarks.

Upstream This fork
ChromaDB — silently deletes data on version bump LanceDB — crash-safe Arrow storage, no version-cliff
"No internet after install" — false mempalace-code init downloads the model explicitly once; cached mining/search use local-only model resolution first
"100% R@5" — unverifiable Number removed. Methodology caveats documented
~30% test coverage 2200+ tests, every feature acceptance-gated
No backup, no recovery backup / restore / export / import
No incremental mining Content-hash incremental: only changed files re-chunked
No code-search code_search — filter by language, symbol, glob
Line-count chunking Language-aware mining: tree-sitter AST for supported grammars, regex structural chunking, YAML-aware Kubernetes/Helm/Ansible splits, prose sections, and adaptive chunks for configs/data

Full audit: docs/UPSTREAM_HARDENING.md.


Benchmarks

Token savings vs grep + read (full methodology)

Project size Median Mean P95 Peak
Small (555 chunks) 13x 19x 42x 59x
Large (19k chunks) 80x 129x 279x 595x

Token savings scale with project size — grep noise grows linearly (more files contain the keyword), while mempalace-code search stays constant (top-5 semantically relevant chunks regardless of corpus size). These numbers are from a 19k-chunk project; larger codebases would push the ratios higher.

Retrieval quality

Benchmark Score
Code retrieval R@5 (MiniLM, 469 chunks) 95.0%
Code retrieval R@10 100%
.NET retrieval R@5 (CleanArchitecture pinned corpus, vector) 90.0%
.NET retrieval R@5 (same corpus, rerank="hybrid") 100%

Upstream LongMemEval result (96.6% R@5 on conversations) retained with methodology caveats.


Installation Details
pip install mempalace-code
# or
uv pip install mempalace-code

Bootstrap script (recommended for servers/CI):

curl -fsSL https://raw.githubusercontent.com/rergards/mempalace-code/main/scripts/bootstrap.sh | bash

Optional extras:

pip install "mempalace-code[treesitter]"  # AST parsing
pip install "mempalace-code[chroma]"      # ChromaDB legacy backend (deprecated, capped below 1.x)
pip install "mempalace-code[spellcheck]"  # autocorrect for room/wing names
pip install "mempalace-code[dev]"         # pytest + ruff + pyright

Requirements: Python 3.11+. ~80 MB embedding model cached once during mempalace-code init or mempalace-code fetch-model; repeated cached runs resolve the model locally.

All CLI Commands
# Setup
mempalace-code init <dir>                              # initialize rooms
mempalace-code init <dir> --detect-entities            # optional prose entity bootstrap

# Mining
mempalace-code mine <dir>                              # mine code project
mempalace-code mine <dir> --wing myapp                 # tag with wing
mempalace-code mine <dir> --mode convos                # mine conversations
mempalace-code mine <dir> --full                       # force full rebuild
mempalace-code mine <dir> --watch                      # auto-incremental on file changes
mempalace-code mine-all <parent-dir>                   # sync all projects incrementally (one wing per project)
mempalace-code mine-all <parent-dir> --new-only        # only mine projects not yet in the palace

# Watch (project auto-sync)
mempalace-code watch <initialized-project>             # watch a single initialized project
mempalace-code watch <parent-dir>                      # watch all initialized projects in a parent directory
mempalace-code watch <parent-dir> schedule             # print launchd/cron daemon snippet
mempalace-code watch <parent-dir> status               # disk-budget + launchd state

# Search
mempalace-code search "query"                          # search everything
mempalace-code search "query" --wing myapp             # scoped to wing
mempalace-code search "query" --room auth              # scoped to room

# Backup & Recovery
mempalace-code backup create                           # create backup (default: <palace_parent>/backups/)
mempalace-code backup list                             # list existing backups
mempalace-code backup schedule --freq daily            # print daily scheduler snippet
mempalace-code restore <archive>                       # restore from backup
mempalace-code export --only-manual                    # JSONL export
mempalace-code import <file>                           # JSONL import
mempalace-code health                                  # probe for fragment corruption
mempalace-code cleanup                                 # reclaim stale Lance versions
mempalace-code repair --rollback                       # roll back to last working version

# Context
mempalace-code wake-up                                 # L0 + L1 context
mempalace-code wake-up --wing myapp                    # project-scoped
mempalace-code status                                  # palace overview

# Model
mempalace-code fetch-model                             # cache or verify model for offline use
Saving Conversation Context

Code mining is automatic via mempalace-code watch. For conversation context (decisions, discussions, debugging notes), the AI uses MCP tools directly — works with any agent (Claude Code, Codex, Cursor, etc.):

  1. Wire the MCP server (see install docs)
  2. Add usage rules to your agent's instructions (CLAUDE.md, system prompt, etc.)
  3. The agent calls mempalace_add_drawer and mempalace_diary_write during sessions

Legacy: Claude Code also supports optional auto-save hooks that remind the AI to save at fixed intervals. These are redundant if MCP + usage rules are set up.

Project Structure
mempalace/
├── mempalace_code/
│   ├── cli.py              ← CLI entry point
│   ├── mcp_server.py       ← MCP server (29 tools, profiled at startup)
│   ├── storage.py          ← LanceDB vector storage
│   ├── miner.py            ← language-aware code chunking
│   ├── convo_miner.py      ← conversation ingest
│   ├── searcher.py         ← semantic search
│   ├── knowledge_graph.py  ← temporal entity graph (SQLite)
│   ├── palace_graph.py     ← room navigation graph
│   └── layers.py           ← 4-layer memory stack
├── mempalace/              ← source-only MCP compatibility shim
├── benchmarks/             ← reproducible benchmark runners
├── hooks/                  ← Claude Code auto-save hooks (legacy, optional)
├── examples/               ← usage examples
└── tests/                  ← 2200+ tests

Contributing

PRs welcome. See CONTRIBUTING.md.

python -m pytest tests/ -x -q    # full suite, all local, no network
python -m pyright --pythonpath "$(python -c 'import sys; print(sys.executable)')"  # type-check baseline

License

Apache 2.0 — see LICENSE and NOTICE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mempalace_code-1.11.0.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mempalace_code-1.11.0-py3-none-any.whl (268.6 kB view details)

Uploaded Python 3

File details

Details for the file mempalace_code-1.11.0.tar.gz.

File metadata

  • Download URL: mempalace_code-1.11.0.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mempalace_code-1.11.0.tar.gz
Algorithm Hash digest
SHA256 d2e59a856c5be44fae8b493e2b4c5681a31b014bc315f8150d38b79bdbd1a263
MD5 4f7c6d4c9f11798c28376ea33fcf1f2e
BLAKE2b-256 58e868b6d0e4a3ec3d94507991b7db46494fbc825473bd1f70766d08239fdeec

See more details on using hashes here.

Provenance

The following attestation bundles were made for mempalace_code-1.11.0.tar.gz:

Publisher: publish.yml on rergards/mempalace-code

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mempalace_code-1.11.0-py3-none-any.whl.

File metadata

  • Download URL: mempalace_code-1.11.0-py3-none-any.whl
  • Upload date:
  • Size: 268.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mempalace_code-1.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 1054cff02305b6f7ac9bab4d21fa34c349cfcde3a8e501d878a1b248df327f1f
MD5 0486e2aa6b100c31db6fd04d041c6077
BLAKE2b-256 93e922849d061ba234b70b1e9bc6bbeefad3515d7e695570501b1534121fdaac

See more details on using hashes here.

Provenance

The following attestation bundles were made for mempalace_code-1.11.0-py3-none-any.whl:

Publisher: publish.yml on rergards/mempalace-code

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page