Skip to main content
MemPalace

MemPalace

Local-first AI memory. Verbatim storage, pluggable backend, 96.6% R@5 raw on LongMemEval — zero API calls.


What it is

MemPalace stores your conversation history as verbatim text and retrieves it with semantic search. It does not summarize, extract, or paraphrase. The index is structured — people and projects become wings, topics become rooms, and original content lives in drawers — so searches can be scoped rather than run against a flat corpus.

The retrieval layer is pluggable. The current default is ChromaDB; the interface is defined in mempalace/backends/base.py and alternative backends can be dropped in without touching the rest of the system.

Nothing leaves your machine unless you opt in.

Architecture, concepts, and mining flows: mempalaceofficial.com/concepts/the-palace.


Install

MemPalace ships a CLI, so install it in an isolated environment to avoid PEP 668 errors on Debian/Ubuntu/Homebrew Pythons and to keep mempalace's deps (chromadb, numpy, grpcio, …) from conflicting with anything else in your global site-packages.

We recommend uv — uv tool install puts the mempalace CLI in an isolated environment on your PATH:

uv tool install mempalace
mempalace init ~/projects/myapp

pipx works the same way if you prefer it: pipx install mempalace.

Prefer plain pip only inside an activated virtualenv where you explicitly want import mempalace available:

python -m venv .venv && source .venv/bin/activate
pip install mempalace

Android / Termux

Native Termux installation is not currently supported because compiled dependencies such as ChromaDB and ONNX Runtime publish Linux wheels, not Android wheels. Android ARM64 users can run the regular Linux packages in an isolated Debian PRoot container instead. See the Termux installation guide for the tested setup and an argv-preserving launcher.

Docker

A container image is also available for running the MCP server or the CLI without a local Python toolchain. Multi-arch (amd64 + arm64), so it runs natively on Apple Silicon:

docker pull ghcr.io/mempalace/mempalace:latest

Everything persists under /data — palace, config, and the cached embedding model — so mount a volume there and reuse it across runs:

# MCP server over stdio — note the `-i` flag (JSON-RPC needs stdin)
docker run -i --rm -v mempalace-data:/data ghcr.io/mempalace/mempalace

# Run any CLI command instead. The container only sees what you mount, so
# mount the directory you want to mine — read-only is enough, mining never
# writes to the source.
docker run --rm -v mempalace-data:/data -v /path/to/project:/work:ro \
  ghcr.io/mempalace/mempalace mine /work
docker run --rm -v mempalace-data:/data ghcr.io/mempalace/mempalace search "why GraphQL"

The first command that needs embeddings downloads the model into /data (~80 MB for the default minilm, ~300 MB for embeddinggemma). It is a one-off as long as the volume persists, but it does mean the first call is slow and needs network — worth knowing before assuming a hung container.

Wire it into an MCP client (e.g. Claude Code) as a stdio server. Mount anything you want the server to be able to mine — it cannot reach your transcripts otherwise:

{
  "mcpServers": {
    "mempalace": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-v", "mempalace-data:/data",
        "-v", "/absolute/path/to/.claude/projects:/transcripts:ro",
        "ghcr.io/mempalace/mempalace"
      ]
    }
  }
}

Use a real absolute path there — ~ and $HOME are not expanded by every MCP client. Paths are container paths from then on: mine /transcripts, not ~/.claude/projects.

Mount permissions on Linux. The image runs as uid 1000 and bind mounts keep their host ownership, so a mounted directory has to be readable by that uid — an ordinary 0755 checkout is fine, a 0700 directory is not, and the failure surfaces as PermissionError: [Errno 13] rather than anything about Docker. Docker Desktop maps uids on macOS and Windows, so this only bites on Linux. Do not work around it with --user: /data is owned by uid 1000 inside the image, so another uid cannot write the palace at all.

docker compose run --rm mcp works too (see docker-compose.yml), and deploy/docker-compose.server.yml stands up the team server. To build the image yourself instead of pulling — required for the GPU variant, which is not published:

docker build -t mempalace .                                  # CPU
docker build --build-arg EXTRAS="extract,spellcheck" -t mempalace .
docker build -f Dockerfile.gpu -t mempalace:gpu .            # CUDA; run with --gpus all

The GPU image is x86_64-only: onnxruntime-gpu publishes no aarch64 Linux wheels, so that last build fails on an ARM host (including Apple Silicon) with a dependency-resolution error rather than an obvious one.

Note that a build from a clone uses whatever branch you checked out; develop is the default branch, so pull the published image if you want the released version.

Storage backends

ChromaDB is the default and needs no configuration. MemPalace also ships a pluggable backend contract, exercised across deliberately different substrates so the contract is never accidentally shaped around one vendor. Every non-default backend is opt-in.

Backend Mode Install Namespaces Lexical Configure with
chroma (default) Local (embedded) bundled – ✓ –
sqlite_exact Local (exact) bundled – ✓ –
milvus Local (Lite) · Server opt-in mempalace[milvus] ✓ ✓ MEMPALACE_MILVUS_URI
qdrant Server (REST) bundled ✓ ✓ MEMPALACE_QDRANT_URL
pgvector Server (Postgres) mempalace[pgvector] ✓ ✓ MEMPALACE_PGVECTOR_DSN

Select with --backend <name>, MEMPALACE_BACKEND=<name>, or "backend": "<name>" in config.json. See Storage backends for connection variables, namespace behavior, and deployment notes.

Quickstart

# Mine content into the palace
mempalace mine ~/projects/myapp                    # project files
mempalace mine ~/.claude/projects/ --mode convos   # Claude Code sessions (scope with --wing per project)

# Search
mempalace search "why did we switch to GraphQL"

# Load context for a new session
mempalace wake-up

For Claude Code, Gemini CLI, Antigravity, MCP-compatible tools, and local models, see mempalaceofficial.com/guide/getting-started.


Benchmarks

All numbers below are reproducible from this repository with the commands in benchmarks/BENCHMARKS.md. Full per-question result files are committed under benchmarks/results_*.

LongMemEval — retrieval recall (R@5, 500 questions):

Mode R@5 LLM required
Raw (semantic search, no heuristics, no LLM) 96.6% None
Hybrid v4, held-out 450q (tuned on 50 dev, not seen during training) 98.4% None
Hybrid v4 + LLM rerank (full 500) ≥99% Any capable model

The raw 96.6% requires no API key, no cloud, and no LLM at any stage. The hybrid pipeline adds keyword boosting, temporal-proximity boosting, and preference-pattern extraction; the held-out 98.4% is the honest generalisable figure.

The rerank pipeline promotes the best candidate out of the top-20 retrieved sessions using an LLM reader. It works with any reasonably capable model — we have reproduced it with Claude Haiku, Claude Sonnet, and minimax-m2.7 via Ollama Cloud (no Anthropic dependency). The gap between raw and reranked is model-agnostic; we do not headline a "100%" number because the last 0.6% was reached by inspecting specific wrong answers, which benchmarks/BENCHMARKS.md flags as teaching to the test.

Other benchmarks (full results in benchmarks/BENCHMARKS.md):

Benchmark Metric Score Notes
LoCoMo (session, top-10, no rerank) R@10 60.3% 1,986 questions
LoCoMo (hybrid v5, top-10, no rerank) R@10 88.9% Same set
ConvoMem (all categories, 250 items) Avg recall 92.9% 50 per category
MemBench (ACL 2025, 8,500 items) R@5 80.3% All categories

We deliberately do not include a side-by-side comparison against Mem0, Mastra, Hindsight, Supermemory, or Zep. Those projects publish different metrics on different splits, and placing retrieval recall next to end-to-end QA accuracy is not an honest comparison. See each project's own research page for their published numbers.

Reproducing every result:

git clone https://github.com/MemPalace/mempalace.git
cd mempalace
uv sync --extra dev   # or: pip install -e ".[dev]"
# see benchmarks/README.md for dataset download commands
uv run python benchmarks/longmemeval_bench.py /path/to/longmemeval_s_cleaned.json

Knowledge graph

MemPalace includes a temporal entity-relationship graph with validity windows — add, query, invalidate, timeline — backed by local SQLite. Usage and tool reference: mempalaceofficial.com/concepts/knowledge-graph.

MCP server

44 MCP tools cover palace reads/writes, knowledge-graph operations, cross-wing navigation, drawer management, agent diaries, and agent coordination (logstream events + artifact handoffs). Installation and the full tool list: mempalaceofficial.com/reference/mcp-tools.

Agents

Each specialist agent gets its own wing and diary in the palace. Discoverable at runtime via mempalace_list_agents — no bloat in your system prompt: mempalaceofficial.com/concepts/agents.

Auto-save hooks

Auto-save hooks for Claude Code, Codex CLI, and Cursor IDE save periodically and before context compression:

If you are installing under time pressure, start with the Claude Code retention setup checklist: wire the hooks, back up existing JSONL transcripts, and backfill them with mempalace mine ~/.claude/projects/ --mode convos.

For per-message recall on top of the file-level chunks the hooks produce, run mempalace sweep <transcript-dir> periodically — it stores one verbatim drawer per user/assistant message, idempotent and resume-safe.


Requirements

  • Python 3.9+
  • A vector-store backend (ChromaDB by default)
  • ~300 MB disk for the embedding model. Onboarding (python -m mempalace.onboarding) offers embeddinggemma-300m (multilingual, 100+ languages, recommended) or all-MiniLM-L6-v2 (English-only, ~30 MB). See the docstring at mempalace/embedding.py for details and migration notes.
  • Optional — compute embeddings on a server instead of locally. Set embedding_model: "openai-compat" in ~/.mempalace/config.json together with embedding_api_url / embedding_api_model (and embedding_api_key if the server needs auth) to use any OpenAI-compatible /v1/embeddings endpoint — LM Studio, llama.cpp, vLLM, Ollama's OpenAI shim, or a self-hosted server (e.g. a larger multilingual or GPU-served embedder). Each key is overridable via the matching MEMPALACE_EMBEDDING_API_* env var. When the endpoint is on your machine or LAN, no content leaves your network. Switching to it requires mempalace repair rebuild-index (different vector space).

No API key is required for the core benchmark path.

Docs

Contributing

PRs welcome. See CONTRIBUTING.md.

License

MIT — see LICENSE.

Metadata

Release files for mempalace 3.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mempalace 3.8.0
File Size Uploaded
mempalace-3.8.0.tar.gz 25.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for mempalace 3.8.0
File Interpreter ABI Platform
mempalace-3.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 26.5 MB

Release files / mempalace-3.8.0.tar.gz

Download URL mempalace-3.8.0.tar.gz
Size 25.8 MB
Tags Source
SHA-256 checksum
How to use checksums
9336425c2e8dabdc028617565e663c22af2179be33ef9f822404aa78221ed21e
BLAKE2b-256 checksum
How to use checksums
466bf771d3ba3ab8393cdea451932df96097585b5659c17853b046b992fa60d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release files / mempalace-3.8.0-py3-none-any.whl

Download URL mempalace-3.8.0-py3-none-any.whl
Size 755.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1c00680aeab8b9cf0f2c94bc79cbabce3fad2ea6586c3fc215d5a6b503e2d9dd
BLAKE2b-256 checksum
How to use checksums
82c7999a2fa64d5875eac50fb87ce2babf1ee033267df2cbe45e033e1afd078f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release history Release notifications | RSS feed

3.10.0

2 release files

3.9.0

2 release files

This release

3.8.0 This release

2 release files

3.7.1

2 release files

3.7.0

2 release files

3.6.0

2 release files

3.5.0

2 release files

3.4.1

2 release files

3.4.0

2 release files

3.3.6

2 release files

3.3.5

2 release files

3.3.4

2 release files

3.3.3

2 release files

3.3.2

2 release files

3.3.1

2 release files

3.3.0

2 release files

3.2.0

2 release files

3.1.0

2 release files

3.0.0

2 release files

2.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page