Memory Path Engine
Local memory with replayable evidence paths — not only top-k chunks. Structured graph retrieval for agents, with a local palace CLI, MCP closed loop, and public-benchmark KPIs.
Memory Path Engine models memory as typed nodes, edges, weights, and MemoryPath objects so a system can retrieve, traverse, and explain how it reached an answer. Product M1 adds an on-disk palace (SQLite) and the mpe CLI; Layer B fixtures remain the architecture proof surface.
System shape (v0 + Memory Palace v1 + Product M1)
Bundled markdown packs are ingested into a graph (MemoryNode / MemoryEdge). Retrievers return a MemoryPath. The CLI persists that graph under .mpe/ and prints answer + hops.
Memory Palace v1 adds a parallel domain (memory_engine.memory) mapped through palace_to_store. See docs/architecture.md and docs/ROADMAP.md.
examples/*_pack ──▶ mpe ingest ──▶ .mpe/store.sqlite (typed graph)
│
┌──────────────┼──────────────┐
▼ ▼ ▼
BaselineTopK other modes WeightedGraph
(flat answers) in `retrieve` (path + scores)
│
▼
mpe search/path → answer + hop list
Why this project is different
Most RAG systems still look like this:
- Split documents into chunks.
- Embed chunks.
- Return
top-kmatches. - Ask the LLM to improvise the reasoning.
This project asks:
Can we retrieve a memory path instead of only retrieving similar chunks?
Three product bets:
structure: typed nodes and edges, not only a flat vector indexweight: importance, risk, novelty, reinforcement / forgettingpath: replayable evidence chains as the default search output
What you can do here
- compare multiple retrieval modes in one codebase
- inspect replayable evidence paths instead of only final answers
- test graph-aware retrieval on contract-like and operational documents
- run repository-owned structured benchmarks instead of toy snippets
Quick start
Maintainers: configure the GitHub link-card image using docs/social-preview.md (docs/assets/open-graph-cover.png).
Install from PyPI (Product M8) or editable clone (see docs/install.md for pipx / uv / Docker):
pip install memory-path-engine
# pip install 'memory-path-engine[embed]' # optional dense backends
# or: python -m pip install --no-build-isolation -e .
mpe doctor
Publish / first tag: docs/publish.md · Changelog: CHANGELOG.md.
Product CLI (mpe) — local palace
mpe init --mode hybrid
mpe ingest examples/runbook_pack/runbooks --pack example_runbook_pack
mpe search "What if rollback does not recover the API?" --mode hybrid
mpe path "What if rollback does not recover the API?"
mpe status
mpe backup
Palace files live in ./.mpe/ (or $MPE_PALACE). Search always prints an answer plus hop citations.
5-minute agent closed loop (MCP + hooks): see docs/getting-started.md.
mpe hooks install
mpe mcp # stdio MCP server for Cursor / Claude
# docker build -t mpe-mcp . && docker run -i --rm -v "$PWD/.mpe:/data/palace" -e MPE_PALACE=/data/palace mpe-mcp
Acceptance checklist: docs/ACCEPTANCE.md.
Current progress vs vision stages: docs/progress.md.
LongMemEval product KPI baseline
mpe bench longmemeval --label tiny --granularity session
mpe bench longmemeval --label tiny --granularity turn
Checked-in tiny reports live under benchmarks/external/longmemeval/baselines/.
Latest public recall (session, LongMemEval-S cleaned):
| Slice | Embedding | Mode | R@5 | R@10 | NDCG@10 |
|---|---|---|---|---|---|
| 50q | ngram | hybrid | 0.980 | 1.000 | 0.948 |
| 50q | ngram | lexical_baseline | 0.980 | 1.000 | 0.948 |
| 500q (full) | ngram | hybrid | 0.950 | 0.974 | 0.863 |
| 500q (full) | ngram | lexical_baseline | 0.950 | 0.974 | 0.863 |
| 500q (full) | fastembed | hybrid | 0.962 | 0.980 | 0.875 |
Artifacts: benchmarks/external/longmemeval/baselines/longmemeval_kpi_{medium30,medium50,medium50_fastembed,full_ngram,full_fastembed}.*.
M7 note: hybrid public ranking now reuses BM25/blend seed scores so NDCG no longer collapses after graph expansion. With fastembed, hybrid beats lexical on full R@5/R@10/NDCG.
HotpotQA mid-slice (64q, evidence-hit, Product M10):
| Embedding | Mode | evidence_hit | comparison hit |
|---|---|---|---|
| ngram | lexical_baseline | 0.406 | 0.562 |
| ngram | hybrid | 0.844 | 0.938 |
| fastembed | embedding_baseline | 0.531 | 0.812 |
| fastembed | hybrid | 0.859 | 0.938 |
Artifacts: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md. Dense (fastembed) especially helps comparison/paraphrase-like questions; set MPE_EMBEDDING_CACHE_DIR to cache vectors on disk.
Optional dense embeddings (Product M5+): default remains dependency-free ngram. Install pip install 'memory-path-engine[embed]' then:
mpe bench longmemeval --label tiny --embedding fastembed
# or: export MPE_EMBEDDING=fastembed
# optional: export MPE_EMBEDDING_MAX_CHARS=4000
Dense backends batch-encode when texts are short, and head+tail-truncate long session memories to fit model context.
Reproduce full LongMemEval-S (public KPI):
python scripts/download_longmemeval.py
mpe bench longmemeval \
--dataset benchmarks/external/longmemeval/data/longmemeval_s_cleaned.json \
--label full --granularity session --limit 0
| Slice | Granularity | How to run | Notes |
|---|---|---|---|
| tiny (2q, checked in) | session / turn | mpe bench longmemeval --label tiny [--granularity turn] |
CI smoke + committed baselines |
| medium (50q) | session | nightly / local with --limit 50 |
positioning |
| full (LongMemEval-S) | session / turn | download + --label full --limit 0 |
headline public KPI |
See benchmarks/external/longmemeval/README.md and docs/ROADMAP.md.
Run the test suite:
python -m unittest discover -s tests -v
Run the runbook demo:
python -m memory_engine.demo --scenario runbook
Terminal-style capture of real stdout (refresh with python scripts/generate_runbook_demo_terminal_svg.py; latency_ms may differ run to run):
Run the research-notes demo:
python -m memory_engine.demo --scenario research
Run the HotpotQA tiny benchmark sanity check:
python scripts/run_hotpotqa_benchmark.py
Download HotpotQA (CMU URL, with HuggingFace Hub fallback) and run the medium64 KPI slice:
python scripts/download_hotpotqa.py
python scripts/run_hotpotqa_benchmark.py \
--dataset benchmarks/external/hotpotqa/data/hotpot_dev_distractor_v1.json \
--limit 64 --top-k 10 \
--modes lexical_baseline,embedding_baseline,weighted_graph,hybrid,activation_spreading_v1 \
--summary-output benchmarks/external/hotpotqa/data/hotpotqa-medium64-summary.json
Optional dense embedding disk cache (Product M10):
export MPE_EMBEDDING_CACHE_DIR="$PWD/.mpe/embed-cache"
export MPE_EMBEDDING=fastembed
Committed mid-slice KPI: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md.
Run the LongMemEval tiny benchmark sanity check:
python scripts/run_longmemeval_benchmark.py
Print compact v1 palace metadata (spaces, routes, memory kinds) per case:
python scripts/run_longmemeval_benchmark.py --v1-recall-summary
Generate a fixed-format Layer B report (path/route/space/lifecycle/activation snapshot):
python scripts/generate_layer_b_report.py --output "benchmarks/structured_memory/layer_b_report.json" --markdown-output "benchmarks/structured_memory/layer_b_report.md"
Generate a fixed-format ablation matrix and latency summary (no-structure / no-weight / no-path-expansion):
python scripts/generate_ablation_report.py --output "benchmarks/structured_memory/ablation_report.json" --markdown-output "benchmarks/structured_memory/ablation_report.md"
Generate a Layer A positioning report (external metrics only; keeps architecture claims in Layer B):
python scripts/generate_layer_a_report.py --slice-profile tiny --markdown-output "benchmarks/external/layer_a_report.md"
Run Layer C transferable stand-in benchmarks:
python scripts/run_layer_c_benchmark.py --markdown-output "benchmarks/layer_c_minimal/layer_c_report.md"
Download the official HotpotQA dev distractor file for local benchmark runs:
python scripts/download_hotpotqa.py
Download the cleaned LongMemEval-S file for local benchmark runs:
python scripts/download_longmemeval.py
What you will see
python -m memory_engine.demo prints a small banner, the query, then path-aware output: a BEST ANSWER line built from the winning walk, and a REPLAY PATH with one line per hop (node id, score, via=<edge type>) plus short scoring reasons on the following lines. With --scenario contract, a BASELINE block (flat top-k answers) appears above the path-aware section for the same query.
Representative runbook excerpt (answer line shortened; latency and hop scores can vary slightly between runs):
========================================================================
Memory Path Engine | demo
scenario: runbook
========================================================================
-------------------------------- QUERY ---------------------------------
What should we do if rollback does not recover the API after a
deployment incident?
----------------- PATH-AWARE weighted graph retrieval -----------------
BEST ANSWER
… stitched runbook units … [latency_ms=…]
REPLAY PATH
1. 01_api_incident_runbook:5 | score=0.500 | via=seed
seed hit semantic=0.501
2. 01_api_incident_runbook:4 | score=0.299 | via=next_unit
expanded at hop 1 total=0.299 exception=0.450 contradiction=0.000
========================================================================
What the demos show
Runbook demo
The runbook demo loads incident and recovery procedures, then asks a multi-step operational question:
What should we do if rollback does not recover the API after a deployment incident?
The output includes:
- a BEST ANSWER line composed from the graph walk
- a REPLAY PATH with per-step scores,
viaedge types, and short reasons
For a representative stdout excerpt, see What you will see (under Quick start).
Contract demo
The contract demo runs the same query through a baseline retriever and the weighted graph retriever. Stdout shows flat top-k answers first, then the path-aware best answer and replay steps, so you can compare shapes of evidence without relying on a single aggregate metric.
Retrieval modes in this repo
| Retriever | What it emphasizes | Useful for |
|---|---|---|
| lexical baseline | keyword overlap | simple lookups and sanity checks |
| embedding baseline | semantic similarity | paraphrases and fuzzy matches |
| structure-only traversal | graph connectivity | linked evidence exploration |
| weighted graph retrieval | structure plus importance weighting | multi-hop retrieval with replayable evidence |
| activation spreading v1 | explicit propagation with decay | graph diffusion experiments |
Why the examples span multiple document types
The core is meant to stay domain-agnostic. The current examples use both contract-like documents and runbooks because together they stress:
- hierarchical structure
- exception and dependency chains
- critical risk-bearing units
- procedural and operational steps
- strong need for evidence-backed reasoning
If the retrieval and replay ideas cannot survive across these document types, they are unlikely to generalize well to other structured knowledge domains.
Repository layout
src/memory_engine: schema, storage, ingestion, retrieval, scoring, replayexamples/contract_pack: contract-like demo pack with dense dependencies and exceptionsexamples/runbook_pack: operational runbook pack for procedural retrievalbenchmarks/structured_memory: typed benchmark fixtures and evaluation assetsdocs: architecture, evaluation, hypotheses, and project visiontests: unit tests for schema, retrieval behavior, and benchmark support
Read this first
docs/vision.md: why this project exists and where it is headingdocs/architecture.md: how the current system is structureddocs/api-tracks.md: legacy vs palace recall entry pointsdocs/evaluation.md: how retrieval modes are compareddocs/benchmark-strategy.md: how public, repo-owned, and private benchmarks should be useddocs/private-contract-dataset-guide.md: how to build and annotate a private contract golden setdocs/hypotheses.md: milestone hypotheses and success criteria
Research hypotheses
The first milestone tests three claims:
H1: graph-aware retrieval beats vanillatop-kretrieval on multi-hop questionsH2: anomaly and importance weighting improve recall of critical evidenceH3: replayable memory paths improve explainability without unacceptable latency
Experimental framework
The retrieval stack separates:
- candidate generation
- semantic similarity backend
- scoring strategy
- path replay
That separation makes it possible to compare lexical baseline, embedding baseline, structure-only traversal, and weighted graph retrieval without rewriting the main search loop.
The evaluation layer can emit detailed per-question reports, which is useful for miss analysis and ablation debugging instead of relying only on a single aggregate score.
The repository also includes a dedicated structured benchmark bounded context with:
- strong pydantic dataset models
- a JSON repository for benchmark fixtures
- application services that load datasets, build stores, and run retrievers end to end
Benchmarks
The benchmark story is intentionally split into three layers:
- External positioning: LongMemEval retrieval-only recall (
R@5,R@10,NDCG@10) at session or turn granularity - Public retrieval sanity: HotpotQA evidence retrieval on distractor-style multi-document questions
- Mechanism validation: repository-owned structured fixtures for path, semantic, contradiction, and dynamic-memory behavior
Current run matrix:
benchmarks/structured_memory/*.json: CIbenchmarks/structured_memory/spatial_recall_benchmark.json,route_replay_benchmark.json,consolidation_gain_benchmark.json,state_transition_benchmark.json,contradiction_tension_benchmark.json: Layer B checks for palace-oriented expectations (space, route shape, diffusion gain, lifecycle) and explicit contradiction / rule-tension pairsbenchmarks/external/hotpotqa/hotpot_tiny_fixture.json: CI sanitybenchmarks/external/hotpotqa/data/*.json: local / nightly (mediumdefault 64 samples,fulloptional)benchmarks/external/longmemeval/longmemeval_tiny_fixture.json: local sanitybenchmarks/external/longmemeval/data/*.json: local / nightly (mediumdefault 50 samples,fulloptional)benchmarks/layer_c_minimal/*: runnable Layer C transfer stand-ins + private annotation templates
What is in scope for v0.4 (Product M3)
- typed
MemoryNode/MemoryEdge/MemoryPathgraph - SQLite-backed local palace (
.mpe/) andmpeCLI - stdio MCP server + Cursor hook templates (
mpe hooks install) hybridretriever mode (lexical + embedding blend → graph expand)- LongMemEval session + turn KPI baselines and full-corpus reproduce recipe
- domain packs for contract / runbook / research documents
What is out of scope for now
- Multi-backend vector zoo / hosted embedding services
- LLM-backed answer synthesis (path reasoning stays deterministic)
- multi-modal memory encoding
- full UI
Planned next steps
See docs/ROADMAP.md. Product M1–M10 are done (memory-path-engine is on PyPI); next levers:
- Organization-side Layer C private gold labels
- Optional HotpotQA full distractor KPI / LongMemEval turn-granularity tables
- Optional path explainability / agent UX polish
For suggested GitHub topic tags (About section), see docs/github-topics.md.
License
MIT. See LICENSE.
Metadata
Release files for memory-path-engine 0.9.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| memory_path_engine-0.9.0.tar.gz | 148.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| memory_path_engine-0.9.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 281.3 kB
Release files / memory_path_engine-0.9.0.tar.gz
| Download URL | memory_path_engine-0.9.0.tar.gz |
|---|---|
| Size | 148.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
004037271d27252d18da339eed1686917f85bea7300c5411ee71386005801055
|
|
BLAKE2b-256 checksum How to use checksums |
4c24d8b30a38c92221f78a8c9c9931206a67deef2cc41edb831f22ea62a7c8db
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / memory_path_engine-0.9.0-py3-none-any.whl
| Download URL | memory_path_engine-0.9.0-py3-none-any.whl |
|---|---|
| Size | 133.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
53755aecdd960f77142c0f0c145509bcfd44b80df15a2b8b21b75c65a2b865fc
|
|
BLAKE2b-256 checksum How to use checksums |
34b345884789eabbb9e8f6e7673ad07befd81ac8b281223b266504bc7f30efb5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log