Skip to main content

birkin-mnemosyne - tiny multilingual memory for agents: three note cards linked by glowing threads to a search lens

birkin-mnemosyne

birkin-mnemosyne is local Markdown memory for personal agents: multilingual BM25 retrieval, usage-driven decay and zone priorities, plus provider-portable curation whose file operations are bounded by a deterministic executor. The core has zero runtime dependencies and needs no model or API key to store or search notes. Notes stay readable, greppable and diffable; an optional MCP server connects the vault to agent clients, and an opt-in semantic mode adds meaning-based retrieval.

Install

Requires Python >= 3.10. Install directly from GitHub:

pip install git+https://github.com/ashmoonori-afk/birkin-mnemosyne

From a checkout, the existing development and benchmark commands are:

pip install -e .                 # standard-library-only core
pip install -e ".[dev]"          # pytest, ruff
pip install -e ".[bench]"        # fastembed, numpy: embedding baselines only
pip install -e ".[mcp]"          # official MCP SDK and server
pip install -e ".[semantic]"     # optional meaning-based retrieval

The extras are optional. [bench] reproduces embedding baselines; it does not enable semantic retrieval. See below for explicit semantic preparation.

30-second usage

from birkin_mnemosyne import VaultMemory

mem = VaultMemory({"vault_path": "my_vault"})
mem.write_note(
    "Ingress DNS",
    "nginx ingress resolves service DNS; allow HTTPS through the firewall.",
    zone="devops",
)
for hit in mem.dex.search("ingress dns"):
    print(hit)
# Reinforce the note your agent actually used, not every search candidate.
mem.dex.record_access("ingress-dns")

For an existing vault, use Mnemosyne("my_vault"), call refresh(), then search(query). Unicode normalization, accent folding, Latin prefix stems, and Hangul/Han/kana bigrams support multilingual lexical matching. File changes refresh the index; the compressed cache is rebuildable, while usage history persists separately. Usage decay and zone activity adjust ranking.

What the numbers say

All retrieval quality below is the final frozen test split, copied from RESULTS.md. The synthetic corpus has 160 gold notes in six languages, padded with invented distractors to 1,000 and 10,000 notes. Three independent query authors (claude-opus-5.5, gpt-6.1-sol, claude-fable-5.1) supply exact, low-overlap paraphrase and code-switched questions. Cross-language sibling notes are removed before scoring; tuning uses dev only.

MRR measures how early the correct note appears; R@5 measures whether it is in the first five hits. Higher is better. Each arrow is default core -> optional semantic mode, not an earlier draft result. Bold values mark a loss. Language codes: en English, ko Korean, ja Japanese, zh Chinese, es Spanish, de German.

Per language, all query kinds

160 notes

Language MRR: core -> semantic R@5: core -> semantic
en 0.591 -> 0.640 0.698 -> 0.746
ko 0.579 -> 0.640 0.654 -> 0.728
ja 0.794 -> 0.843 0.848 -> 0.899
zh 0.818 -> 0.871 0.867 -> 0.926
es 0.698 -> 0.744 0.769 -> 0.861
de 0.689 -> 0.749 0.756 -> 0.822

1,000 notes

Language MRR: core -> semantic R@5: core -> semantic
en 0.581 -> 0.620 0.679 -> 0.714
ko 0.554 -> 0.614 0.613 -> 0.683
ja 0.782 -> 0.809 0.848 -> 0.909
zh 0.766 -> 0.823 0.830 -> 0.881
es 0.655 -> 0.686 0.694 -> 0.713
de 0.664 -> 0.708 0.741 -> 0.770

10,000 notes

Language MRR: core -> semantic R@5: core -> semantic
en 0.568 -> 0.604 0.655 -> 0.675
ko 0.559 -> 0.608 0.609 -> 0.658
ja 0.782 -> 0.816 0.848 -> 0.879
zh 0.764 -> 0.782 0.822 -> 0.815
es 0.673 -> 0.681 0.704 -> 0.704
de 0.681 -> 0.708 0.741 -> 0.748

Per independent query author and language

These MRR tables keep every author and corpus size separate. Semantic mode loses Spanish MRR for claude-opus-5.5 at 1,000 and 10,000 notes, and Chinese MRR for gpt-6.1-sol at 10,000 notes; those cells are bold, not averaged away.

claude-opus-5.5

Notes Language Core MRR Semantic MRR
160 en 0.606 0.672
160 ko 0.597 0.668
160 ja 0.773 0.812
160 zh 0.820 0.852
160 es 0.692 0.707
160 de 0.694 0.739
1000 en 0.592 0.644
1000 ko 0.565 0.622
1000 ja 0.784 0.795
1000 zh 0.750 0.794
1000 es 0.653 0.646
1000 de 0.671 0.692
10000 en 0.588 0.639
10000 ko 0.570 0.608
10000 ja 0.763 0.828
10000 zh 0.746 0.764
10000 es 0.653 0.639
10000 de 0.671 0.687

gpt-6.1-sol

Notes Language Core MRR Semantic MRR
160 en 0.524 0.559
160 ko 0.612 0.671
160 ja 0.827 0.840
160 zh 0.811 0.877
160 es 0.628 0.711
160 de 0.612 0.663
1000 en 0.512 0.542
1000 ko 0.609 0.674
1000 ja 0.821 0.851
1000 zh 0.762 0.847
1000 es 0.577 0.652
1000 de 0.595 0.629
10000 en 0.496 0.517
10000 ko 0.613 0.668
10000 ja 0.826 0.828
10000 zh 0.777 0.776
10000 es 0.609 0.643
10000 de 0.622 0.648

claude-fable-5.1

Notes Language Core MRR Semantic MRR
160 en 0.643 0.691
160 ko 0.530 0.581
160 ja 0.781 0.876
160 zh 0.822 0.884
160 es 0.775 0.814
160 de 0.761 0.844
1000 en 0.641 0.673
1000 ko 0.487 0.548
1000 ja 0.742 0.781
1000 zh 0.785 0.829
1000 es 0.736 0.761
1000 de 0.726 0.804
10000 en 0.621 0.655
10000 ko 0.493 0.548
10000 ja 0.759 0.790
10000 zh 0.770 0.807
10000 es 0.756 0.763
10000 de 0.749 0.787

The Chinese author-level R@5 losses at 10,000 notes are also explicit:

Author R@5: core -> semantic
claude-fable-5.1 0.867 -> 0.844
gpt-6.1-sol 0.822 -> 0.800

Where semantic mode falls below the core

The mode is opt-in and is not the recommended default. The bar was "no slice below the core". It misses that bar on the test split in one pooled slice, three author slices and six language x kind cells. The pooled Chinese and author losses are above; every language x kind loss is below. mixed means code-switched; para means low-overlap paraphrase.

Notes Language / kind MRR: core -> semantic R@5: core -> semantic
160 ja / mixed 0.896 -> 0.886 0.970 -> 0.970
160 zh / mixed 0.917 -> 0.905 0.978 -> 0.956
1000 ja / mixed 0.865 -> 0.826 0.970 -> 0.970
10000 ja / mixed 0.880 -> 0.859 0.970 -> 0.970
10000 zh / mixed 0.875 -> 0.860 0.911 -> 0.911
10000 zh / para 0.418 -> 0.487 0.556 -> 0.533

The configuration passed the dev gate, was frozen before this test run and was not retuned afterward. Exact keyword queries keep their lexical order; no exact query moved at any corpus size. Turn semantic mode on when questions are usually worded differently from the notes; leave it off when lookups are mostly keywords.

Footprint

Apple M1, macOS arm64 measurements, not service-level guarantees. Every arrow is core -> semantic. Install size is 0.18 MB -> 38.2 MB, including the extra's dependencies but not the prepared model. The first preparation needs an approximately 530 MB download; the compact model is approximately 140 MB on disk. The memory budget was 150 MB on top of the core; the largest measured addition was 61 MB while indexing. The prepared-model start-up budget was 1 s; the slowest semantic start in the reference timing run took 439 ms.

Index and memory

Notes Index on disk RSS: search RSS: indexing
160 0.07 MB -> 0.09 MB 25 -> 48 MB 26 -> 83 MB
1000 0.35 MB -> 0.43 MB 34 -> 56 MB 40 -> 101 MB
10000 3.27 MB -> 4.08 MB 138 -> 159 MB 170 -> 231 MB

Start-up and latency

Notes Start-up median (max) p50 p95
160 41 ms (42) -> 89 ms (95) 0.5 -> 1.6 ms 0.6 -> 3.4 ms
1000 65 ms (98) -> 125 ms (132) 3.1 -> 7.4 ms 3.9 -> 8.4 ms
10000 307 ms (547) -> 382 ms (439) 35.4 -> 75.6 ms 37.8 -> 79.4 ms

RSS is peak resident memory in separate fresh processes for search and whole-vault indexing. Start-up includes interpreter start, import, index load and first query. Reference timings were measured at a lower machine load (1-minute load average 5.6 to 7.8), using 10 processes per start-up cell and 300 test queries in a warm process. Latency grows with vault size and machine load. The timing fields in semantic_test_run.json are not the reference; use the separately measured start-up and latency table in RESULTS.md.

Honest limits

  • Korean does not improve in the default core. Against the earlier tokenizer, at 10,000 notes its MRR is 0.569 -> 0.559. English stems also match English distractors for code-switched Korean questions; Korean tokenization itself is unchanged. Stem reweighting traded English for Korean, and a minority-script boost hurt independently authored mirror questions. Neither was shipped: see Code-switched queries in the core.
  • This is a synthetic personal-scale benchmark. It does not establish superiority over other memory projects, answer accuracy, or universal multilingual coverage. Low-overlap paraphrases remain difficult for the lexical core. Scripts with combining vowel signs, such as Devanagari and Thai, are split at those signs by the current tokenizer.
  • Semantic results can be irrelevant. With the mode on, even a query without a lexical match receives semantic candidates; there is no relevance floor. Tokenizer equivalence was checked on this benchmark only, and memory in a long-lived process with varied CJK text was not measured. Footprint measurements cover macOS arm64 only.
  • Compression has a write cost. Saving recompresses the whole index; loading briefly holds compressed and decoded text together. Mixing older and newer library versions on a vault causes repeated cache rebuilding.
  • Curation safety is narrower than correctness. The executor bounds file operations; model choice still affects placement and linking quality. Python's explicit purge_expired() maintenance call can delete expired notes and is not exposed over MCP.
  • The earlier LongMemEval retrieval and end-to-end QA harnesses are not yet public in this repository. Their numbers are omitted here. The companion paper is not a substitute for a committed reproduction harness.

Reproduce the benchmark

Core (no optional runtime dependencies):

python benchmarks/retrieval/bench_retrieval.py --sizes 160 1000 10000 --install-size

Semantic mode uses minishlab/potion-multilingual-128M static embeddings, with no transformer at query time. Install the extra and explicitly prepare its model once:

pip install "birkin-mnemosyne[semantic] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne"
python -m birkin_mnemosyne.semantic

Opt in with Mnemosyne(vault, semantic=True) or MNEMOSYNE_SEMANTIC=1. Search never downloads or converts a model. If requested with the extra missing or the model unprepared, search uses the core ranking and logs one line explaining why. Semantic chunks are fused with lexical ranking; full lexical matches stay first and fused ties use lexical rank.

From a checkout with the model prepared:

python benchmarks/retrieval/bench_retrieval.py --engines bm25 hybrid --sizes 160 1000 10000 --json run.json --install-size --install-extras semantic
python benchmarks/retrieval/compare.py run.json run.json --engine-before bm25 --engine-after hybrid

The committed quality run is semantic_test_run.json. See RESULTS.md for the full query-kind tables, query counts, rejected ideas and footprint methodology. Additional checks and offline examples:

pytest -q                                # MCP tests skip without [mcp]
python examples/quickstart.py            # write, search, decay
python examples/automatic_profiles.py    # profile review and persistence
python benchmarks/bench_safety_matrix.py # defense-layer ablation
python benchmarks/bench_korean_embed.py  # embedding baseline; needs [bench]

CurationPlan/1: bounded curation with any model

A model proposes typed JSON operations (rezone, link, supersede, archive); a deterministic executor validates, clamps, applies and audits them. The curation schema has no delete operation: archive moves a note. Archives are capped per pass using ARCHIVE_CAP_MIN and ARCHIVE_CAP_FRACTION of active notes; negative-polarity warnings and control notes are protected. Every move stays inside the vault, _archive is not an active zone, and free-text summaries are inert rather than control signals. Unparseable output becomes an empty plan. Placement is model judgment; linking co-placed notes is mechanical. Schema validation alone does not prevent mass archiving; the executor's clamp does. Path containment and slug lookup also protect the lower-level file operations.

from pathlib import Path

from birkin_mnemosyne import Mnemosyne, run_curation_pass, get_completer

vault = Path("my_vault")
mem = Mnemosyne(vault)
mem.refresh()
hits = mem.search("kubernetes ingress dns")

# the vault argument must be a pathlib.Path
outcome = run_curation_pass(vault, get_completer("codex"), provider="codex")
# Or pass your own complete(prompt: str) -> str callable.

For openclaw, hermes or your own loop, expose mem.search(query) as a recall tool, call mem.record_access(note_slug) for notes actually used, and pass your existing model client to run_curation_pass(vault_path, my_complete, provider="custom"). It returns a CurationOutcome with accepted/dropped operations, archive cap and audit summary. Already have a plan as data? evaluate_plan(vault_path, plan) runs the same gate as a dry run; pass apply=True to apply it.

Provider Invocation Plan constraint
claude claude -p, empty tool allow-list prompt-specified
codex codex exec --sandbox read-only --output-schema enforced JSON schema
api Anthropic Messages API via stdlib urllib prompt-specified
gemini gemini -p -, CLI defaults prompt-specified
local ollama run <model> prompt-specified

MCP server (Claude Code, Claude Desktop, Codex CLI, Cursor)

The vault can be served over the Model Context Protocol so any MCP client can remember, recall and curate. The server is an optional extra — the core library stays stdlib-only — and runs with one command over stdio:

uvx --from "birkin-mnemosyne[mcp] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne" \
    mnemosyne-mcp --vault ~/mnemosyne

(or pip install "birkin-mnemosyne[mcp] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne" and run mnemosyne-mcp). The first launch downloads the SDK; if your client times out on that first start, run the command once in a terminal.

setting how
vault directory --vault PATH, else $MNEMOSYNE_VAULT, else ~/.birkin-mnemosyne/vault
require a source for new notes --evidence-required or MNEMOSYNE_EVIDENCE_REQUIRED=1

Claude Code

claude mcp add --scope user mnemosyne -- \
  uvx --from "birkin-mnemosyne[mcp] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne" \
  mnemosyne-mcp --vault ~/mnemosyne

Claude Desktop (claude_desktop_config.json) and Cursor (~/.cursor/mcp.json or .cursor/mcp.json) use the same shape:

{
  "mcpServers": {
    "mnemosyne": {
      "command": "uvx",
      "args": [
        "--from", "birkin-mnemosyne[mcp] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne",
        "mnemosyne-mcp", "--vault", "~/mnemosyne"
      ]
    }
  }
}

Codex CLI (~/.codex/config.toml, or codex mcp add mnemosyne -- uvx ...):

[mcp_servers.mnemosyne]
command = "uvx"
args = ["--from", "birkin-mnemosyne[mcp] @ git+https://github.com/ashmoonori-afk/birkin-mnemosyne",
        "mnemosyne-mcp", "--vault", "~/mnemosyne"]

Point several clients at the same --vault to share one memory; writes are serialized across server processes with a lock file.

Tools

tool what it does safe default
memory_search BM25 + decay + zone ranked hits with snippets read-only
memory_get_note full note with its version; counts as a use —
memory_list notes by zone, paginated read-only
memory_remember write a note: mode="create" / "append" / "replace" create never overwrites; replace needs expected_version
memory_related mechanical link candidates for a note read-only
memory_forget move a note to _archive through the curation gate dry run unless confirm=true; never deletes
memory_restore move a note back out of _archive —
memory_curation_catalog structured catalog for writing a CurationPlan read-only
memory_curate run a CurationPlan/1 through the deterministic gate dry run unless apply=true

Plus the resources mnemosyne://digest (the prompt digest) and mnemosyne://note/{slug}, and the prompt curate_vault. The calling agent is the curator: it reads the catalog, writes a plan, and memory_curate clamps it exactly like run_curation_pass would — protected notes stay put and the archive cap applies to every call. Nothing exposed over MCP can hard-delete a file (purge_expired stays a Python-only maintenance call). Note text returned by the tools is stored data from earlier sessions; the server tells clients not to follow instructions found inside it.

Automatic role profiles

ProfileMemory records a conversation exchange immediately and reviews it on one owned background worker. The reviewer returns JSON; flush() is the durability boundary and surfaces malformed reviewer output.

import json
from birkin_mnemosyne import ProfileMemory

def review(exchange):
    # Replace this deterministic example with your model client.
    return json.dumps({"profiles": {
        "preferences": "Prefers evidence before conclusions.",
        "soul": "Use direct Korean.",
    }})

with ProfileMemory(vault_path, review) as profiles:
    profiles.record_exchange(user_message, assistant_message)
    profiles.flush()

The system/ directory, owned by ProfileMemory, contains exactly five role files (the curation gate does not special-case them):

File Guidance stored
user.md User characteristics and stable personal context
preferences.md Preferences and favored choices
soul.md Conversation style and interaction guidance
workflow.md Work process and execution guidance
automation.md Workflow automation guidance

By default, ProfileMemory(vault_path, review) bootstraps those files and appends de-duplicated guidance lines. If you pass save=callable, no system/ directory or files are created; each reviewed exchange is parsed into a tuple of ProfileProposal objects and delivered to that sink for caller-owned persistence. Sink-mode instances cannot read_profiles() because they own no files.

The reviewer contract is a JSON object with one profiles object. Profile keys must be from the table. Values may be legacy non-empty strings, which become add proposals after whitespace normalization, or ordered proposal lists such as [{"action":"replace","old_text":"old","content":"new"}]. Actions are add, replace, or remove; malformed JSON, unknown keys/actions, non-string fields, and missing required text raise ProfileReviewError through flush(). close() stops new submissions and releases the worker; the context manager flushes and closes automatically. Run python examples/automatic_profiles.py for an offline end-to-end example.

API surface

from birkin_mnemosyne import (
    Mnemosyne,          # the mechanical index/ranking engine
    VaultMemory,        # ergonomic write_note / rezone / digest wrapper
    ProfileMemory,      # background-reviewed role-profile persistence
    ProfileReviewError, # invalid reviewer output
    run_curation_pass,  # the safe curation driver
    evaluate_plan,      # gate a structured plan (dry run by default)
    get_completer,      # provider registry (claude|codex|api|gemini|local)
    validate_clamp,     # the gate, if you want to inspect a plan without applying
    build_plan_prompt, extract_plan, mechanical_catalog,
    slug, tokenize, bm25_scores,
)

Key Mnemosyne methods: refresh(), search(query, limit, zone), related(slug), record_access(slug), stale(), rezone(slug, zone), zone_priorities(), stats().

Vault layout and configuration

Notes are slug-named Markdown files with YAML frontmatter and [[wikilinks]]. Zones are one-level directories; the vault root is the inbox.

my_vault/
  inbox-note.md
  devops/
    ingress-dns.md
  people/
  projects/
  identity/
  knowledge/
  journal/
  system/                       # role files owned by ProfileMemory
  _archive/                     # soft-forgotten notes
  .mnemosyne-index.json.z        # rebuildable compressed index cache
  .mnemosyne-dynamics.json       # persistent usage state

VaultMemory({"vault_path": "my_vault"}) selects the vault. The legacy vault configuration key is also accepted; without either key the default is ./vault. Mnemosyne takes the path directly. Set zone= when writing to choose placement; otherwise note types map to zones:

Note type Default zone
person people
project projects
preference identity
fact, topic knowledge
session journal

_archive is not an active curation zone. The cache can be rebuilt without discarding usage state; the legacy .mnemosyne-index.json cache is removed on the next save. Mixing older and newer library versions on one vault causes repeated cache rebuilding. MCP path and evidence settings are listed in the MCP server reference.

Credits

Extracted from the Birkin personal agent. The CJK-bigram BM25 approach was adopted by oh-my-openagent's memory recall in PR #9209, written by the same author. The lexical tie-break and keeping full lexical matches first follow that PR's review.

The hero image was AI-generated (ChatGPT image generation).

Local file memory, memory palaces, BM25, Ebbinghaus forgetting and plan-then-execute safety have prior art; the individual ingredients are not new. birkin-mnemosyne combines a standard-library lexical substrate with provider-portable, plan-only curation bounded by code. It is unrelated to the concurrent graph-memory project sharing the mythological name. The Python package imports as birkin_mnemosyne.

License

MIT; see LICENSE and NOTICE.

Metadata

Release files for birkin-mnemosyne 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for birkin-mnemosyne 0.4.0
File Size Uploaded
birkin_mnemosyne-0.4.0.tar.gz 100.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for birkin-mnemosyne 0.4.0
File Interpreter ABI Platform
birkin_mnemosyne-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 168.5 kB

Release files / birkin_mnemosyne-0.4.0.tar.gz

Download URL birkin_mnemosyne-0.4.0.tar.gz
Size 100.3 kB
Tags Source
SHA-256 checksum
How to use checksums
80e289841c1febb8899ac57531f4bf5d28ae2bfa97e5cf5d4637da2a8576fc3f
BLAKE2b-256 checksum
How to use checksums
435c7a08adfe83f5f9b7250a0534e1db05c2252f43b974a07a27b40c15f27dfe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / birkin_mnemosyne-0.4.0-py3-none-any.whl

Download URL birkin_mnemosyne-0.4.0-py3-none-any.whl
Size 68.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
52598fa6636de56577a8cf74b450883e3c9042ff74fbf6b6dc7a087a0e59f4db
BLAKE2b-256 checksum
How to use checksums
ed3b61cbd0d7d2b23656defe23684b15da4d6e4c60d47173fa4610f5573fa26d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page