Skip to main content

Local-first AI memory — semantic retrieval over your own SQLite + HNSW index.

Project description

Mnemonics

/nɪˈmɒnɪks/, ni-MON-iks (the "m" is silent, like "memories" with an N).

Agent memory infrastructure. Local-first, tier-aware, MCP-ready.

Mnemonics is a memory layer for AI agents: ingest conversations or facts, retrieve the relevant ones at inference time, decay old entries the way a brain does. Works as a Python library, a REST server, or an MCP server, plug it into any agent framework in under 10 lines.

pip install mnemonics

Agent integration guide · Benchmarks · MCP setup

Why

Most AI memory tools push your data to a hosted service. Mnemonics doesn't. Your index, your DB, your machine. Every retrieval is fully transparent: you see the raw cosine score, the decay factor, the age, and the tier on every result, nothing is silently demoted behind your back.

Works with LangChain, CrewAI, AutoGen, LlamaIndex, or any framework that can call a Python function or hit an HTTP endpoint. See AGENTS.md for drop-in examples.

Install

pip install mnemonics

CLI alias: both mnemonics and the shorter mnem are available after install.

Quick start

# Store something
mnem ingest "The Eiffel Tower is 330 meters tall and located in Paris."

# Retrieve (decay applied by default)
mnem retrieve "how tall is the Eiffel Tower"
#   [0.912] [raw=0.912 decay=1.00 age=0d tier=def] The Eiffel Tower is 330 meters tall ...

Tiers and decay

Every memory has a tier that controls how aggressively its score fades over time:

Tier Label Half-life Use for
0 pinned no decay decisions, key facts, things you must keep
1 default 90 days general notes (the default)
2 ambient 14 days low-confidence observations, chit-chat

Final score on every retrieval is:

score = raw_cosine × exp(-ln(2) × age_days / half_life)

Pinned (tier 0) entries always score at full weight. Tier-1 entries lose half their weight after 90 days, tier-2 after 14. Disable decay anytime with --no-decay (CLI) or decay=false (REST/MCP).

mnem pin <id>             # tier=0, never decays
mnem tier <id> 2          # tier=2, ambient (fast decay)
mnem retrieve "..." --no-decay   # show raw cosine scores

Maintenance

# Garbage-collect tier-2 rows not accessed in 30+ days (dry-run by default)
mnem gc --ns default --age-days 30
mnem gc --ns default --age-days 30 --apply   # actually delete

# Sweep tier-1 rows older than 60 days
mnem gc --ns sessions --age-days 60 --tier 1 --apply

# Bulk-delete an entire namespace (dry-run by default)
mnem forget --ns old-project
mnem forget --ns old-project --apply         # actually delete
mnem forget --ns archive --before 2026-01-01 # date-filtered

# Store health check (DB integrity, index vs SQL counts, orphan files)
mnem doctor
mnem doctor --json                           # JSON output for scripting

# Repair orphan vectors without re-encoding
# (Use when doctor shows "orphan vector(s)", loads vectors from old index by ID)
mnem rebuild-index --ns <ns>

# One-shot auto-repair: fixes all orphan vectors + removes orphan .bin files
mnem doctor --fix

Raw + summary

You can attach an optional one-line summary alongside the raw text. The embedding is still computed from the raw chunk, so vector recall is unchanged, but the BM25 mirror indexes both columns, so a keyword can land a hit via the gist even when the raw transcript reads differently.

mnem ingest "The full session transcript, jargon-heavy, full of code refs." \
    --summary "Zeus HTTP timeout bumped to 300s"

mnem retrieve "timeout" --hybrid
#   [...] Zeus HTTP timeout bumped to 300s
#       └─ raw: The full session transcript, jargon-heavy, full of code refs.

Useful when you want to keep the original wording (no summarization loss) but still let keyword search reach the row through a shorter label.

Every retrieval bumps access_count and last_accessed for the rows it returned, which sets up future reinforcement scoring without any caller action.

Python API

from mnemonics.store import Store
from mnemonics.ingest import ingest
from mnemonics.retrieve import retrieve

store = Store("~/.mnemonics")

ingest(["Paris is the capital of France.", "Rome is the capital of Italy."], store)

result = retrieve("what is the capital of France", store, top_k=3)
for r in result["results"]:
    print(f"[{r['score']:.3f}] tier={r['tier']} age={r['age_days']:.0f}d  {r['text']}")

store.pin(id) and store.set_tier(id, tier) change a memory's tier directly.

REST server

mnem serve --port 7810

The server binds to 127.0.0.1 only, no external interface, no telemetry.

Method Path Body
POST /ingest {"texts": [...], "ns": "default"}
POST /retrieve {"query": "...", "top_k": 5, "decay": true}
POST /search-bm25 {"query": "...", "ns": "default", "top_k": 5}
GET /stats (per-namespace tier breakdown)
POST /pin {"id": N}
POST /tier {"id": N, "tier": 0|1|2}
POST /rebuild-index {"ns": "..."}
POST /gc {"ns": "...", "age_days": 30, "tier": 2, "dry_run": true}
POST /forget {"ns": "...", "before": "2026-01-01", "dry_run": true}
POST /forget-ns {"ns": "...", "before": "...", "tier": N, "dry_run": true}
POST /repair (auto-fix orphan vectors and index files)
GET /health
GET /doctor (DB integrity, index vs SQL counts, capacity)
GET /namespaces
GET /count?ns=default
GET /memories?ns=default&limit=20&offset=0&tier=0 (browse newest-first, optional tier filter)
GET /memory/<id> (single memory by id)
DELETE /memory/<id>

MCP (Claude Code / Cursor / Metis)

mnem mcp
{
  "mcpServers": {
    "mnemonics": {
      "command": "mnem",
      "args": ["mcp"]
    }
  }
}

Tools exposed:

  • mnemonics_ingest
  • mnemonics_retrieve (decay-aware hybrid search, supports decay: false for raw cosine)
  • mnemonics_bm25 (pure keyword search, instant, no encoding, covers text + summary)
  • mnemonics_list (browse a namespace newest-first, with limit/offset/tier filter)
  • mnemonics_get (fetch a single memory by id)
  • mnemonics_forget (delete by id)
  • mnemonics_forget_ns (bulk delete a namespace; dry-run by default)
  • mnemonics_rebuild_index (rebuild index for a namespace from SQL source of truth)
  • mnemonics_pin
  • mnemonics_tier
  • mnemonics_gc (supports tier: 1 for default-tier rows)
  • mnemonics_stats (per-namespace tier breakdown: pin/def/amb counts)
  • mnemonics_health (DB integrity + index health report as JSON)
  • mnemonics_repair (auto-fix orphan vectors and orphan index files)

Namespaces

Isolate memories by project, user, or any key:

mnem ingest "project notes..." --ns work
mnem retrieve "deadlines" --ns work

Architecture

texts -> chunk (200w / 40w overlap) -> embed (all-MiniLM-L6-v2)
      -> hnswlib cosine index (per namespace)
      -> SQLite metadata store (id, ns, text, summary, meta, created,
                                tier, last_accessed, access_count)
      -> FTS5 mirror (text + summary), BM25 keyword surface

retrieve -> embed query -> knn search -> tier-aware decay -> ranked results
         -> UPDATE last_accessed, access_count on retrieved rows
         -> --hybrid: fuse vector top-k and BM25 top-k via RRF

Storage layout under ~/.mnemonics:

memories.db        SQLite (text, meta, tier, access counters, timestamps)
index_<ns>.bin     hnswlib index for each namespace

Privacy

  • The REST server binds to 127.0.0.1 only. There is no 0.0.0.0 flag.

  • The mnemonics package contains no outbound HTTP, no telemetry, no analytics.

  • First-run network: sentence-transformers downloads the all-MiniLM-L6-v2 model (~90 MB) from Hugging Face Hub on the first ingest or retrieve. The model caches under ~/.cache/huggingface/. After that first download, you can pin the package fully offline:

    export TRANSFORMERS_OFFLINE=1
    export HF_HUB_OFFLINE=1
    
  • DB encryption-at-rest is opt-in via SQLCipher (added in 0.3.0). The default install keeps ~/.mnemonics/memories.db as plaintext, so full-disk encryption (FileVault, LUKS) still matters for the default path. To turn it on:

    # macOS: brew install sqlcipher first (Linux: apt install libsqlcipher-dev)
    pip install 'mnemonics[encrypt]'
    mnemonics encrypt-db          # one-shot migrate ~/.mnemonics/memories.db
    export MNEMONICS_ENCRYPT=1    # add to your shell rc so future runs see it
    

    The migration auto-generates a 256-bit key and stores it in the OS keyring (macOS Keychain, Linux SecretService, Windows DPAPI). Pass --key <hex> to supply your own, or --no-keyring to print the key for manual storage. The pre-encryption DB is preserved as memories.db.preencrypt-<timestamp> so the operation is reversible. Losing the key after the original backup is gone loses the DB.

Benchmarks

Benchmark Metric Result
LongMemEval-S R@1 (det.) 0.958
LongMemEval-S R@5 1.000
LongMemEval-S R@10 1.000
LoCoMo Accuracy 82.9%

Results from the env-gated deterministic HNSW pipeline. Enable via MNEMONICS_DETERMINISTIC=1 with gte-Qwen2-7b-instruct (gate) and multilingual-e5-large-instruct (embed). Full reproduction steps in benchmarks/SOTA_PROOF.md.

License

Apache 2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mnemonics-0.4.0.tar.gz (231.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mnemonics-0.4.0-py3-none-any.whl (130.9 kB view details)

Uploaded Python 3

File details

Details for the file mnemonics-0.4.0.tar.gz.

File metadata

  • Download URL: mnemonics-0.4.0.tar.gz
  • Upload date:
  • Size: 231.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.5

File hashes

Hashes for mnemonics-0.4.0.tar.gz
Algorithm Hash digest
SHA256 97be542ca81b0cd6dd0e29312703d64dc9794416b6e4ebe9f2730338f4503262
MD5 a6a0dd1b003c27ad1f04869b8ec3acba
BLAKE2b-256 711eea3dd778e113e35d758febe5d9cca3439ac44a785d2c3050aa60f9e40e09

See more details on using hashes here.

File details

Details for the file mnemonics-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: mnemonics-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 130.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.5

File hashes

Hashes for mnemonics-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ab32a6c767eb4d1a6444bd3e09e8d887002cc42ec3b23bf293ff47de9657a229
MD5 1223a7358978a1c2fe4312dcc60c5410
BLAKE2b-256 b6e2236dacf3a8bd389db28715a980e0d8ff3e261d1fb1cc8b695653430ca42a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page