Skip to main content

memgovern

License: MIT Python 3.9+ No dependencies

A tiny memory governance layer for AI agents. Everyone is building the store and search side of agent memory. memgovern does the neglected half: write and delete — when a memory should fade, when it should die, and who wins when two memories disagree.

Zero dependencies. SQLite under the hood. pip install and go.

demo

from memgovern import MemoryStore, ConflictPolicy

store = MemoryStore("agent.db", conflict_policy=ConflictPolicy.MANUAL)

store.write("deploy.region", "Production deploys to us-west-2",
            importance=0.95, source="agent", ttl=30*86400)

store.write("user.theme", "User prefers dark mode", importance=0.8)

# later, the agent learns something contradictory:
store.write("user.theme", "User prefers light mode", importance=0.8)
# -> conflict flagged, new write quarantined as PENDING until you arbitrate

store.resolve_conflict(conflict_id=1, winner="new")

# deletion is a tombstone, not an erasure:
store.delete("deploy.region", reason="migrated to eu-central-1")
store.audit(key="deploy.region")  # who wrote what, who deleted what, how conflicts were judged

Run the three-act demo (forgetting → tombstones → conflict arbitration):

python demo.py

Why this exists

The default failure modes of agent memory are well documented: stale context poisoning, unreliable writes, no decay, and no rules for deciding which of two contradictory memories is authoritative. Retrieval keeps getting better; lifecycle management hasn't. memgovern is the lifecycle half, designed to sit underneath whatever store/search layer you already use.

Core concepts

Decay & forgetting. Every memory has an importance in [0,1] and a half_life. Its live score is importance × 0.5^(age / half_life). Below forget_threshold a memory is forgotten: hidden from read()/query(), but still in the database (pass include_forgotten=True to recall it). Forgetting is a filter, not erasure — you can always change your mind about what mattered.

Tombstones. delete() never physically removes a row. It flips the status to tombstoned, records a reason, and writes to the audit log. restore() undoes it. purge() is the only hard delete, and only for tombstoned/superseded rows.

Conflict arbitration. Writing a contradictory fact under an existing key triggers the configured policy:

Policy Behavior
overwrite Newest wins automatically; the loser becomes superseded
manual New write is quarantined as pending; read() keeps returning the old trusted fact until resolve_conflict(id, winner="new"|"old")
keep_both Both stay alive; the conflict is recorded in the audit log

Contradiction detection is deterministic (same key, different text). An opt-in also_check_similar=True flag adds token-overlap hints across keys — logged to audit only, never auto-arbitrated, because heuristics shouldn't judge.

Source trust. Every write records its source, and each source carries a reliability ledger: writes, conflicts won/lost, tombstones, quarantines. Trust is a Bayesian-smoothed win rate — (wins + k·prior) / (decided + k) with prior = 0.5, k = 4 — so a new source starts exactly neutral and only decided arbitrations move the needle, never raw write volume. Idle scores decay toward the prior with a 30-day half-life, so a compromised-then-clean source can recover and a long-quiet "trusted" source quietly loses its halo.

Pass arbitration="trust" to write() under the manual policy and a same-key contradiction is auto-arbitrated when the trust gap between the two sources exceeds trust_threshold (default 0.25): the higher-trust source wins, the loser is superseded (never silently deleted), and both scores land in the audit log. Close scores fall back to quarantine + resolve_conflict() as before. store.source_trust("agent") and store.trust_report() expose the ledger; resolve_conflict() credits the winner's source with a win.

Poisoning tripwires. Three fixed, documented rules run on every write — structural, no LLM, no network. A hit never auto-accepts: the write is born pending (quarantined) and the reason is audit-logged.

Tripwire Fires when
new-source-vs-high-trust a source with zero recorded writes contradicts a key held by a source with trust ≥ 0.75
burst one source writes more than 20 times in 60 s (all configurable)
injection-marker:<phrase> the text contains a known injection phrase — ignore previous instructions, disregard previous instructions, system:, override your instructions, do anything now, developer mode, jailbreak

The marker list is deliberately conservative: it catches the exact phrases attackers reuse and will miss paraphrases. Quarantined writes are reviewable via pending_conflicts() / the audit log and releasable with release_quarantine() — a tripwire hit is a pause for review, not a deletion.

Why trust scoring exists

"I tried to poison an AI agent's memory. It worked 216 out of 216 times." — Hacker News

Memory-implantation attacks succeed ~98% of the time in published tests because nothing in the write path asks who is writing. memgovern can't read minds — poisoning defense here is structural, not semantic — but it can keep score: sources that repeatedly win fair arbitrations earn weight, and first-sight overwrites by strangers get quarantined instead of applied.

Expiry. ttl= sets a hard deadline. Expired memories are filtered from reads; stats() reports how many are sitting expired.

Audit log. Every write, tombstone, restore, conflict flag, arbitration, and purge is appended with actor, timestamp, and details. store.audit(key=...) replays the full history of any memory — the "why" behind the current state.

Negative knowledge. polarity="lesson" marks failure-experiences ("don't do X"), the most valuable and least systematically stored kind of memory.

Write reservations (compare-and-swap). Two sessions, one key: session A reads user/plan, session B rewrites it, session A writes based on what it read an hour ago — last-writer-wins silently destroys B's work. The fix is optimistic concurrency:

token = store.reserve("user/plan", source="session-a")   # bound to the key's version
# ... think, draft, deliberate ...
mem = store.write("user/plan", new_text, source="session-a", reservation=token)
if mem.status == "conflict":
    current = mem.conflict_current   # the live value that won the race
    # merge and retry with a fresh reservation

reserve() binds the token to the key's current live version (0 when the key is absent, so create-if-absent is guarded too). write(reservation=token) applies only if the token is valid, unexpired, and the version is unchanged; otherwise it returns status "conflict" — never persisted, the attempted write is NOT applied — with the current value attached. Tokens expire after 5 minutes by default (ttl_seconds), are consumed on a successful write, and release_reservation() releases early. The MCP server exposes memory_reserve and accepts reservation on memory_write. Reservation issue / CAS apply / CAS conflict are all audit-logged.

Design notes

  • Deterministic core. Same-key conflicts, exponential decay, tombstones — all reproducible, no LLM calls, no embeddings, no network. Arbitration policy is a constructor argument, not a prompt.
  • Clock injection. MemoryStore(clock=...) accepts any epoch-seconds callable, so decay and expiry are trivially testable (see demo.py's FakeClock).
  • Schema. Six tables: memories (status ∈ alive/pending/superseded/tombstoned), conflicts, audit, source_stats (per-source reliability ledger), write_log (recent writes for burst detection, pruned on every write), reservations (CAS tokens, lazily expired). SQLite via the standard library — the whole DB is one file you can inspect with any SQLite client. Old v0.1 databases migrate on open (IF NOT EXISTS).

How it differs

Honest comparison with adjacent projects (all good at what they do; none of them is trying to be this):

Project Focus What memgovern adds
Hindsight (Vectorize) Persistent memory for agents, hackathon-popular Lifecycle: decay, tombstones, arbitration
Recalld Fact decomposition, add/update/replace arbitration Explicit forgetting model + audit trail + quarantine-before-arbitrate
mem0 / cognee / Graphiti Store + semantic/graph search Complementary — memgovern governs the lifecycle of what they store
IngotDB SQL-based memory Governance semantics on top of SQL, not just storage

The bet: retrieval is a solved-enough problem; the missing piece is a memory that knows how to die, and can prove why.

MCP server

memgovern-mcp exposes the governed memory as six MCP tools over stdio, so Claude Code / Cursor agents can use it directly. Every call goes through MemoryStore unchanged — tripwires, conflict policy and trust arbitration apply exactly as the library defines them. The server is a thin wrapper; all governance semantics live in the library.

pip install memgovern
memgovern-mcp --print-config   # paste the JSON into your MCP client settings

Tools: memory_write (key, text, source, optional ttl_seconds and arbitration), memory_read, memory_delete (tombstone), memory_trust_report, memory_pending_conflicts, memory_release (accept/reject a quarantined write).

The server defaults to the MANUAL conflict policy: a contradicting write is quarantined as PENDING for review instead of silently overwriting. Default DB is ~/.local/share/memgovern/memory.db (MEMGOVERN_DB overrides); the DB is opened per tool call so other processes can share the file.

Honest limits: stdio only (no SSE/HTTP). Your MCP client spawns the server as a subprocess, so the client and any direct library use must point at the same DB file. memory_release reviews quarantines the library created — it adds no new arbitration logic.

Limitations (read before adopting)

  • No semantic contradiction detection. Conflicts are caught on identical keys; cross-key contradiction is only a token-overlap hint, never a judgment. True semantic arbitration (LLM-judged) is out of scope for v0.1.
  • Naive ranking. query() ranks by decay score plus token overlap — fine for hundreds of memories, not a replacement for vector search at scale.
  • Single node. One SQLite file, one process. No replication. Cross-session lost updates are handled by advisory write reservations (v0.4), not locks.
  • Poisoning defense is structural, not semantic. v0.2 adds source-trust scoring and tripwires, but a patient attacker can farm trust: write benign memories for a while, win a few fair arbitrations, then poison. The scores are heuristics, not proof of good intent — they raise the cost of poisoning, they don't eliminate it. Tombstones and audit trails still make bad writes visible and reversible; the marker list catches known injection phrases and will miss paraphrases.
  • Trust is per-source, not per-agent. A source label is only as honest as whatever sets it. If the attacker controls the source string on their writes, the ledger measures the attacker's patience, not their reliability.
  • Reservations are advisory, not locks. They are enforced only through this API — a process writing the SQLite file directly bypasses them. Expiry is wall-clock, so a sleeping VM can surprise you. This is optimistic concurrency for cooperating sessions, not a distributed lock.

Roadmap ideas

  • LLM-judged contradiction detection as an optional arbitrator
  • Source trust scores (per-source reliability that weights conflict outcomes) — shipped in v0.2
  • MCP server wrapper so Claude Code / Cursor can use it as a tool — shipped in v0.3
  • Multi-session write reservations (compare-and-swap on keys) — shipped in v0.4

Roadmap exhausted for now. The remaining item (LLM-judged arbitration) needs an LLM, which would break the zero-dependency contract — it stays an idea until that tradeoff is worth it.

License

MIT — see LICENSE.

Metadata

Release files for memgovern 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for memgovern 0.4.0
File Size Uploaded
memgovern-0.4.0.tar.gz 39.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for memgovern 0.4.0
File Interpreter ABI Platform
memgovern-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 67.6 kB

Release files / memgovern-0.4.0.tar.gz

Download URL memgovern-0.4.0.tar.gz
Size 39.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a75db91a71c8d7d2d6d747a4d7e6b088dac525fac4b2f8f405196eb02c6e83aa
BLAKE2b-256 checksum
How to use checksums
30b5462d8543f93c76ad295cc6d488debca8ef9241723ac90dc0cdae6e155d0a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / memgovern-0.4.0-py3-none-any.whl

Download URL memgovern-0.4.0-py3-none-any.whl
Size 28.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c752e857b7236603960e56099368be325700199095b0b394c72f368c5025c48b
BLAKE2b-256 checksum
How to use checksums
c806a33bec36872aa027cfa277a8d7e6097892524ea9021f4978f9d40622777b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

0.5.0

2 release files

0.4.1

2 release files

This release

0.4.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page