Skip to main content

Experience-driven memory for autonomous agents: learn from past successes and failures.

Project description

Mimir

PyPI Python License: MIT

Experience-driven memory for autonomous agents. Mimir helps agents learn from their past successes and failures instead of starting from scratch on every task.

Named after Mímir, the keeper of wisdom in Norse mythology.


The problem

Today's agents have memory, but they don't really learn.

Most frameworks store one of two things:

  • Conversation history (LangGraph memory, buffer memory)
  • Vector embeddings of documents (RAG, AGENTS.md, CLAUDE.md)

Both let an agent remember information. Neither lets it remember experience.

Task:     Fix authentication latency
Action:   Added Redis cache
Outcome:  Success

A month later, the agent has no meaningful understanding that this strategy worked. It solves the same class of problem from zero, every time.

The idea

Instead of storing documents, embeddings, and metadata, Mimir stores experiences:

Problem  →  Action  →  Outcome  →  Confidence  →  Context  →  Time

From a stream of experiences, Mimir reflects, extracts reusable strategies, and recommends actions for new tasks, so the agent gets measurably better over time.

from mimir import Mimir

memory = Mimir()

# Record what happened
memory.record(
    task="Fix authentication timeout",
    action="Implemented Redis caching",
    outcome="success",
    score=0.95,
)

# Recall relevant past experience
past = memory.recall("authentication latency")

# Get a recommended strategy with confidence
strategy = memory.recommend("login timeout")
# → Strategy: "Redis caching"  |  confidence: 0.87  |  based on 23 successes / 2 failures

How it differs from AGENTS.md / CLAUDE.md

AGENTS.md / CLAUDE.md Mimir
Knowledge type Static, hand-written rules Dynamic, learned from outcomes
Updates Manually edited Updates itself from results
Example "Use FastAPI. Use PostgreSQL." "Redis caching solved auth latency 23/25 times (92%)."
Failures Not tracked First-class: agents stop repeating mistakes

AGENTS.md answers "What should the agent remember?" Mimir answers "How does an agent accumulate experience and become wiser over time?"

Design principles

Mimir is built as a modular monolith Python library, not a microservice swarm or managed cloud product. The library is the product.

  • No LLM and no web server required for v1. Storage, retrieval, and ranking come first. Reflection via an LLM is added later, behind an interface.
  • Pluggable seams. Storage, embeddings, and the write path are interfaces, so scaling up (SQLite → Postgres → async reflection → Redis cache) is a swap, never a rewrite.
  • Derived knowledge is rebuildable. Strategies and reflections are computed from raw experiences and can always be regenerated.
  • Failures are first-class. Learning from what didn't work is treated as importantly as what did.

Architecture

┌────────────────────────────────────────────────────────────┐
│  Public API   Mimir()  .record() .recall() .recommend()     │
├────────────────────────────────────────────────────────────┤
│  Write chokepoint   ──►  [validation / provenance hook]      │   single write path
├──────────────┬───────────────┬─────────────────────────────┤
│  Episodic    │  Reflection   │  Recommendation             │
│  Engine      │  Engine       │  Engine                     │
│  (record/    │  (reflect/    │  (recommend / rank /         │
│   recall)    │   extract)    │   confidence)               │
├──────────────┴───────────────┴─────────────────────────────┤
│  Retrieval layer        (keyword + optional vector hybrid)   │
├────────────────────────────────────────────────────────────┤
│  Storage interface      SQLite (v1) · Postgres (v2) · …      │   pluggable
├────────────────────────────────────────────────────────────┤
│  Embedding provider     none (default) · local · API        │   pluggable
└────────────────────────────────────────────────────────────┘

Data model

Experience
  id, task, action, outcome (success|failure|partial),
  score (0..1), context (json), embedding (nullable),
  created_at, superseded_by (nullable)

Strategy   (derived)  problem_pattern, recommended_action, confidence,
                      success_count, failure_count, source_experience_ids
Reflection (derived)  summary, pattern, supporting_experience_ids, created_at

Installation

pip install mimir-learn

The distribution is named mimir-learn on PyPI, but you import it as mimir:

from mimir import Mimir

Optional extras (keyword recall works without either):

pip install "mimir-learn[embeddings]"  # local embeddings for semantic recall
pip install "mimir-learn[vector]"      # sqlite-vec ANN index for fast vector search

For development:

git clone https://github.com/AshNicolus/mimir.git
cd mimir
pip install -e ".[dev]"

Requirements: Python 3.10+, tested on 3.10, 3.11, and 3.12 (Linux, macOS, Windows). This matters for agents, which often run on the Python version their host ships, and 3.10 is still the default on several current Linux distributions. v1 has no required external services: storage is a local SQLite file. Semantic search and reflection are optional extras.

Quick start

from mimir import Mimir

memory = Mimir(db_path="mimir.db")

memory.record(
    task="Fix login latency",
    action="Added Redis cache in front of session lookups",
    outcome="success",
    score=0.9,
    context={"service": "auth", "language": "python"},
)

memory.record_failure(
    task="Throttle abusive clients",
    action="Added a fixed-window rate limiter",
    reason="WebSocket traffic wasn't handled; limiter only saw HTTP",
)

for exp in memory.recall("authentication is slow", k=5):
    print(exp.action, exp.outcome, exp.score)

print(memory.recommend("login times out under load"))

Recommendations

recommend() aggregates past experiences for a task and returns the action with the strongest track record, ranked by a relevance-weighted Wilson lower bound. An action proven on closely matching tasks outranks an equally successful one proven on loosely related tasks. Reported counts always cover the full matching population; weighting only affects ranking, and you can turn it off to compare:

memory.recommend("login times out under load", weight_by_relevance=False)

By default actions are grouped by normalized text, so "Added Redis cache" and "use redis caching" count as separate strategies. Plug in a clusterer to merge equivalent phrasings and pool their evidence:

from mimir import Mimir, EmbeddingClusterer

memory = Mimir(clusterer=EmbeddingClusterer(my_embedder))

ExactClusterer is the default and needs no embeddings. Any other strategy can implement ActionClusterer.

Hybrid recall and the vector index

Configure an embedder and recall becomes hybrid: keyword (SQLite FTS5) and vector candidates are fused with reciprocal-rank fusion, so an experience can be found by meaning even when it shares no words with the query. Install the vector extra to back vector search with a sqlite-vec ANN index; without it, recall falls back to a dependency-free Python cosine scan, so the behavior is identical and only the speed differs.

Superseding stale experiences

Knowledge goes stale. Mark an old experience as replaced, and it drops out of recall and recommendation by default while staying retrievable by id:

# record a replacement and link it in one call
new = memory.record("auth is slow", "add a read cache", supersedes=old_id)

# or link two experiences that already exist
memory.supersede(old_id, new.id)

Pass include_superseded=True to recall() or recommend() to see superseded rows anyway, which is useful for studying concept drift and decay.

Roadmap

Phase Goal Status
1: Episodic memory record() / recall(), outcome tracking, SQLite backend ✅ Done
2: Failure memory record_failure(), failures queried separately ✅ Done
3: Reflection engine reflect(): cluster experiences, synthesize patterns (LLM) Planned
4: Strategy extraction Turn experiences into reusable strategies with confidence Planned
5: Recommendation engine recommend(): rank strategies for a new task ✅ Relevance-weighted aggregation, pluggable action clustering (non-LLM)
6: Shared org memory Multiple agents learn from a shared store Future
Hybrid retrieval Keyword + vector recall, optional sqlite-vec ANN index ✅ Done
Reliability Versioned schema with a migration runner; staleness via superseded_by ✅ Done
Concurrency Per-thread connections so reads scale under WAL, writes serialized ✅ Done
Runtime support Run on the Python versions agent hosts actually ship, across Linux, macOS, and Windows Python 3.10–3.12

Scaling path

Mimir starts as a single SQLite file and grows by swapping seams, no rewrites:

  1. v1: SQLite, in-process; WAL with per-thread connections so reads scale across threads while writes serialize.
  2. v2: Postgres + pgvector backend for concurrent multi-agent writes.
  3. v3: extract the (slow, batch) reflection engine into an async worker.
  4. v4: Redis cache for hot/recent experiences on the read path.

Status

Alpha (0.1.3): published on PyPI as mimir-learn. Episodic and failure memory are complete and tested, and the recommendation engine works today via relevance-weighted outcome aggregation with pluggable action clustering (no LLM). 0.1.3 adds hybrid keyword + vector recall (optional sqlite-vec ANN index), schema versioning with a migration runner, concurrent reads under WAL, experience superseding for staleness, and an outcome/score consistency check. APIs may still change before 1.0. Feedback and ideas welcome.

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mimir_learn-0.1.3.tar.gz (31.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mimir_learn-0.1.3-py3-none-any.whl (24.6 kB view details)

Uploaded Python 3

File details

Details for the file mimir_learn-0.1.3.tar.gz.

File metadata

  • Download URL: mimir_learn-0.1.3.tar.gz
  • Upload date:
  • Size: 31.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mimir_learn-0.1.3.tar.gz
Algorithm Hash digest
SHA256 744c254e8cb77170ad0e6ba81450d9bcf876008fb3953d0a6d6d050a74c0da07
MD5 5998a4963896074913a9931e2c161bd7
BLAKE2b-256 628f946a4c9f8292ad9c969aad46b29b34dab6f5839e9b1af7d8f80f292d4a0d

See more details on using hashes here.

Provenance

The following attestation bundles were made for mimir_learn-0.1.3.tar.gz:

Publisher: publish.yml on AshNicolus/mimir

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mimir_learn-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: mimir_learn-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 24.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mimir_learn-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 b0521ada94def478403af2574a062ea0521a47fcdba980ccc2e34392b6221a32
MD5 d60d637fb9bdc7d4dba7f683d31c37a6
BLAKE2b-256 1506e7096f566e4cc12ad9ec3d84af5e97a23d0658c38ab5bdd3eaabe7c4e6a2

See more details on using hashes here.

Provenance

The following attestation bundles were made for mimir_learn-0.1.3-py3-none-any.whl:

Publisher: publish.yml on AshNicolus/mimir

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page