Skip to main content

Persistent memory service for AI coding agents — PostgreSQL + pgvector backend

Project description

Epimneme

Pronounced ep-im-NEE-mee. From Greek: ἐπί (epi-, "upon") + μνήμη (mnēmē, "memory") — meta-memory, a layer that sits upon memory itself.

Persistent memory for AI coding agents. PostgreSQL + pgvector backend, accessed via MCP or REST. Stop re-explaining your codebase every session.

License Python CI


Why Epimneme?

Every new chat with your coding agent starts from zero. You re-paste the same design decisions, re-explain the same gotchas, re-answer the same questions. Epimneme gives agents a long-term memory: facts, decisions, procedures, and a knowledge graph — all searchable, versioned, and deduplicated.

Agents connect via MCP (VS Code, Cursor, Claude Desktop, etc.) or through a plain REST API. Both use Bearer-token auth scoped per-project for multi-tenant safety.

Benchmarks

On LongMemEval (500 questions, 6 question types) with no LLM reranking and no benchmark-specific tuning:

Metric Epimneme (no LLM)
Recall @ 1 79.0%
Recall @ 5 96.0%
Recall @ 10 98.8%
Recall @ 50 100%
NDCG @ 10 0.887
LLM cost $0 / query

See benchmarks/BENCHMARK_RESULTS.md for the full write-up, including LoCoMo numbers and a per-category breakdown against MemPalace.

Quick Start

# 1. Clone
git clone https://github.com/alewman/epimneme.git
cd epimneme

# 2. Configure
cp .env.example .env
# Edit .env — at minimum set EPIMNEME_PG_PASSWORD to a strong value

# 3. Bring up Postgres + Epimneme
cp docker-compose.example.yml docker-compose.yml   # or edit in place
docker compose up -d --build

# 4. Wait for health, then create an admin API key
curl http://localhost:8000/health
docker exec epimneme python -m epimneme.manage create-key \
  --name admin --role admin

# Save the key it prints — it will not be shown again.

# 5. Create your first project
docker exec epimneme python -m epimneme.manage create-project \
  --name my-project --description "My awesome project"

Epimneme now listens on http://localhost:8000. See Connecting an Agent below.

Connecting an Agent

VS Code / Cursor / Claude Desktop (MCP over SSE)

{
  "mcpServers": {
    "epimneme": {
      "url": "http://localhost:8000/sse",
      "headers": {
        "Authorization": "Bearer epimneme_YOUR_KEY_HERE"
      }
    }
  }
}

Once connected, the agent sees these tools:

Tool Purpose
session_start Begin a session, receive previous context (summary, decisions, issues, entities)
session_end Close session with summary + handoff notes for the next agent
remember Store a memory (fact, decision, procedure, pattern, preference, observation, issue)
recall Search memories by semantic similarity + keyword. deep=true for graph traversal
project_status Get a project overview. Auto-registers new projects
entity_track Add a node to the knowledge graph (file, module, concept, tool, person, library, config, command)
entity_relate Add an edge (depends_on, part_of, uses, implements, …)
entity_explore Traverse the graph from an entity (configurable depth + direction)

curl

curl -H "Authorization: Bearer epimneme_YOUR_KEY" \
  "http://localhost:8000/api/memories/search?query=database+config&project=my-project"

Full REST reference: docs/API.md (or run the server and visit /docs for the live OpenAPI UI).

What Gets Stored

Seven memory kinds. Pick the right one and recall works much better later.

Kind Use for
fact Discrete knowledge — "CLI entry point is main.py"
decision Why something was done — "Chose PostgreSQL because pgvector + recursive CTEs cover both search and graph needs"
procedure Step-by-step instructions
pattern Recurring conventions — "All tests use pytest fixtures from conftest.py"
preference Working style — "Prefers small PRs with single-purpose commits"
observation General notes
issue Known bugs, limitations, tech debt

Plus a knowledge graph of entities (files, modules, concepts) and relationships (depends_on, uses, part_of, implements, …) that recall can traverse when deep=true.

Features

  • Hybrid retrieval — pgvector HNSW semantic search fused with PostgreSQL full-text + trigram keyword search via Reciprocal Rank Fusion.
  • FSRS-inspired decay — memories fade without access, stabilize with repetition. The retrieval ranker boosts well-used memories.
  • Dual dedup — SimHash (O(1) Hamming) catches minor rewordings; semantic cosine catches reworded but equivalent facts.
  • Versioningupdate_memory creates a new version instead of overwriting. Full history preserved.
  • Conflict surfacing — when you store a new fact/decision similar to an old one, the response flags the potential conflict so the agent can resolve it.
  • Periodic reflection — background job garbage-collects low-retrievability memories, consolidates clusters, and resolves detected conflicts. Pinned, persistent-project, decision, and procedure memories are exempt.
  • Multi-tenant — projects are namespaces; API keys are project-scoped (or global admin).
  • Dual auth — Bearer tokens for agents/programmatic access, OAuth passthrough (X-Forwarded-User) for browsers behind a reverse proxy.
  • Rate limited — per-IP token bucket, honours X-Forwarded-For.
  • Activity stream — in-memory ring buffer + text log for auditing.
  • Dashboard — self-contained web UI for inspecting memories, entities, activity, and backups.

Architecture

┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│  VS Code /   │     │   n8n /      │     │  Browser     │
│  Cursor      │     │   Scripts    │     │  (OAuth)     │
│  (MCP/SSE)   │     │  (REST API)  │     │              │
└──────┬───────┘     └──────┬───────┘     └──────┬───────┘
       │  Bearer            │  Bearer            │  Reverse proxy OAuth
       └────────────┬───────┴────────────────────┘
                    │
          ┌─────────▼─────────┐
          │   FastAPI + MCP   │  ← /api/*, /sse, /messages, /, /health
          │   (port 8000)     │
          └─────────┬─────────┘
                    │
          ┌─────────▼─────────┐
          │   MemoryManager   │  ← sessions, memories, entities, decay, dedup
          └─────────┬─────────┘
                    │
          ┌─────────▼─────────┐
          │  PostgreSQL 16    │
          │  + pgvector       │
          │                   │
          │  vectors +        │
          │  full-text +      │
          │  graph (rCTE)     │
          └───────────────────┘

Full details: ARCHITECTURE.md.

Configuration

All settings are environment variables (EPIMNEME_*). See .env.example and the Configuration section of ARCHITECTURE.md for the full list.

Minimum required:

Variable Notes
EPIMNEME_PG_PASSWORD Must not be the literal string epimneme. The server refuses to start otherwise.

Deploying Behind a Reverse Proxy

The bundled docker-compose.example.yml has commented-out Traefik labels. Uncomment and adjust for your proxy. Recommended setup:

  • Browser requests (/*) go through your OAuth/SSO middleware — users authenticate as admin.
  • API / SSE requests (/api/*, /sse, /messages, /health) skip OAuth and use Bearer tokens instead. Do not apply gzip compression to /sse — it breaks SSE streaming.

Development

# Editable install with dev extras
pip install -e '.[dev]'

# Run tests
make test                # local (uses mocks, no DB needed)
make test-cov            # with coverage report

# Lint
make lint                # ruff check
make lint-fix            # ruff check --fix

See CONTRIBUTING.md for full contributor guidelines.

Security

If you discover a security vulnerability, please open a GitHub Security Advisory rather than a public issue. See SECURITY.md.

License

Apache License 2.0 — see LICENSE and NOTICE.

Credits

  • Author & maintainer: Alewman
  • Design & implementation assistance: Anthropic's Claude (via GitHub Copilot Chat and Claude Code). Large portions of the code, tests, migrations, reranking, and documentation were authored in close collaboration with Claude across many sessions — many of which Epimneme itself made possible by persisting the context.
  • Built on FastAPI, PostgreSQL + pgvector, sentence-transformers, FlashRank, and the Model Context Protocol.
  • Benchmarked against LongMemEval and LoCoMo; compared to MemPalace.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

epimneme-0.7.0.tar.gz (254.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

epimneme-0.7.0-py3-none-any.whl (146.6 kB view details)

Uploaded Python 3

File details

Details for the file epimneme-0.7.0.tar.gz.

File metadata

  • Download URL: epimneme-0.7.0.tar.gz
  • Upload date:
  • Size: 254.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.0

File hashes

Hashes for epimneme-0.7.0.tar.gz
Algorithm Hash digest
SHA256 575a956548a1aef3794d8d78db83aa1d922babc0a80bd519026ec4565cdd34db
MD5 3045951fb65e0012546b9e9ac5567ea4
BLAKE2b-256 0718cebe1623f7b75bda7773c3450ddfff020033c54a548ab5c1706de00e5d6c

See more details on using hashes here.

File details

Details for the file epimneme-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: epimneme-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 146.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.0

File hashes

Hashes for epimneme-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a4e47e887d1e04d6d039c058eb8641231e46c05e337f4ba52f330feea6aa6a7a
MD5 604463af15e611ff281e248755802c90
BLAKE2b-256 e30f5037db1e214f3f5ea2ec9a98bcd748bc190cbd97b50f3ca32a83dcd5505d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page