Brain-inspired persistent memory layer for AI applications
Project description
hippomem
Brain-inspired persistent memory for AI applications.
hippomem is a memory layer that sits alongside your AI applications and gives them persistent, structured, evolving memory — across sessions and, eventually, across applications. The long-term goal is a single shared memory store that any AI application can read from and write to: one place where a user's knowledge, preferences, and context accumulate over time and remain accessible to any tool they use.
This is the first step toward that vision. Today, hippomem gives individual AI applications long-term memory that persists between conversations, builds up structured knowledge about the user, and stays coherent over time through a consolidation process.
What hippomem stores:
- Episodic memory — facts, preferences, and events from conversations
- Entity memory — people, pets, and organizations the user mentions
- Self memory — stable traits and facts about the user (e.g. job, location, habits)
All stored locally in SQLite + FAISS. No data leaves your machine.
hippomem does not try to remember everything. Unlike a fact store or a rolling message log, it is modeled on how human memory actually works: selective, lossy, and shaped by relevance. Memories that are used get reinforced; memories that go untouched decay. The hypothesis hippomem is built on is that this lossiness is not a weakness — it is what makes memory useful. A system that forgets selectively surfaces what matters, rather than drowning every response in accumulated context.
Note: hippomem is an actively evolving open-source project. It is functional and being used, but you should expect rough edges in both the implementation and documentation. If you find gaps or bugs, please raise a GitHub issue or open a pull request — contributions will directly shape what gets built next.
For detailed documentation, visit the docs on GitHub.
Install
pip install hippomem
Requires Python 3.11+.
Usage modes at a glance
| Mode | How |
|---|---|
| Daemon | hippomem serve — standalone service + Studio UI |
| Library | from hippomem import MemoryService — runs in your process |
| Client | from hippomem.client import HippoMemClient — connects to daemon over HTTP |
Quickstart: Daemon mode
Run hippomem as a persistent local service. Multiple apps can share one memory store, and you get the Studio UI for free.
cp .env.example .env
# Edit .env — set LLM_API_KEY at minimum
hippomem serve
# → API + Studio UI at http://localhost:8719
Options:
hippomem serve --port 8719 --host 127.0.0.1
Studio UI
The Studio UI is available at http://localhost:8719 when the daemon is running:
| Tab | What it shows |
|---|---|
| Dashboard | Memory counts, token usage, cost |
| Chat | Test the decode → LLM → encode loop interactively |
| Memory Explorer | List, grid, and D3 graph of stored episodic memories and entity profiles |
| Self | Self-traits learned about the user, grouped by category (goals, preferences, personality, etc.) |
| Inspector | Per-operation LLM traces with prompts, responses, token/cost/latency |
| Settings | Configure LLM connection, feature toggles, and advanced memory tuning |
Connecting your app
Use HippoMemClient to connect from any application:
from hippomem.client import HippoMemClient
async with HippoMemClient("http://localhost:8719") as mem:
result = await mem.decode("user_123", "What was I working on?")
# inject result.context into your LLM system prompt
response = await your_llm(system=result.context, message=user_message)
await mem.encode("user_123", user_message, response, decode_result=result)
HippoMemClient is included in the standard pip install hippomem.
Quickstart: Library mode
Embed memory directly in your app process. No separate service needed.
import asyncio
from hippomem import MemoryService, MemoryConfig
async def main():
memory = MemoryService(
llm_api_key="sk-...",
llm_base_url="https://openrouter.ai/api/v1", # or any OpenAI-compatible URL
)
async with memory:
user_id = "user_123"
history = []
user_message = "I'm building a FastAPI app with JWT auth."
# 1. Retrieve relevant memory before your LLM call
result = await memory.decode(user_id, user_message, conversation_history=history)
# 2. Inject result.context into your LLM system prompt
response = await your_llm(system=result.context, message=user_message)
# 3. Store the exchange after your LLM responds
await memory.encode(user_id, user_message, response, decode_result=result)
history.append((user_message, response))
asyncio.run(main())
See examples/demo.py and examples/chat_server.py for full working examples.
Direct memory search
Use retrieve() when you want raw search results rather than synthesized LLM context — for example to build your own UI, run analysis, or power a search feature:
result = await memory.retrieve(
user_id,
"FastAPI JWT auth",
mode="hybrid", # "hybrid", "faiss", or "bm25"
top_k=5,
)
for episode in result.episodes:
print(episode.core_intent, episode.entities)
retrieve() is independent of the decode/encode loop — you can call it at any time without affecting normal memory operation.
Configuration
Set environment variables in .env (copy from .env.example):
| Variable | Required | Default | Description |
|---|---|---|---|
LLM_API_KEY |
Yes | — | API key for your LLM provider |
LLM_BASE_URL |
No | https://openrouter.ai/api/v1 |
OpenAI-compatible base URL |
LLM_MODEL |
No | google/gemini-3.1-flash-lite-preview |
Model for hippomem's internal operations |
CHAT_MODEL |
No | Same as LLM_MODEL |
Model for /chat endpoint (daemon mode) |
SYSTEM_PROMPT |
No | Built-in default | Base system prompt in daemon mode |
DB_URL |
No | sqlite:///.hippomem/hippomem.db |
SQLite database path |
VECTOR_DIR |
No | .hippomem/vectors |
FAISS vector index directory |
For library mode, pass llm_api_key and llm_base_url directly to MemoryService. Everything else can be tuned via MemoryConfig:
from hippomem import MemoryService, MemoryConfig
config = MemoryConfig(
llm_model="x-ai/grok-4.1-fast",
db_url="sqlite:///my_app.db",
vector_dir="./my_vectors",
enable_entity_extraction=True, # extract people, orgs, pets (default: True)
enable_self_memory=True, # track stable user traits (default: True)
enable_background_consolidation=False, # periodic decay + clustering
)
memory = MemoryService(llm_api_key="sk-...", llm_base_url="...", config=config)
How it works
hippomem uses a cascade of LLM-powered steps inspired by how the hippocampus encodes and retrieves episodic memory:
- decode (before your LLM call): checks recent context continuity → retrieves relevant memory → synthesizes a context string ready to inject into your system prompt
- encode (after your LLM response): extracts new information → creates or updates memory engrams → links related memories via graph edges → updates entity and self-memory if enabled
- consolidate (periodic): compresses accumulated episode facts into clean baselines, enriches entity profiles, prunes stale self-traits, and synthesizes active user signals into a structured identity persona
- retrieve (optional): direct search API that returns raw structured results — episodes with linked entities and graph-connected neighbors — independent of the decode/encode lifecycle
Docs
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hippomem-0.3.0.tar.gz.
File metadata
- Download URL: hippomem-0.3.0.tar.gz
- Upload date:
- Size: 1.9 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
934c2e01cd9b833088c5a1a016b5bd0a7a208528249fefcec406c1f62ce1fcce
|
|
| MD5 |
9f29f823d40607e78fb7d4a7d6736349
|
|
| BLAKE2b-256 |
aecaa6d8c543af2161239cc35c1ca2cc394a96df7aae5ed5cddf3e5b2d040fcf
|
File details
Details for the file hippomem-0.3.0-py3-none-any.whl.
File metadata
- Download URL: hippomem-0.3.0-py3-none-any.whl
- Upload date:
- Size: 1.9 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dbc5d2d5f4b184482e98090d181619a5de8af394862d6fae419e46839f8846c1
|
|
| MD5 |
0554cc0acab88cff071a645844c7dcc9
|
|
| BLAKE2b-256 |
fcbccd527d3b5de5a9575a72b2f38856c01470259dd484fef7a8efc52dbef626
|