llama-index-memory-engram
Durable, explainable memory for LlamaIndex agents — powered by Engram.
EngramMemory is a BaseMemory implementation that replaces LlamaIndex's
built-in chat-history buffer with Engram's hybrid retrieval pipeline
(BM25 + vector + knowledge graph + reranker). Every message your agent
sees is persisted to an Engram bucket; reads come back in chronological
order, and semantic recall is one call away.
Install
pip install llama-index-memory-engram
Usage
import os
from llama_index.llms.openai import OpenAI
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.memory.engram import EngramMemory
os.environ["ENGRAM_API_KEY"] = "eng_live_..." # or pass api_key=... explicitly
memory = EngramMemory.from_defaults(
bucket="user-42", # one bucket per user / session / agent
read_limit=50, # how many recent messages get() returns
)
agent = FunctionAgent(
llm=OpenAI("gpt-4o"),
tools=[...],
)
# IMPORTANT: pass memory to .run(), NOT to FunctionAgent(...)
response = await agent.run(
"What did we decide about the Q3 launch?",
memory=memory,
)
⚠️ Pass
memorytoagent.run(), not the constructorModern LlamaIndex
FunctionAgent.__init__accepts arbitrary**kwargsbut silently swallowsmemory=— your memory backend will appear configured but never actually be touched. You must passmemoryto eachagent.run(..., memory=memory)invocation. This is a LlamaIndex API quirk, not an issue with this package; we caught it in our e2e test and it's the single most common pitfall when wiringEngramMemory.If you want it to feel like a constructor arg, wrap once:
async def chat(prompt: str) -> str: r = await agent.run(prompt, memory=memory) return str(r)
Get an API key at https://lumetra.io. Keys look like eng_live_....
Direct semantic recall
get() returns the most recent read_limit messages, which is what
agents expect from chat history. When you want hybrid retrieval over the
entire bucket, call query() directly:
result = memory.query("regulatory risks we discussed last quarter")
print(result["answer"])
print(result["memories_found"])
Bucket scoping
Pick a bucket name per logical conversation scope:
EngramMemory(bucket=f"user-{user_id}") # per user
EngramMemory(bucket=f"session-{session_id}") # per session
EngramMemory(bucket=f"agent-{agent_id}") # per agent
Buckets are created on first write — no admin call needed.
Self-hosted Engram
EngramMemory(
bucket="ops",
base_url="https://engram.internal.example.com",
api_key="...",
)
API reference
| Method | Behavior |
|---|---|
put(message) |
Append one ChatMessage to the bucket. |
put_messages(messages) |
Append many. |
get(input=None) |
Return the most recent read_limit messages, oldest-first. |
get_all() |
Same as get(). |
set(messages) |
Clear the bucket, then write messages. |
reset() |
Clear the bucket. |
query(question) |
Hybrid retrieval over the entire bucket. Returns the raw Engram response. |
list_buckets(limit, offset) |
List buckets visible to this API key. |
delete_memory(memory_id) |
Delete a single memory by id. |
All methods have async equivalents (aput, aget, ...) inherited from
BaseMemory; they currently run the sync implementation in a thread.
Configuration
| Constructor arg | Env var | Default |
|---|---|---|
api_key |
ENGRAM_API_KEY |
required |
bucket |
— | "default" |
base_url |
— | "https://api.lumetra.io" |
read_limit |
— | 50 |
timeout |
— | 120.0 |
License
MIT — see LICENSE.
For data-handling details see PRIVACY.md and https://lumetra.io/privacy.
Release files for llama-index-memory-engram 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_memory_engram-0.1.2.tar.gz | 7.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_memory_engram-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 16.1 kB
Release files / llama_index_memory_engram-0.1.2.tar.gz
| Download URL | llama_index_memory_engram-0.1.2.tar.gz |
|---|---|
| Size | 7.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
05b76bdd7947ccd6724186eae31c0372900633d4b6ee35567c610488754bb4dd
|
|
BLAKE2b-256 checksum How to use checksums |
dc0ec4ede76cc53f4d2cb7e1724c3c973d923bec79a18db0f92313c76fe79864
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.2
|
Release files / llama_index_memory_engram-0.1.2-py3-none-any.whl
| Download URL | llama_index_memory_engram-0.1.2-py3-none-any.whl |
|---|---|
| Size | 8.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
57c1fec9e7084d24ffd91e93e3baa77c24afc2e41a11c46104ecac51aa4957e6
|
|
BLAKE2b-256 checksum How to use checksums |
57f84172b591e1f88a7d570fcf3cf507702081789ebc355dc8be576064bf7375
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.2
|