Skip to main content

llama-index-memory-engram

Durable, explainable memory for LlamaIndex agents — powered by Engram.

EngramMemory is a BaseMemory implementation that replaces LlamaIndex's built-in chat-history buffer with Engram's hybrid retrieval pipeline (BM25 + vector + knowledge graph + reranker). Every message your agent sees is persisted to an Engram bucket; reads come back in chronological order, and semantic recall is one call away.

Install

pip install llama-index-memory-engram

Usage

import os
from llama_index.llms.openai import OpenAI
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.memory.engram import EngramMemory

os.environ["ENGRAM_API_KEY"] = "eng_live_..."   # or pass api_key=... explicitly

memory = EngramMemory.from_defaults(
    bucket="user-42",         # one bucket per user / session / agent
    read_limit=50,            # how many recent messages get() returns
)

agent = FunctionAgent(
    llm=OpenAI("gpt-4o"),
    tools=[...],
)

# IMPORTANT: pass memory to .run(), NOT to FunctionAgent(...)
response = await agent.run(
    "What did we decide about the Q3 launch?",
    memory=memory,
)

⚠️ Pass memory to agent.run(), not the constructor

Modern LlamaIndex FunctionAgent.__init__ accepts arbitrary **kwargs but silently swallows memory= — your memory backend will appear configured but never actually be touched. You must pass memory to each agent.run(..., memory=memory) invocation. This is a LlamaIndex API quirk, not an issue with this package; we caught it in our e2e test and it's the single most common pitfall when wiring EngramMemory.

If you want it to feel like a constructor arg, wrap once:

async def chat(prompt: str) -> str:
    r = await agent.run(prompt, memory=memory)
    return str(r)

Get an API key at https://lumetra.io. Keys look like eng_live_....

Direct semantic recall

get() returns the most recent read_limit messages, which is what agents expect from chat history. When you want hybrid retrieval over the entire bucket, call query() directly:

result = memory.query("regulatory risks we discussed last quarter")
print(result["answer"])
print(result["memories_found"])

Bucket scoping

Pick a bucket name per logical conversation scope:

EngramMemory(bucket=f"user-{user_id}")           # per user
EngramMemory(bucket=f"session-{session_id}")     # per session
EngramMemory(bucket=f"agent-{agent_id}")         # per agent

Buckets are created on first write — no admin call needed.

Self-hosted Engram

EngramMemory(
    bucket="ops",
    base_url="https://engram.internal.example.com",
    api_key="...",
)

API reference

Method Behavior
put(message) Append one ChatMessage to the bucket.
put_messages(messages) Append many.
get(input=None) Return the most recent read_limit messages, oldest-first.
get_all() Same as get().
set(messages) Clear the bucket, then write messages.
reset() Clear the bucket.
query(question) Hybrid retrieval over the entire bucket. Returns the raw Engram response.
list_buckets(limit, offset) List buckets visible to this API key.
delete_memory(memory_id) Delete a single memory by id.

All methods have async equivalents (aput, aget, ...) inherited from BaseMemory; they currently run the sync implementation in a thread.

Configuration

Constructor arg Env var Default
api_key ENGRAM_API_KEY required
bucket — "default"
base_url — "https://api.lumetra.io"
read_limit — 50
timeout — 120.0

License

MIT — see LICENSE.

For data-handling details see PRIVACY.md and https://lumetra.io/privacy.

Release files for llama-index-memory-engram 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-index-memory-engram 0.1.2
File Size Uploaded
llama_index_memory_engram-0.1.2.tar.gz 7.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-index-memory-engram 0.1.2
File Interpreter ABI Platform
llama_index_memory_engram-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 16.1 kB

Release files / llama_index_memory_engram-0.1.2.tar.gz

Download URL llama_index_memory_engram-0.1.2.tar.gz
Size 7.4 kB
Tags Source
SHA-256 checksum
How to use checksums
05b76bdd7947ccd6724186eae31c0372900633d4b6ee35567c610488754bb4dd
BLAKE2b-256 checksum
How to use checksums
dc0ec4ede76cc53f4d2cb7e1724c3c973d923bec79a18db0f92313c76fe79864
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.2

Release files / llama_index_memory_engram-0.1.2-py3-none-any.whl

Download URL llama_index_memory_engram-0.1.2-py3-none-any.whl
Size 8.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
57c1fec9e7084d24ffd91e93e3baa77c24afc2e41a11c46104ecac51aa4957e6
BLAKE2b-256 checksum
How to use checksums
57f84172b591e1f88a7d570fcf3cf507702081789ebc355dc8be576064bf7375
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.2

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page