korely-memory
The Python SDK for Korely Agents: memory for AI agents, with bi-temporal typed facts and contradiction checking built in.
A typed, zero-dependency client over the Korely REST API. Every method maps 1:1 onto an endpoint, so anything you can do with curl you can do here, and the JSON shapes in the API reference are the attribute shapes you get back. All the intelligence (embeddings, entity and typed-fact extraction, contradiction checking, bi-temporal validity) runs server-side, so your install stays small and your process stays light.
Install
pip install korely-memory
Python 3.9 or later, for the SDK, the korely CLI and the korely-mcp stdio
server alike: the MCP server needs no extra package since 0.1.16
(pip install 'korely-memory[mcp]' still works, the extra is empty).
Quickstart
from korely_memory import Korely
korely = Korely(api_key="kor_live_...", region="eu")
# or read the key from the environment (KORELY_API_KEY)
korely = Korely(region="eu")
# Remember: the write path extracts facts and resolves contradictions
korely.add("Maria lives in Rome", user_id="maria")
korely.add("Maria moved to Milan", user_id="maria")
# Recall the raw memories, ranked by meaning. Both come back: memories are
# kept as written, it is the facts extracted from them that get superseded.
for hit in korely.search("where does Maria live", user_id="maria", limit=5):
print(hit.id, hit.score, hit.snippet)
# One-call, prompt-ready context for your LLM. Its "Known facts" are the
# current ones, so the Rome fact, once superseded, is not among them.
# (Facts are extracted a few seconds after each write on the hosted service.)
ctx = korely.get_context("where should I send the package?", user_id="maria",
token_budget=800)
messages = [{"role": "system", "content": f"You are helpful.\n\n{ctx.context}"}]
Methods
Every method wraps exactly one REST endpoint.
| Method | Endpoint |
|---|---|
add(content, *, agent_id=, user_id=, run_id=, metadata=, timestamp=) |
POST /v1/memories |
search(query, *, user_id=, agent_id=, run_id=, metadata=, limit=) |
POST /v1/memories/search |
get_all(*, user_id=, agent_id=, run_id=, limit=, offset=) |
GET /v1/memories |
get(memory_id) |
GET /v1/memories/:id |
update(memory_id, *, content, expected_updated_at=) |
PATCH /v1/memories/:id |
delete(memory_id) |
DELETE /v1/memories/:id |
delete_all(*, user_id) |
DELETE /v1/users/:user_id/memories |
history(memory_id) |
GET /v1/memories/:id/history |
users(*, agent_id=, limit=, offset=) |
GET /v1/users |
list_agents(*, limit=, offset=) |
GET /v1/agents |
delete_agent(agent_id) |
DELETE /v1/agents/:agent_id |
get_facts(*, subject=, entity=, predicate=, predicate_family=, include_invalidated=, as_of=, …) |
GET /v1/facts |
add_fact_triple(subject, predicate, object, *, user_id=, valid_from=, tense=, …) |
POST /v1/facts |
correct_fact(fact_id, *, subject=, predicate=, object=) |
PATCH /v1/facts/:id |
forget_fact(fact_id, *, at=) |
POST /v1/facts/:id/forget |
get_profile(*, user_id, agent_id=, as_of=) |
GET /v1/profile |
get_context(*, query, user_id=, agent_id=, token_budget=) |
GET /v1/context |
events(*, user_id=, status=, limit=) |
GET /v1/events |
batch(memories) |
POST /v1/batch |
batch_status(job_id) |
GET /v1/batch/:id |
AsyncKorely has the same methods, awaitable.
add(..., timestamp="2026-01-15") backfills the past: facts extracted inherit
the timestamp as their valid_from, so as_of point-in-time queries reflect
when things were true, not when they were ingested. Each item of batch() takes
the same timestamp key, so a migration keeps its real dates:
korely.batch([
{"content": "Franco signed up on the Pro plan.", "user_id": "franco", "timestamp": "2026-01-15"},
{"content": "Franco downgraded to Free.", "user_id": "franco", "timestamp": "2026-06-20"},
])
A timestamp that is not an ISO 8601 date or datetime refuses the whole batch
with a 422 naming the item (memories[1].timestamp), before anything is queued.
list_agents() / delete_agent(agent_id) manage your agent namespaces: call
list_agents() after an agent_cap_exceeded error to reuse an existing
agent_id, or delete_agent() to purge a throwaway one. The page's total
counts the namespaces of this key's project; used counts the cap slots taken
across the account, which is what the 403 compares with cap. delete_agent()
raises NotFoundError for a name this project does not use, and its receipt's
slot_freed says whether the slot is free now (it is not while another project
of the account still uses the name).
delete_all(user_id=) answers with memories_deleted and facts_deleted, the
rows physically erased. memories_forgotten and facts_invalidated carry the
same numbers under their old names and are deprecated.
correct_fact() returns the new fact, whose invalidated lists every fact the
correction superseded (the corrected one, plus any the contradiction check
closed). A correction that restates the fact as it already stands supersedes
nothing: the same fact comes back, reconfirmed, with invalidated == [].
get_facts() returns a list of Fact that also carries .total, the number of
facts matching the filters across all pages, so offset knows when to stop.
get_all(), get_facts(), users(), list_agents() and events() take a
limit up to 200.
Bi-temporal facts
The differentiator: typed (subject, predicate, object) facts with validity over
time. Ask what was true on any date.
# Current state
facts = korely.get_facts(entity="Northwind Hosting")
print(facts[0].object) # 50 euro per month
print(facts[0].invalid_at) # None: active
# Point-in-time: what did we believe on June 1?
facts = korely.get_facts(entity="Northwind Hosting", as_of="2026-06-01")
print(facts[0].object) # 40 euro per month
Scoping
Three identifiers, three levels of scope, the same everywhere (SDK, REST, MCP):
agent_id: your application or agent (one namespace per product surface)user_id: your end user (free-form string; unlimited on every tier)run_id: one session or run (sub-scope inside a user)
korely.add("Asked to be contacted on Slack", agent_id="support-bot", user_id="customer-4812")
results = korely.search("contact preference", user_id="customer-4812")
Always pass
user_idon reads in multi-tenant products. Filters are additive (AND); a search withoutuser_idspans every end user in the namespace.
Error handling
Every error the server answers with is an APIError carrying the stable code
and the message of the REST error envelope ({"code", "message"}), so you can
branch on err.code; a self-hosted install that answers FastAPI's detail is
read the same way, and err.body keeps the response as it came. The common
statuses also have their own subclass. Everything subclasses KorelyError,
which is also what a client-side problem raises (no key, a connection error, a
timeout).
import time
from korely_memory import Korely, AuthenticationError, NotFoundError, QuotaExceededError
korely = Korely(api_key="kor_live_...")
try:
memory = korely.get("mem_8f2c1a")
except AuthenticationError:
raise # 401: check or rotate the key
except NotFoundError:
memory = None # 404: forgotten or never existed
except QuotaExceededError as err: # 429
if err.retry_after is None:
raise # monthly quota used up: nothing to wait for
time.sleep(err.retry_after) # rate limit: wait as long as the server said
memory = korely.get("mem_8f2c1a")
| Exception | Status | Typical code |
|---|---|---|
AuthenticationError |
401 | invalid_key |
NamespaceForbiddenError |
403 | agent_cap_exceeded, missing scope |
NotFoundError |
404 | not_found |
StaleWriteError |
409 | stale_write |
QuotaExceededError (.retry_after) |
429 | rate_limit_exceeded (has retry_after), quota_exceeded (monthly, retry_after is None) |
APIError |
any other, and the base of all of the above | invalid_request (422), search_unavailable / model_unavailable (503, safe to retry) |
The SDK does not retry on its own.
LangGraph
pip install 'korely-memory[langgraph]' # Python 3.10+, as LangGraph itself
korely_memory.integrations.langgraph gives a graph three ways to use Korely.
Take the ones you need: import korely_memory loads none of them, so the core
package keeps zero dependencies.
Context before the model answers
korely_context(client, user_id, query) makes one GET /v1/context call and
returns the text for a SystemMessage: the user's current facts and the
memories relevant to the question, within token_budget, under a
Current date: YYYY-MM-DD line. In our measurements the model answers better
when its prompt carries the date; include_date=False leaves it out and
today= sets it. Write each turn back with add(), and the next turn finds
it, on any thread.
from dataclasses import dataclass
from langchain.chat_models import init_chat_model
from langchain_core.messages import SystemMessage
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import START, MessagesState, StateGraph
from langgraph.runtime import Runtime
from korely_memory import Korely
from korely_memory.integrations.langgraph import korely_context
korely = Korely() # reads KORELY_API_KEY
model = init_chat_model("provider:model") # any chat model LangChain supports
@dataclass
class Context:
user_id: str
def call_model(state: MessagesState, runtime: Runtime[Context]):
user_id = runtime.context.user_id
question = state["messages"][-1].text
memory = korely_context(korely, user_id, question, token_budget=800)
reply = model.invoke([SystemMessage(memory), *state["messages"]])
korely.add([{"role": "user", "content": question},
{"role": "assistant", "content": reply.text}], user_id=user_id)
return {"messages": [reply]}
builder = StateGraph(MessagesState, context_schema=Context)
builder.add_node(call_model)
builder.add_edge(START, "call_model")
graph = builder.compile(checkpointer=InMemorySaver())
graph.invoke(
{"messages": [{"role": "user", "content": "Where should I send the package?"}]},
{"configurable": {"thread_id": "1"}},
context=Context(user_id="maria"),
)
The checkpointer keeps one thread's messages; Korely keeps what the user said
across all of them. akorely_context() is the same call for an async node.
Tools
from korely_memory.integrations.langgraph import create_korely_tools
tools = create_korely_tools(korely, user_id="maria") # [search_memory, save_memory]
model_with_tools = model.bind_tools(tools)
search_memory(query) answers with the same block as korely_context();
save_memory(content) stores one memory. The app binds the user (and
agent_id=) when it creates the tools: neither is in the tools' schema, so
the model can neither see nor change them. Create the tools per user, for
instance in the node that calls the model. A Korely error reaches the model as
the tool's answer instead of ending the run.
Store
from korely_memory.integrations.langgraph import KorelyStore
graph = builder.compile(checkpointer=InMemorySaver(), store=KorelyStore(korely))
# in a node: runtime.store.put(("memories", user_id), key, {"content": "..."})
KorelyStore is a LangGraph BaseStore, for code that expects one:
runtime.store, LangMem's memory tools. Each item is a Korely memory of the
user, so Korely extracts facts from it and korely_context() and the tools
find it. The memory's text is the value's content, text, memory or
data string, else one key: value line per field (index=[...] on a put,
or index_fields=, picks other fields). The value itself travels in the
memory's metadata and comes back exactly. Each put is a write like any
other: it is digested into facts and counts in your plan's writes, so the
store suits what a user said or decided, not caches or scratch state.
| Operation | On Korely |
|---|---|
put, get, delete |
Yes. The API has no lookup by your key, so each pages through the namespace's items, one request per 200. A put on an existing key stores the new memory, then forgets the old one. |
search(ns, query=...) |
Yes, one POST /v1/memories/search, with scores. offset + limit at most 50; filter takes equality on top-level fields. |
search(ns) without a query |
Yes, newest first, with every filter operator. |
list_namespaces, a prefix such as ("memories",), ttl, index=False |
No: NotImplementedError, before any request. |
The namespace ("memories", user_id) is that Korely end user;
KorelyStore(korely, agent_id="support-bot") adds the agent, and
namespace_to_scope= maps other shapes (a function returning the user id or
(user_id, agent_id)). A search covers exactly the namespace it names, never
the ones below it. Each namespace is also a Korely run,
run_id="langgraph:memories.maria": that is how the store reads its own items
and nothing else of the user's, while the rest of Korely reads them as the
user's memories. A value must fit in a memory's metadata, 8 KB of JSON on the
current server. delete forgets as delete() does; erasure is
delete_all(user_id=). Two writers on one key at the same instant can leave
two memories: reads take the newest, and the next put or delete removes
the other.
MCP server
pip install 'korely-memory[mcp]' # Python 3.9+ since 0.1.16
korely-mcp is a stdio MCP server with four tools (korely_get_context,
korely_add, korely_search, korely_get_facts), the same four the hosted
server at https://api.korely.ai/agent/mcp offers, with the same arguments
(korely_add takes timestamp) and the same fact lines: a fact whose end date
is still to come reads [until 2027-01-01], not [superseded ...], and dates
are UTC days. It reads the key from KORELY_API_KEY or from the file
korely init saved.
Links
MIT licensed.
Metadata
Release files for korely-memory 0.1.16
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| korely_memory-0.1.16.tar.gz | 77.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| korely_memory-0.1.16-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 130.6 kB
Release files / korely_memory-0.1.16.tar.gz
| Download URL | korely_memory-0.1.16.tar.gz |
|---|---|
| Size | 77.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c857bd9dc81a2703e2b6f2da30cc57dbf0884ac910a789b8a494e7c92107389f
|
|
BLAKE2b-256 checksum How to use checksums |
1a2712d9bd0923121d6c5fb24b1467c25e183215dc706d6f15fd3c6ebdcfea26
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|
Release files / korely_memory-0.1.16-py3-none-any.whl
| Download URL | korely_memory-0.1.16-py3-none-any.whl |
|---|---|
| Size | 53.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3da3ba6d766d80e72540b374806b2b51bc61937d7025d90e08970f9d105c54eb
|
|
BLAKE2b-256 checksum How to use checksums |
83954cd8a12fc3bf1f331ee5ba31f8ab47af32b7be55f49554874aabc9805e3c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|