Skip to main content

Anona Memory SDK

Official SDKs for Anona Memory — managed AI memory for intelligent agents. Record, retrieve, and reason over memories per user/space via a simple client, or auto-inject memory into LiteLLM calls with one line.

Documentation — Python SDK · TypeScript SDK · Quickstart · API reference · MCP server · Framework integrations

import { Anona } from "@anona-labs/memory";

const anona = new Anona({ apiKey: process.env.ANONA_API_KEY! });
await anona.record({ spaceId: "support", content: "Alice prefers email" });
const hits = await anona.retrieve({ spaceId: "support", query: "how to contact Alice" });
npm install @anona-labs/memory

Install

pip install anona

With the LiteLLM integration (or mcp for the MCP server):

pip install "anona[litellm]"

Quickstart

from anona import AnonaClient

# base_url defaults to https://api.anonalabs.com — pass it only to override.
client = AnonaClient(api_key="anona_live_...")

# Record a memory
client.record(space_id="space_123", content="User prefers dark mode.")

# Retrieve memories
results = client.retrieve(space_id="space_123", query="UI preferences", limit=5)
for r in results:
    print(r["relevance_score"], r["content"])

# Reason: a synthesized insight across memories
summary = client.reason(space_id="space_123", query="What do we know about this user?")
print(summary)

client.close()

Async ingestion (don't block on a write)

Recording runs fact extraction, so a normal record() takes a moment. In a chat loop or any latency-sensitive path, queue the write with background=True and poll the returned job instead:

import time

job = client.record(
    space_id="space_123",
    content="User prefers dark mode.",
    background=True,        # returns a job_id, doesn't wait
)

while True:
    status = client.get_job(space_id="space_123", job_id=job["job_id"])
    if status["status"] in ("completed", "failed", "cancelled", "not_found"):
        break
    time.sleep(2)

# Backfill many memories at once (always queued, up to 100 per call):
batch = client.record_batch(
    space_id="space_123",
    items=[
        {"content": "User is on the Pro plan."},
        {"content": "Signed up in 2024.", "timestamp": "2024-03-01T00:00:00Z"},
    ],
)
print(batch["accepted"], "queued as job", batch["job_id"])

Async variants (async_record, async_retrieve, async_reason) are available on the same client, or use it as a context manager:

async with AnonaClient(api_key="...") as client:
    await client.async_record(space_id="space_123", content="...")

API

AnonaClient(api_key, base_url="https://api.anonalabs.com")

  • record(space_id, content, metadata=None, background=False, timestamp=None) -> dict — store a memory; background=True queues it and returns a job_id; timestamp (ISO 8601) is when the event happened, for importing history
  • record_batch(space_id, items) -> dict — bulk-ingest up to 100 items (always queued); returns a job_id
  • get_job(space_id, job_id) -> dict — poll a queued job's status (free); status is one of pending / processing / completed / failed / cancelled / not_found
  • retrieve(space_id, query, limit=10, as_of=None, query_timestamp=None, occurred_after=None, occurred_before=None) -> list[dict] — see Time travel for the temporal arguments
  • reason(space_id, query) -> str | None
  • list_spaces() -> list[dict]
  • get_user_profile(space_id, user_id, *, limit=None, offset=None, memory_type=None, format=None, context_max_tokens=None) -> dict — everything the space has learned about one end user; see User profiles
  • ask_about_user(space_id, user_id, query, *, model=None) -> dict — one synthesised answer, drawn from that user's memories only
  • upload_file(space_id, file, *, filename=None, strategy=None, tags=None) -> dict — upload a file (path / bytes / file-like) so retrieval can draw on its content; ingested asynchronously, returns job_ids. PDF, DOCX, PPTX, XLSX, images (OCR), HTML, TXT/MD, CSV, audio. Files over 25 MB are rejected client-side.
  • list_documents(space_id, limit=100, offset=0) -> list[dict]
  • delete_document(space_id, document_id) -> None — remove a document and the memories extracted from it
  • get_graph(space_id, limit=500, min_count=1) -> dict — entity relationship graph (nodes + co-occurrence edges)
  • list_entities(space_id, limit=100, offset=0) -> list[dict]
  • get_entity(space_id, entity_id) -> dict — one entity + its observations
  • get_extraction_settings(space_id) -> dict / set_extraction_settings(space_id, mode=None, guidance=None, custom_prompt=None) -> dict / reset_extraction_settings(space_id) -> None — steer what a write keeps; see Extraction settings
  • get_chat_settings(space_id) -> dict / set_chat_settings(space_id, memory_limit=None, memory_token_budget=None, auto_record=None, memory=None) -> dict / reset_chat_settings(space_id) -> None — per-space defaults for the drop-in proxy endpoints
  • create_webhook(space_id, url, event_types=None, enabled=True) -> dict — the response carries secret, returned only on create
  • list_webhooks(space_id) -> list[dict], update_webhook(space_id, webhook_id, url=None, event_types=None, enabled=None) -> dict, delete_webhook(space_id, webhook_id) -> None
  • list_webhook_deliveries(space_id, webhook_id, limit=50, cursor=None) -> dict — recent attempts, for debugging a receiver
  • async_record(...), async_record_batch(...), async_get_job(...), async_retrieve(...), async_reason(...), async_get_user_profile(...), async_ask_about_user(...), async_list_spaces(...), async_upload_file(...), async_list_documents(...), async_delete_document(...), async_get_graph(...), async_list_entities(...), async_get_entity(...), async_get_extraction_settings(...), async_set_extraction_settings(...), async_reset_extraction_settings(...), async_get_chat_settings(...), async_set_chat_settings(...), async_reset_chat_settings(...), async_create_webhook(...), async_list_webhooks(...), async_update_webhook(...), async_delete_webhook(...), async_list_webhook_deliveries(...) — async equivalents
  • close() / aclose() — release underlying HTTP clients

Errors raise AnonaError(status_code, detail).

Time travel

Three arguments, and the distinction between them is the whole point: one is about when something happened, the other two are about what you knew and when.

# Importing history: date the memory to the event, not to the import run.
client.record(
    space_id="support",
    content="Renewed the enterprise contract",
    timestamp="2025-06-14T10:00:00Z",
)

# What did the space know in June, ignoring everything learned since?
client.retrieve(space_id="support", query="contract status", as_of="2026-06-01T00:00:00Z")

# Same corpus, but score recency and resolve "last June" against a past instant.
client.retrieve(space_id="support", query="what changed last June",
                query_timestamp="2026-01-01T00:00:00Z")

# Everything that HAPPENED in June 2025, however recently it was imported.
client.retrieve(space_id="support", query="contract status",
                occurred_after="2025-06-01T00:00:00Z",
                occurred_before="2025-06-30T23:59:59Z")
  • timestamp on record is when the event occurred. It feeds recency ranking, comes back on a result as timestamp, and is what the event-time window below matches against. It does not change when the memory was recorded, so it has no effect on as_of. It is not copied into occurred_start / occurred_end: those are filled only from a date found in the memory's own text, so a memory whose text names no date has both of them null while carrying a perfectly good timestamp.
  • as_of on retrieve is a hard cutoff: only memories recorded at or before that instant come back. A backdated import is recorded today no matter what timestamp it carries.
  • query_timestamp on retrieve only re-ranks. It moves the "now" that recency and relative dates are measured against, and never removes a result. Reach for as_of when you need the cutoff actually enforced.
  • occurred_after / occurred_before on retrieve bound when the thing happened, which is the one question as_of cannot answer: a year of history imported this morning has one record time and twelve months of event time. Either bound alone is an open-ended window, and the test is an overlap, so an event straddling an edge is inside.

User profiles

If you scope writes with user_id, a space accumulates a per-user history with nowhere to address it directly. These two read it back, instead of you reconstructing it from a wide-open search.

profile = client.get_user_profile("support", "alice_123")
print(profile["memory_count"], profile["last_active"])

# Prompt-ready instead of a list
block = client.get_user_profile(
    "support", "alice_123", format="block", context_max_tokens=500
)
messages = [{"role": "system", "content": block["context"]}, ...]

answer = client.ask_about_user(
    "support", "alice_123", "How does she prefer to be contacted?"
)
print(answer["insights"], answer["model"])

user_id has to be the same value your writes are scoped with — a typo on either side looks like an empty profile, not an error.

Three things worth knowing:

  • An unknown user is not a 404. A user_id is a scope tag created by the first write naming it, not a resource you register, so there is no valid set for a typo to fall outside of. A user nobody has recorded under comes back with memory_count of 0 and an empty memories list. An unknown space is still a 404.
  • memory_count can go down. The default view collapses layers: several raw facts become one synthesised note, and the note is what gets counted. A profile read during an import can genuinely go 115 → 67 → 15 while the corpus behind it grows the whole time. Read it as how many distinct things are currently known about this user, never as an ingestion counter, and never as the basis for a progress bar.
  • ask_about_user returns the whole response, not the answer string reason returns, because model reports which LLM actually answered. That is the rate the call was charged at, and it differs from what you asked for whenever you asked for nothing.

Extraction settings

Recording a memory is not storage. Your text goes through one pass that decides which facts are worth keeping and how they are phrased, and only what survives is stored — so a detail dropped there is not ranked low later, it is not there at all. These settings point that pass at your domain.

client.set_extraction_settings(
    "engineering",
    guidance=(
        "Engineering log. Always capture service names, metric values with "
        "units, and the named owner. Treat incidents as dated events. "
        "Skip standup small talk."
    ),
)
  • guidance is added to the standard rules and applies in every mode. Reach for it first — naming your vocabulary and the fields that always matter is what turns a good guess into a reliable one. Max 4,000 characters.
  • mode is concise (the default), verbose (keep every specific), verbatim (store the text as written, derive only the metadata around it), or custom.
  • custom_prompt replaces the standard rules, and only applies while mode is "custom". Max 8,000 characters.

Two things worth knowing. set_extraction_settings replaces the record, so anything you leave out is cleared. And settings apply to writes made after the call — memories already stored are never re-extracted, so changing these is safe and never rewrites history. To see the effect, save, record a representative piece of text, and read the memory back.

Unhelpful guidance produces no error; extraction simply keeps different things.

Framework adapters

Anona plugs into the Python agent frameworks through optional extras. Every adapter handles recall and storage for you, and scopes memories per end user.

Framework Install Import
LangChain / LangGraph pip install 'anona[langchain]' anona.integrations.langchain
CrewAI pip install 'anona[crewai]' anona.integrations.crewai
LlamaIndex pip install 'anona[llamaindex]' anona.integrations.llamaindex
Google ADK pip install 'anona[adk]' anona.integrations.google_adk
Microsoft Agent Framework pip install 'anona[msagent]' anona.integrations.ms_agent
AWS Strands pip install 'anona[strands]' anona.integrations.strands

All six are built on one MemoryBridge, which owns scope resolution and the failure contract:

from anona.integrations import MemoryBridge
from anona.integrations.langchain import AnonaMemory

bridge = MemoryBridge(
    api_key="anona_live_...",
    space_id="my-space",
    user_id="customer-42",   # optional scope: this user's memories only
)

agent = create_agent(model="gpt-4o-mini", middleware=[AnonaMemory(bridge=bridge)])

Memory failures never raise into your agent. A failed recall or store is logged and the agent runs on without memory, rather than taking your application down.

Each adapter's own module docstring documents its scoping, failure behaviour and per-call cost. Full docs: https://docs.anonalabs.com/integrations/langchain

Runnable end-to-end scripts for all six live in examples/.

LiteLLM integration

Auto-inject relevant memories into every litellm.completion() call, and auto-store the resulting Q&A pair:

from anona.integrations.litellm import AnonaMemory

mem = AnonaMemory(
    api_key="anona_live_...",
    space_id="space_123",
    recall_limit=5,       # how many memories to retrieve per call
    inject_mode="system", # "system" or "user"
    store_after=True,     # auto-store the exchange after each call
)
mem.enable()

# All subsequent litellm.completion() calls now auto-recall + auto-store.
import litellm
litellm.completion(model="gpt-4o", messages=[{"role": "user", "content": "..."}])

MCP server

The SDK ships an MCP server so any MCP client — Claude Desktop, Claude Code, Cursor — can read and write Anona memory as native tools: record, retrieve, list_spaces, and reason.

Install the extra:

pip install "anona[mcp]"

Claude Desktop / Cursor — add to claude_desktop_config.json (or ~/.cursor/mcp.json), then restart:

{
  "mcpServers": {
    "anona": {
      "command": "uvx",
      "args": [
        "--from",
        "anona[mcp]",
        "anona-mcp"
      ],
      "env": {
        "ANONA_API_KEY": "anona_live_...",
        "ANONA_SPACE_ID": "space_123"
      }
    }
  }
}

Claude Code — one command:

claude mcp add anona \
  --env ANONA_API_KEY=anona_live_... \
  --env ANONA_SPACE_ID=space_123 \
  -- uvx --from "anona[mcp]" anona-mcp

ANONA_SPACE_ID sets the default space so you can just say "remember this" without naming one; override it per call with the space_id argument. The key is personal — the server only reaches spaces you are a member of.

Requirements

  • Python >= 3.10
  • httpx >= 0.24
  • litellm >= 1.0 (optional, only for the LiteLLM integration)
  • mcp >= 1.2 (optional, only for the MCP server)
  • one of langchain, crewai, llama-index-core, google-adk, agent-framework-core, strands-agents (optional, only for the matching framework adapter — see the extras above for the verified version floors)

Documentation

Full guides and the complete API reference live at docs.anonalabs.com:

Page What it covers
Python SDK Every AnonaClient method, with examples
TypeScript SDK The Anona client and its adapters
Quickstart First key to first memory
Concepts Spaces, scoping, observations, credits
API reference The REST API these SDKs wrap
MCP server Local stdio and remote transports
Framework integrations LangChain, CrewAI, LlamaIndex, Google ADK and more

License

MIT

Metadata

Release files for anona 0.11.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for anona 0.11.1
File Size Uploaded
anona-0.11.1.tar.gz 114.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for anona 0.11.1
File Interpreter ABI Platform
anona-0.11.1-py3-none-any.whl Python 3 none any Details

Total release size: 176.6 kB

Release files / anona-0.11.1.tar.gz

Download URL anona-0.11.1.tar.gz
Size 114.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a6c30bc0c0ddb657805e1461dbc6f376d8b2eab5d7ad0ff3392ab61e5309e75c
BLAKE2b-256 checksum
How to use checksums
c120d67c3567faeb696ae5765a7c2d4b8d19be7557547db4649e503d19e4ee55
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / anona-0.11.1-py3-none-any.whl

Download URL anona-0.11.1-py3-none-any.whl
Size 62.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
161202578a9d127ad510360eabe13e3d7c5b10fe5e233fa6e33b692f216bb87f
BLAKE2b-256 checksum
How to use checksums
1bffd5beb308d39f2f51bc9d3498443b74279bae5217497bae741bbfa0774cd1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

0.16.0

2 release files

0.13.2

2 release files

0.12.3

2 release files

0.12.2

2 release files

This release

0.11.1 This release

2 release files

0.10.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page