Skip to main content

Nervon — Reasoning-Native Memory Framework for AI Agents

Project description

Nervon

Reasoning-Native Memory for AI Agents

Most AI memory systems are just vector databases with extra steps — store embeddings, retrieve by similarity, done. Nervon is different. It uses LLM reasoning at every stage: what to remember, when to update, what contradicts old knowledge, and what to forget.

Why Not Just Use a Vector DB?

Vector DB approach Nervon
Store Embed everything LLM extracts facts first
Retrieve Cosine similarity Similarity + temporal filtering
Update Append-only LLM decides: ADD / UPDATE / DELETE
Conflict Last-write-wins Temporal versioning (old → retired)
Context Raw JSON dump Prompt-ready get_context()

Architecture

Three memory tiers, inspired by how humans actually remember:

┌─────────────────────────────────────────────┐
│  Working Memory   — always loaded, O(1)     │
│  (key-value blocks, max 10 per user)        │
├─────────────────────────────────────────────┤
│  Semantic Store   — search by meaning       │
│  (vector + temporal versioning)             │
├─────────────────────────────────────────────┤
│  Episodic Log     — search by time          │
│  (conversation summaries, append-only)      │
└─────────────────────────────────────────────┘

Install

pip install nervon

Quick Start

from nervon import MemoryClient

# Initialize (stores in SQLite locally)
memory = MemoryClient(user_id="user-1")

# Add memories from conversation
memory.add("My name is Alice and I live in New York")

# Or from a message list
memory.add([
    {"role": "user", "content": "I just switched to Python 3.12"},
    {"role": "assistant", "content": "Nice upgrade!"}
])

# Search by meaning
results = memory.search("where does the user live")
for r in results:
    print(f"{r.content} (score: {r.score:.2f})")

# Get prompt-ready context (the killer feature)
context = memory.get_context("Tell me about the user")
print(context)
# Output:
# ## WORKING MEMORY
# • preferences: Likes dark mode
#
# ## RELEVANT MEMORIES
# • User's name is Alice (score: 0.92)
# • User lives in New York (score: 0.87)
#
# ## RECENT CONTEXT
# • [2024-03-19] Discussed Python upgrade (topics: python, upgrade)

How It Works

When you call memory.add():

  1. Extract — LLM pulls atomic facts from the conversation
  2. Compare — Each fact is checked against existing memories via embedding similarity
  3. Decide — LLM reasons about what to do:
    • ADD — New information, store it
    • UPDATE — Changed info (e.g., user moved cities) → retire old memory, create new
    • DELETE — Contradicted or irrelevant → retire the memory
    • NOOP — Already known, skip
  4. Summarize — Conversation is summarized and stored as an episode

Old memories aren't deleted — they're retired with a valid_until timestamp. You get full history.

Working Memory

For information that should always be in context (user preferences, active tasks):

memory.set_working_memory("preferences", "Prefers concise responses")
memory.set_working_memory("current_task", "Building a REST API")

# Always included in get_context(), no search needed
blocks = memory.get_working_memory()

Max 10 blocks per user. Think of it as the agent's "scratchpad."

Episodes

Every add() also creates an episode — a timestamped summary of the conversation:

episodes = memory.get_episodes(limit=5)
for ep in episodes:
    print(f"[{ep.occurred_at}] {ep.summary}")
    print(f"  Topics: {', '.join(ep.key_topics)}")

Configuration

memory = MemoryClient(
    user_id="user-1",
    db_path="my_app.db",           # SQLite path (default: nervon.db)
    llm_model="openai/gpt-4o-mini", # Any litellm-supported model
    embedding_model="openai/text-embedding-3-small",
    embedding_dim=1536,
)

Nervon uses litellm under the hood, so any provider works: OpenAI, Anthropic, Ollama, Azure, etc.

API Reference

Method Description
add(messages) Extract facts, compare, store. Returns list of memory IDs
search(query, limit=5) Semantic search over memories
get_context(query, max_tokens=2000) Prompt-ready string with all 3 tiers
set_working_memory(name, content) Upsert a working memory block
get_working_memory() Get all working memory blocks
get_episodes(limit=10) Get recent episode summaries
reset() Clear all data for this user

vs Mem0

Nervon was built after studying Mem0 (the $150M-valued AI memory startup). Key differences:

  • 3-tier vs flat — Mem0 has one memory tier. Nervon separates working memory, semantic memory, and episodic memory by access pattern.
  • Temporal versioning — Mem0 overwrites. Nervon retires old memories with timestamps — you can see what changed and when.
  • Prompt-ready output — Mem0 returns raw JSON. get_context() returns a formatted string you can drop straight into a system prompt.
  • No vendor lock-in — Pure Python + SQLite + litellm. No cloud service required.

Requirements

  • Python ≥ 3.11
  • An LLM API key (OpenAI, Anthropic, etc.)

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nervon-0.1.0.tar.gz (22.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nervon-0.1.0-py3-none-any.whl (19.1 kB view details)

Uploaded Python 3

File details

Details for the file nervon-0.1.0.tar.gz.

File metadata

  • Download URL: nervon-0.1.0.tar.gz
  • Upload date:
  • Size: 22.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for nervon-0.1.0.tar.gz
Algorithm Hash digest
SHA256 854f06c8a8bc2715584ed2400b6bf4e3aaa3acd714a511ebbe6719fd60a7fb0e
MD5 bf901bf946c2f8b6a5c59388d981e9db
BLAKE2b-256 81641ca8a40ea76ab945e6da3f170b3564591ce68f17bb41ba95b44175647818

See more details on using hashes here.

File details

Details for the file nervon-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: nervon-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 19.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for nervon-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7f127b32c76bf8c4b039f81cd7542f2b3e74e6f39a82eb067b35896e83e228e8
MD5 e04b3861964ddac80368b125dd9a1bc1
BLAKE2b-256 696fa8b852647429ec8af4e587fa70d353690534aea902ac4b55a955d2cfe385

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page