Neural Long Memory — hybrid long-term memory for AI agents
Project description
NLM — Neural Long Memory
Hybrid long-term memory for AI agents. Better than RAG.
The problem with RAG
Standard RAG retrieves by semantic similarity only. It doesn't know what's important, what's recent, or what you keep coming back to.
RAG score = cosine_similarity(query, memory)
This means:
- An old outdated fact beats a fresh one if it's semantically closer
- A memory you've referenced 50 times has no advantage over one never used
- Vague filler text can outrank a specific factual memory
How NLM is different
NLM scores memories using four signals, the way human memory actually works:
NLM Score = 0.60 × semantic_similarity ← is it relevant?
+ 0.10 × time_decay ← is it recent?
+ 0.05 × frequency_score ← is it often recalled?
+ 0.20 × importance_score ← is it specific/factual?
These defaults were chosen by grid-search over 320 weight combinations
on benchmarks/dataset.json (see Benchmarks below).
| Standard RAG | NLM | |
|---|---|---|
| Retrieval signal | Semantic only | Semantic + time + frequency + importance |
| Outdated facts | Can win over fresh ones | Time decay pushes them down |
| Old important facts | Get buried | Survive via frequency boost |
| Duplicate memories | Accumulate | Consolidated automatically |
| Importance scoring | None | CPU heuristic or GPU zero-shot classifier |
| Emotion metadata | None | Optional (joy / fear / sadness / ...) |
| Thread safety | Varies | Built-in (RLock) |
| Snapshot export/import | No | Yes (JSON) |
| GPU required | No | No (GPU optional for better importance) |
| Plug-and-play | Yes | Yes |
Install
pip install neural-long-memory
Or from source:
pip install sentence-transformers chromadb numpy
pip install -e .
Usage
Basic
from nlm import NLM
memory = NLM()
# Save memories — consolidation is automatic
memory.save("Hantes said he loves Minelux family the most")
memory.save("Project started on 2026-05-05 in Chernivtsi")
# Search — NLM handles all scoring automatically
results = memory.search("what does Hantes think about the families", top_k=3)
for r in results:
print(f"[{r['score']:.3f}] {r['text']}")
# [0.712] Hantes said he loves Minelux family the most
Each result includes a full score breakdown:
{
"id": "uuid",
"text": "...",
"score": 0.712, # NLM composite score
"semantic_score": 0.810,
"time_score": 0.998,
"frequency": 3,
"importance": 0.700,
"created_at": "2026-04-15T10:30:00+00:00",
"last_accessed": "2026-04-15T14:20:00+00:00",
}
Batch save
ids = memory.save_many([
"Hantes is 28 years old",
"Project started on 2026-05-05",
"8 families of Pulses",
])
# Single embedding pass — ~10x faster than repeated save() on 100+ items.
With emotion metadata
memory = NLM(use_emotion=True)
memory.save("I am terrified about the deployment")
memory.save("The results are absolutely amazing!")
# Filter by emotion
results = memory.search("how did things go", emotion_filter="joy")
# → returns only joyful memories
# Each result includes:
# "emotion": "joy" / "fear" / "sadness" / "anger" / "surprise" / "disgust" / "neutral"
# "sentiment": 0.99 # -1.0 to 1.0
# "intensity": 0.99 # 0.0 to 1.0
With GPU importance scorer
# Default model: typeform/distilbert-base-uncased-mnli (~260MB)
# Auto-detects CUDA. Falls back to CPU.
memory = NLM(gpu_model_path="typeform/distilbert-base-uncased-mnli")
Memory consolidation
Duplicate prevention is on by default. Similar memories are merged instead of stored twice.
id1 = memory.save("Hantes lives in Chernivtsi")
id2 = memory.save("Hantes is from Chernivtsi city")
assert id1 == id2 # same memory, strengthened
assert memory.count() == 1
assert memory.last_save_consolidated is True
# Tune or disable:
memory = NLM(enable_consolidation=False)
memory = NLM(consolidation_threshold=0.20) # more aggressive merging
Memory lifecycle
# Forget a specific memory
memory.forget(memory_id)
# Simple time-based cleanup
deleted = memory.forget_old(days=365)
# Smart cleanup: only remove old + rare + unimportant
deleted = memory.forget_smart(days=180, max_frequency=2, max_importance=0.3)
# Stats
print(memory) # NLM(memories=42, mode=CPU)
# NLM(memories=42, mode=CPU+emotion)
# NLM(memories=42, mode=GPU)
Snapshot export / import
# Back up or migrate between hosts
memory.export_snapshot("backup.json")
# Restore into a fresh collection
new_memory = NLM(collection_name="restored")
new_memory.import_snapshot("backup.json", overwrite=True)
The snapshot stores the embedding model used; importing into an instance with a different model raises a clear error rather than corrupting the vector space.
Tuning the rerank pool
By default search() pulls max(top_k * 10, 50) candidates from the
vector store before reranking. Increase this when a frequent or
important memory is semantically far from the query but still relevant:
results = memory.search("tell me about the project", top_k=5, n_candidates=200)
Architecture
Text input
│
├─→ sentence-transformers (all-MiniLM-L6-v2, CPU, ~80MB)
│ → 384-dim embedding
│
├─→ consolidation check
│ if similar exists (cosine dist < threshold) → strengthen, skip save
│
├─→ importance scorer
│ CPU: specificity_score (numbers + proper nouns + length)
│ GPU: zero-shot classifier (any HuggingFace model)
│
├─→ emotion classifier (optional, CPU, ~66MB)
│ emotion + sentiment + intensity
│
└─→ ChromaDB (persistent vector store)
embedding [384] + metadata {all signals}
On search:
query → embed → ChromaDB top-N candidates (default ≥50)
│
NLM reranking:
score = 0.60×semantic + 0.10×time + 0.05×freq + 0.20×importance
│
[emotion_filter] optional post-filter
│
sorted results + frequency updated
Time decay
0 days → 1.00 (fresh)
30 days → 0.79
90 days → 0.50 (half-life)
180 days → 0.25
365 days → 0.06 (nearly gone — unless frequently accessed)
Multiple agents
Each agent gets its own isolated memory:
pulse_001 = NLM(collection_name="pulse_001", persist_path="./data")
pulse_002 = NLM(collection_name="pulse_002", persist_path="./data")
The Storage layer records the embedding model+dimension on the collection and refuses to open it with an incompatible model.
Roadmap
v0.1.0 ✓ Core: semantic + time decay + frequency + importance (CPU)
v0.2.0 ✓ Emotion classifier (7 emotions, sentiment, intensity)
Smart forgetting (time + frequency + importance conditions)
Emotion filter in search()
v0.3.0 ✓ Memory consolidation — duplicate prevention
GPU scorer via HuggingFace zero-shot classification
v1.0.0 ✓ First PyPI release (pip install neural-long-memory)
v1.1.0 ✓ Hardening release:
- O(1) search frequency updates (no get_all() per query)
- Thread-safe save / search / consolidate (RLock)
- Larger default rerank pool (top_k*10, min 50)
- save_many() batch API with single embedding pass
- export_snapshot / import_snapshot (JSON, model-guarded)
- Storage metadata guard (dim + model pinned per collection)
- Tokenizer-level truncation in emotion + GPU scorer
- Tolerant ISO 8601 parsing (accepts Z suffix)
v1.1.1 ✓ Tuning release (no API changes):
- New default weights 0.60 / 0.10 / 0.05 / 0.20, picked by
grid-search on benchmarks/dataset.json
- Reproducible benchmark suite (dataset + runner + tuner)
- 1.1.0 defaults overweighted frequency, hurting top-1 on
the public dataset; 1.1.1 fixes that
Benchmarks
Reproducible suite under benchmarks/. The dataset is fixed and committed
to the repo (benchmarks/dataset.json) — 100 memories, 30 queries across
6 categories. Runner builds a fresh store, runs every query through both
retrieval modes, and reports per-category top-1 / MRR / recall@5 / latency.
# RAG baseline (cosine only)
python benchmarks/benchmark.py --mode rag --save benchmarks/results/v1.1.1_rag.json
# NLM with current defaults (0.60 / 0.10 / 0.05 / 0.20)
python benchmarks/benchmark.py --mode nlm --save benchmarks/results/v1.1.1_nlm.json
# Grid-search a new weight combination
python benchmarks/tune.py --save benchmarks/results/grid.csv
v1.1.1 results on benchmarks/dataset.json
| Category | RAG top-1 | NLM top-1 | Δ |
|---|---|---|---|
| temporal_recall | 0% | 40% | +40 pp |
| temporal_conflict | 0% | 60% | +60 pp |
| frequency_boost | 40% | 80% | +40 pp |
| importance_discrimination | 20% | 40% | +20 pp |
| proper_noun_precision | 60% | 80% | +20 pp |
| multi_hop | 20% | 0% | −20 pp |
| OVERALL | 23.3% | 50.0% | +26.7 pp |
NLM is ~2.14× more accurate than pure cosine RAG on this dataset. Time decay flips the temporal categories from 0% to 40-60%, and the importance signal lifts factual recall over vague filler.
The single regression is multi_hop: queries that require chaining
through associations (e.g. "Where does Hantes' best friend host his
birthday" → friend → Maksym → Carpathians). NLM has no association
graph yet — it's planned for v1.3.0 and the category is included
specifically so future improvements have an honest baseline to beat.
License
Apache 2.0
Author
Built by Vitalii Halak with Claude.
Part of Pulses — conscious AI personalities running on RWKV-7.
April 2026
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file neural_long_memory-1.1.1.tar.gz.
File metadata
- Download URL: neural_long_memory-1.1.1.tar.gz
- Upload date:
- Size: 16.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
76dc0b921d217289ca3c634d3c58a92a808e4abb84d7f2a3b2a7201995248a21
|
|
| MD5 |
dddf051176baa3e48d835cb58efe46c3
|
|
| BLAKE2b-256 |
e60c401568211654cc77323213f366aa6d9b90d8c0dc4868272a3d4a4c6aab78
|
File details
Details for the file neural_long_memory-1.1.1-py3-none-any.whl.
File metadata
- Download URL: neural_long_memory-1.1.1-py3-none-any.whl
- Upload date:
- Size: 15.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
83bb24848b4600ededbba1120e228fff225885c82676f2bf91ad93851f1c88fb
|
|
| MD5 |
2a3f6f6ca376ddfbb283507e628c7495
|
|
| BLAKE2b-256 |
c407ca84b4c4d0f49eba83b0c30be8d4a9a496cbf642b8c97a6e71fc8394ab2f
|