hebb-memory
Memory written into a frozen model, one page per memory. Teach an open model new facts after it ships, without touching its weights. Delete any one of those facts exactly. See which fact an answer came from. Keep each customer's facts where only that customer's questions can read them.
import hebb_memory as hebb
mem = hebb.attach("Qwen/Qwen2.5-0.5B") # downloads the trained memory once
r = mem.remember("acme", "refund window", "60 days") # written into its own page
mem.ask("acme", "refund window") # the model answers, reading that page
mem.trace("acme", "refund window") # which pages the answer read, and how much
mem.forget(r) # gone, exactly
mem.forget("acme") # or everything about acme
Why
A model is frozen the day it ships. Teaching it something new usually means fine-tuning, and a fine-tune is the wrong shape for facts that belong to one customer and may have to be deleted:
- It is big. A LoRA adapter for one customer on Qwen2.5-0.5B is 8,798,208 bytes. Three of that customer's facts as pages are 6,948 bytes, 1,266x smaller. A page does not grow with the model; an adapter does.
- It can't forget one thing. Facts trained into shared weights are smeared across all of them. Here each fact lives in one page, so deleting it means zeroing that page and nothing else.
- It can't say where an answer came from. Here every read goes through pages, so the pages it read are the answer's sources.
Install
pip install hebb-memory
Or the latest from GitHub: pip install git+https://github.com/NeilGilani/hebb-memory.
Python 3.10+, PyTorch and Hugging Face transformers. It attaches to causal LMs whose decoder
layers sit at model.layers (Qwen, Llama, Mistral, Gemma), transformer.h (GPT-2),
model.decoder.layers (OPT) or gpt_neox.layers (Pythia). The measured results are on
Qwen2.5-0.5B only.
Try it in a few minutes, on a CPU
hebb-memory train --model toy # a tiny model, about 3 minutes on a laptop CPU
hebb-memory demo --model toy # remember, ask, trace and forget a few facts
What it printed on our machine:
before anything is written:
acme refund window -> 30 days
remembered acme refund window = 60 days page 0
remembered acme support channel = phone page 1
remembered globex refund window = 14 days page 2
remembered globex support channel = email page 3
ask acme refund window -> 60 days right read: page 0 (acme refund window) 0.89, page 1 (acme support channel) 0.11
ask acme support channel -> phone right read: page 1 (acme support channel) 0.86, page 0 (acme refund window) 0.14
ask globex refund window -> 30 days WRONG read: page 2 (globex refund window) 0.78, page 3 (globex support channel) 0.22
ask globex support channel -> email right read: page 3 (globex support channel) 0.74, page 2 (globex refund window) 0.26
forgot acme: pages [0, 1] zeroed and freed
acme refund window -> 30 days (no acme pages left; this is the model alone)
globex refund window -> 30 days
Every question read its own page first, and forgetting acme put its answer back to what the bare model says. One answer is wrong: the toy is a 128-wide, 4-layer model trained for three minutes, there to show the API working, not to judge accuracy.
Use it on a real model
A memory is trained once per base model: a key head that turns a sentence into an address, and read heads that let the frozen model attend to pages. The base model's weights never change.
You usually don't train anything. The first time you attach a model that has a published
memory, attach downloads it into ~/.cache/hebb-memory (about 175 MB for Qwen2.5-0.5B),
checks it against its published sha256, and uses the cached copy from then on. Published
memories are the files in this repository's
pretrained release, each
trained by the pretrain workflow and self-checked before it
is uploaded. Set HEBB_MEMORY_OFFLINE=1 (or pass download=False) to never download.
For any other model, train it yourself:
hebb-memory train --model Qwen/Qwen2.5-0.5B --ckpt-dir ckpt # about an hour on an RTX 4070
hebb-memory check --model Qwen/Qwen2.5-0.5B # attach it and run the self-check
--ckpt-dir lets an interrupted run pick up where it stopped. To share what you trained, put
the file on the Hugging Face Hub and point attach at it:
mem = hebb.attach("your-model", checkpoint="hf:<you>/<repo>") # downloads memory.pt from it
The API
mem.remember(subject, attribute, value) |
Write one fact into a fresh page. A new value for the same subject and attribute replaces the old page. Returns a Record. |
mem.ask(subject, attribute) |
The model's answer, reading only that subject's pages. |
mem.choose(subject, attribute, options) |
Which option the model finds most likely, with scores. This is how the accuracy below is measured. |
mem.trace(subject, attribute) |
The pages that question reads, with their weights: the answer's sources. |
mem.forget(record | id | subject) |
Zero and free a page, or all of a subject's pages. |
mem.records(subject=None), mem.subjects() |
What is stored. |
mem.save(path, subject=None), mem.load(path) |
Move memories between processes or machines. A file holds pages, not a model: about 2 KB per fact. |
hebb.selfcheck(mem) |
Are the heads trained, do writes change pages, do questions route to their page, does forgetting empty it? |
A subject is whoever the facts belong to: a customer, a user, a project. Reads are limited to the subject's own pages before anything is compared, so one subject's question cannot read another's page at all.
What is exact, and what is not
Exact, and tested bit for bit (tests/test_memory.py):
- Deletion. Write A, B and C, forget B, and the model's scores on every question are identical, to the last bit, to a model that was only ever given A and C.
- Attribution.
traceis the read path itself (the same addresses, top-k and weights that produce the answer), not an explanation made up afterwards. - Isolation. A read for one subject only ever has that subject's pages as candidates.
Not exact:
- Answers can still be wrong. On the measured model, accuracy is well above chance and above the alternatives below, but not perfect.
- Putting the facts in the prompt is still more accurate at small scale. In the same experiment, pasting every fact into the prompt scored 1.000. Pages win on size, deletion, attribution and isolation, not yet on accuracy against prompting; where prompting stops working as facts pile up has not been measured on a real model.
- Facts are subject, attribute, value. That is the shape the memory was trained and measured on. Free-form notes are not supported in this version.
Results
Qwen2.5-0.5B. Each customer has three facts that conflict with other customers' facts. Two seeds, 800 meta-training episodes, forced choice among the possible values. Leakage (the share of answers that were another customer's value) is in parentheses.
| customers | chance | facts in the prompt | LoRA per customer | pages |
|---|---|---|---|---|
| 8 | 0.382 | 1.000 | 0.646 (0.354) | 0.854 (0.146) |
| 32 | 0.214 | 1.000 | 0.490 (0.510) | 0.677 (0.323) |
| state per customer | LoRA | pages | ratio |
|---|---|---|---|
| Qwen2.5-0.5B | 8,798,208 B | 6,948 B | 1,266x |
These were measured with the research code this library is taken from. One difference: the
research code kept updating the mean of the key addresses as it ran, and this library freezes
it at the trained value, because otherwise every write and every question nudges shared state
and deletion stops being exact. Run hebb-memory check to measure your own setup.
How it works
The base model is frozen. Every other decoder layer gets a small read head that attends over memory tokens. A page is a few slots of those tokens, and each fact gets its own.
- Writing a fact runs a few gradient steps on that page alone, so the frozen model, reading the page, says the fact. Nothing outside the page changes.
- Reading turns the question into an address with a trained key head, picks the closest pages among the subject's own, and lets the read heads attend to them.
- Training (once per base model) meta-learns the read heads, the page initialisation and the write step on facts about made-up companies, so that writing works on facts never seen in training.
The paper: One Page per Memory: Allocation Makes Parametric Memory Persistent, Deletable and Attributable, Given the Address (preprint in preparation).
License
Apache-2.0. Built by Hebb.
Metadata
Release files for hebb-memory 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hebb_memory-0.2.0.tar.gz | 63.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hebb_memory-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 130.7 kB
Release files / hebb_memory-0.2.0.tar.gz
| Download URL | hebb_memory-0.2.0.tar.gz |
|---|---|
| Size | 63.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
531a7b11b017638bf332c7eafcb525d5fc9088dadcc814799be0e2eccd14a952
|
|
BLAKE2b-256 checksum How to use checksums |
9b37fdda291935468b340632eb99d84a0a1c6656ea056791a9aec14f04856078
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / hebb_memory-0.2.0-py3-none-any.whl
| Download URL | hebb_memory-0.2.0-py3-none-any.whl |
|---|---|
| Size | 66.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
23ebc7b5af23c2950cd036482c6fb585cad0c9bd13ccdc8bd84e64b1da7ba2a7
|
|
BLAKE2b-256 checksum How to use checksums |
c67f7ce41bd3daa49fb6da1e85031ac780a13b365753820b65508dfc4f98d0ba
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log