hebb-memory
Teach a frozen model new facts after it ships. Each fact is learned, by gradient descent, into its own small set of neural weights that the model reads as it thinks; the model's original weights never change. Delete any one fact exactly. See which fact an answer came from. Keep each customer's facts where only that customer's questions can use them.
import hebb_memory as hebb
mem = hebb.attach("Qwen/Qwen2.5-0.5B") # downloads the trained memory once
r = mem.remember("acme", "refund window", "60 days") # learned into its own small set of weights
mem.ask("acme", "refund window") # the model answers, using what it learned for acme
mem.trace("acme", "refund window") # which learned facts the answer used, and how much
mem.forget(r) # gone, exactly
mem.forget("acme") # or everything about acme
Why
A model is frozen the day it ships. Teaching it something new usually means fine-tuning, and a fine-tune is the wrong shape for facts that belong to one customer and may have to be deleted:
- It is big. A LoRA adapter for one customer on Qwen2.5-0.5B is 8,798,208 bytes. Three of that customer's facts learned by Hebb are 6,948 bytes, 1,266x smaller. A fact's weights do not grow with the model; an adapter does.
- It can't forget one thing. Facts trained into shared weights are smeared across all of them. Here each fact has its own weights, so deleting it means zeroing those and nothing else.
- It can't say where an answer came from. Here every answer reads through the learned facts' weights, so the facts it read are the answer's sources.
Install
pip install hebb-memory
Or the latest from GitHub: pip install git+https://github.com/NeilGilani/hebb-memory.
Python 3.10+, PyTorch and Hugging Face transformers. It attaches to causal LMs whose decoder
layers sit at model.layers (Qwen, Llama, Mistral, Gemma), transformer.h (GPT-2),
model.decoder.layers (OPT) or gpt_neox.layers (Pythia). The measured results are on
Qwen2.5-0.5B only.
Try it in a few minutes, on a CPU
hebb-memory train --model toy # a tiny model, about 3 minutes on a laptop CPU
hebb-memory demo --model toy # remember, ask, trace and forget a few facts
What it printed on our machine:
before anything is written:
acme refund window -> 30 days
remembered acme refund window = 60 days page 0
remembered acme support channel = phone page 1
remembered globex refund window = 14 days page 2
remembered globex support channel = email page 3
ask acme refund window -> 60 days right read: page 0 (acme refund window) 0.89, page 1 (acme support channel) 0.11
ask acme support channel -> phone right read: page 1 (acme support channel) 0.86, page 0 (acme refund window) 0.14
ask globex refund window -> 30 days WRONG read: page 2 (globex refund window) 0.78, page 3 (globex support channel) 0.22
ask globex support channel -> email right read: page 3 (globex support channel) 0.74, page 2 (globex refund window) 0.26
forgot acme: pages [0, 1] zeroed and freed
acme refund window -> 30 days (no acme pages left; this is the model alone)
globex refund window -> 30 days
Every question read its own page first, and forgetting acme put its answer back to what the bare model says. One answer is wrong: the toy is a 128-wide, 4-layer model trained for three minutes, there to show the API working, not to judge accuracy.
Use it on a real model
A memory is trained once per base model: a key head that turns a sentence into an address, and read heads that let the frozen model attend to pages. The base model's weights never change.
You usually don't train anything. The first time you attach a model that has a published
memory, attach downloads it into ~/.cache/hebb-memory (about 175 MB for Qwen2.5-0.5B),
checks it against its published sha256, and uses the cached copy from then on. Published
memories are the files in this repository's
pretrained release, each
trained by the pretrain workflow and self-checked before it
is uploaded. Set HEBB_MEMORY_OFFLINE=1 (or pass download=False) to never download.
For any other model, train it yourself:
hebb-memory train --model Qwen/Qwen2.5-0.5B --ckpt-dir ckpt # about an hour on an RTX 4070
hebb-memory check --model Qwen/Qwen2.5-0.5B # attach it and run the self-check
--ckpt-dir lets an interrupted run pick up where it stopped. To share what you trained, put
the file on the Hugging Face Hub and point attach at it:
mem = hebb.attach("your-model", checkpoint="hf:<you>/<repo>") # downloads memory.pt from it
The API
mem.remember(subject, attribute, value) |
Write one fact into a fresh page. A new value for the same subject and attribute replaces the old page. Returns a Record. |
mem.ask(subject, attribute) |
The model's answer, reading only that subject's pages. |
mem.choose(subject, attribute, options) |
Which option the model finds most likely, with scores. This is how the accuracy below is measured. |
mem.trace(subject, attribute) |
The pages that question reads, with their weights: the answer's sources. |
mem.forget(record | id | subject) |
Zero and free a page, or all of a subject's pages. |
mem.records(subject=None), mem.subjects() |
What is stored. |
mem.save(path, subject=None), mem.load(path) |
Move memories between processes or machines. A file holds pages, not a model: about 2 KB per fact. |
hebb.selfcheck(mem) |
Are the heads trained, do writes change pages, do questions route to their page, does forgetting empty it? |
A subject is whoever the facts belong to: a customer, a user, a project. Reads are limited to the subject's own pages before anything is compared, so one subject's question cannot read another's page at all.
What is exact, and what is not
Exact, and tested bit for bit (tests/test_memory.py):
- Deletion. Write A, B and C, forget B, and the model's scores on every question are identical, to the last bit, to a model that was only ever given A and C.
- Attribution.
traceis the read path itself (the same addresses, top-k and weights that produce the answer), not an explanation made up afterwards. - Isolation. A read for one subject only ever has that subject's pages as candidates.
Not exact:
- Answers can still be wrong. On the measured model, accuracy is well above chance and above the alternatives below, but not perfect.
- Putting the facts in the prompt is still more accurate at small scale. In the same experiment, pasting every fact into the prompt scored 1.000. Hebb wins on size, deletion, attribution and isolation, not yet on accuracy against prompting; where prompting stops working as facts pile up has not been measured on a real model.
- Facts are subject, attribute, value. That is the shape the memory was trained and measured on. Free-form notes are not supported in this version.
Results
Measured on this library, through its public API, with hebb-memory bench: Qwen2.5-0.5B with
the published memory. Each customer has three facts that conflict with other customers' facts;
every fact is learned with remember and asked with choose, among every value any customer
holds for it. Two evaluation seeds, so 48 questions at 8 customers and 192 at 32, run on GitHub's
CPU runners by the bench workflow.
| customers | chance | model alone | with hebb-memory | wrong answers |
|---|---|---|---|---|
| 8 | 0.408 | 0.458 | 0.896 | 0.104 |
| 32 | 0.209 | 0.208 | 0.750 | 0.250 |
Every option in this test is some customer's value, so every wrong answer is another customer's value. None of them came from another customer's memory: a question only ever reads its own customer's. Learning one fact took about 21 seconds on a 2-core CPU runner.
The paper ran the same test (the seed-0 customers and facts are identical; chance 0.382 and 0.214) with the research code, beside the alternatives:
| customers | facts in the prompt | LoRA per customer | Hebb, research code |
|---|---|---|---|
| 8 | 1.000 | 0.646 | 0.854 |
| 32 | 1.000 | 0.490 | 0.677 |
| state per customer | LoRA | Hebb | ratio |
|---|---|---|---|
| Qwen2.5-0.5B | 8,798,208 B | 6,948 B | 1,266x |
Pasting every fact into the prompt is still the most accurate at this size. The library freezes the mean of the key addresses at its trained value (the research code kept updating it), because otherwise every write and every question nudges shared state and deletion stops being exact; the numbers above are with it frozen. To measure your own setup:
hebb-memory bench --model Qwen/Qwen2.5-0.5B --customers 8 32 --seeds 0 1
How it works
The base model is frozen. Every other decoder layer gets a small read head that attends over memory tokens. A page is a few slots of those tokens, and each fact gets its own.
- Writing a fact runs a few gradient steps on that page alone, so the frozen model, reading the page, says the fact. Nothing outside the page changes.
- Reading turns the question into an address with a trained key head, picks the closest pages among the subject's own, and lets the read heads attend to them.
- Training (once per base model) meta-learns the read heads, the page initialisation and the write step on facts about made-up companies, so that writing works on facts never seen in training.
The paper: One Page per Memory: Allocation Makes Parametric Memory Persistent, Deletable and Attributable, Given the Address (preprint in preparation).
License
Apache-2.0. Built by Hebb.
Metadata
Release files for hebb-memory 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hebb_memory-0.3.0.tar.gz | 66.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hebb_memory-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 135.7 kB
Release files / hebb_memory-0.3.0.tar.gz
| Download URL | hebb_memory-0.3.0.tar.gz |
|---|---|
| Size | 66.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4c30c1784b60e15d098076125f97b83d27897a55d2643034cbb05f76aa94dbfe
|
|
BLAKE2b-256 checksum How to use checksums |
61041f090179f7bf31712b68746192071835af09a148248b59b0275895a3cfd1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / hebb_memory-0.3.0-py3-none-any.whl
| Download URL | hebb_memory-0.3.0-py3-none-any.whl |
|---|---|
| Size | 69.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
545b139a7dbaa645d80208f5d8ec893baf9a5e0640edc0dd3238148d5aed2f3f
|
|
BLAKE2b-256 checksum How to use checksums |
c9ce4e1f0f2a5151d0a6f1115830652d6bda75e0bca6ed46b5f6942f455642cf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log