Skip to main content

hebb-memory

Memory written into a frozen model, one page per memory. Teach an open model new facts after it ships, without touching its weights. Delete any one of those facts exactly. See which fact an answer came from. Keep each customer's facts where only that customer's questions can read them.

import hebb_memory as hebb

mem = hebb.attach("Qwen/Qwen2.5-0.5B")                   # downloads the trained memory once

r = mem.remember("acme", "refund window", "60 days")    # written into its own page
mem.ask("acme", "refund window")                         # the model answers, reading that page
mem.trace("acme", "refund window")                       # which pages the answer read, and how much
mem.forget(r)                                            # gone, exactly
mem.forget("acme")                                       # or everything about acme

Why

A model is frozen the day it ships. Teaching it something new usually means fine-tuning, and a fine-tune is the wrong shape for facts that belong to one customer and may have to be deleted:

  • It is big. A LoRA adapter for one customer on Qwen2.5-0.5B is 8,798,208 bytes. Three of that customer's facts as pages are 6,948 bytes, 1,266x smaller. A page does not grow with the model; an adapter does.
  • It can't forget one thing. Facts trained into shared weights are smeared across all of them. Here each fact lives in one page, so deleting it means zeroing that page and nothing else.
  • It can't say where an answer came from. Here every read goes through pages, so the pages it read are the answer's sources.

Install

pip install hebb-memory

Or the latest from GitHub: pip install git+https://github.com/NeilGilani/hebb-memory.

Python 3.10+, PyTorch and Hugging Face transformers. It attaches to causal LMs whose decoder layers sit at model.layers (Qwen, Llama, Mistral, Gemma), transformer.h (GPT-2), model.decoder.layers (OPT) or gpt_neox.layers (Pythia). The measured results are on Qwen2.5-0.5B only.

Try it in a few minutes, on a CPU

hebb-memory train --model toy          # a tiny model, about 3 minutes on a laptop CPU
hebb-memory demo --model toy           # remember, ask, trace and forget a few facts

What it printed on our machine:

before anything is written:
  acme refund window -> 30 days
remembered  acme    refund window    = 60 days  page 0
remembered  acme    support channel  = phone    page 1
remembered  globex  refund window    = 14 days  page 2
remembered  globex  support channel  = email    page 3
ask         acme    refund window    -> 60 days  right   read: page 0 (acme refund window) 0.89, page 1 (acme support channel) 0.11
ask         acme    support channel  -> phone    right   read: page 1 (acme support channel) 0.86, page 0 (acme refund window) 0.14
ask         globex  refund window    -> 30 days  WRONG   read: page 2 (globex refund window) 0.78, page 3 (globex support channel) 0.22
ask         globex  support channel  -> email    right   read: page 3 (globex support channel) 0.74, page 2 (globex refund window) 0.26
forgot acme: pages [0, 1] zeroed and freed
  acme refund window -> 30 days (no acme pages left; this is the model alone)
  globex refund window -> 30 days

Every question read its own page first, and forgetting acme put its answer back to what the bare model says. One answer is wrong: the toy is a 128-wide, 4-layer model trained for three minutes, there to show the API working, not to judge accuracy.

Use it on a real model

A memory is trained once per base model: a key head that turns a sentence into an address, and read heads that let the frozen model attend to pages. The base model's weights never change.

You usually don't train anything. The first time you attach a model that has a published memory, attach downloads it into ~/.cache/hebb-memory (about 175 MB for Qwen2.5-0.5B), checks it against its published sha256, and uses the cached copy from then on. Published memories are the files in this repository's pretrained release, each trained by the pretrain workflow and self-checked before it is uploaded. Set HEBB_MEMORY_OFFLINE=1 (or pass download=False) to never download.

For any other model, train it yourself:

hebb-memory train --model Qwen/Qwen2.5-0.5B --ckpt-dir ckpt   # about an hour on an RTX 4070
hebb-memory check --model Qwen/Qwen2.5-0.5B                   # attach it and run the self-check

--ckpt-dir lets an interrupted run pick up where it stopped. To share what you trained, put the file on the Hugging Face Hub and point attach at it:

mem = hebb.attach("your-model", checkpoint="hf:<you>/<repo>")   # downloads memory.pt from it

The API

mem.remember(subject, attribute, value) Write one fact into a fresh page. A new value for the same subject and attribute replaces the old page. Returns a Record.
mem.ask(subject, attribute) The model's answer, reading only that subject's pages.
mem.choose(subject, attribute, options) Which option the model finds most likely, with scores. This is how the accuracy below is measured.
mem.trace(subject, attribute) The pages that question reads, with their weights: the answer's sources.
mem.forget(record | id | subject) Zero and free a page, or all of a subject's pages.
mem.records(subject=None), mem.subjects() What is stored.
mem.save(path, subject=None), mem.load(path) Move memories between processes or machines. A file holds pages, not a model: about 2 KB per fact.
hebb.selfcheck(mem) Are the heads trained, do writes change pages, do questions route to their page, does forgetting empty it?

A subject is whoever the facts belong to: a customer, a user, a project. Reads are limited to the subject's own pages before anything is compared, so one subject's question cannot read another's page at all.

What is exact, and what is not

Exact, and tested bit for bit (tests/test_memory.py):

  • Deletion. Write A, B and C, forget B, and the model's scores on every question are identical, to the last bit, to a model that was only ever given A and C.
  • Attribution. trace is the read path itself (the same addresses, top-k and weights that produce the answer), not an explanation made up afterwards.
  • Isolation. A read for one subject only ever has that subject's pages as candidates.

Not exact:

  • Answers can still be wrong. On the measured model, accuracy is well above chance and above the alternatives below, but not perfect.
  • Putting the facts in the prompt is still more accurate at small scale. In the same experiment, pasting every fact into the prompt scored 1.000. Pages win on size, deletion, attribution and isolation, not yet on accuracy against prompting; where prompting stops working as facts pile up has not been measured on a real model.
  • Facts are subject, attribute, value. That is the shape the memory was trained and measured on. Free-form notes are not supported in this version.

Results

Qwen2.5-0.5B. Each customer has three facts that conflict with other customers' facts. Two seeds, 800 meta-training episodes, forced choice among the possible values. Leakage (the share of answers that were another customer's value) is in parentheses.

customers chance facts in the prompt LoRA per customer pages
8 0.382 1.000 0.646 (0.354) 0.854 (0.146)
32 0.214 1.000 0.490 (0.510) 0.677 (0.323)
state per customer LoRA pages ratio
Qwen2.5-0.5B 8,798,208 B 6,948 B 1,266x

These were measured with the research code this library is taken from. One difference: the research code kept updating the mean of the key addresses as it ran, and this library freezes it at the trained value, because otherwise every write and every question nudges shared state and deletion stops being exact. Run hebb-memory check to measure your own setup.

How it works

The base model is frozen. Every other decoder layer gets a small read head that attends over memory tokens. A page is a few slots of those tokens, and each fact gets its own.

  • Writing a fact runs a few gradient steps on that page alone, so the frozen model, reading the page, says the fact. Nothing outside the page changes.
  • Reading turns the question into an address with a trained key head, picks the closest pages among the subject's own, and lets the read heads attend to them.
  • Training (once per base model) meta-learns the read heads, the page initialisation and the write step on facts about made-up companies, so that writing works on facts never seen in training.

The paper: One Page per Memory: Allocation Makes Parametric Memory Persistent, Deletable and Attributable, Given the Address (preprint in preparation).

License

Apache-2.0. Built by Hebb.

Metadata

Release files for hebb-memory 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hebb-memory 0.2.0
File Size Uploaded
hebb_memory-0.2.0.tar.gz 63.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hebb-memory 0.2.0
File Interpreter ABI Platform
hebb_memory-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 130.7 kB

Release files / hebb_memory-0.2.0.tar.gz

Download URL hebb_memory-0.2.0.tar.gz
Size 63.8 kB
Tags Source
SHA-256 checksum
How to use checksums
531a7b11b017638bf332c7eafcb525d5fc9088dadcc814799be0e2eccd14a952
BLAKE2b-256 checksum
How to use checksums
9b37fdda291935468b340632eb99d84a0a1c6656ea056791a9aec14f04856078
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / hebb_memory-0.2.0-py3-none-any.whl

Download URL hebb_memory-0.2.0-py3-none-any.whl
Size 66.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
23ebc7b5af23c2950cd036482c6fb585cad0c9bd13ccdc8bd84e64b1da7ba2a7
BLAKE2b-256 checksum
How to use checksums
c67f7ce41bd3daa49fb6da1e85031ac780a13b365753820b65508dfc4f98d0ba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page