Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Reverie

Precise precedent retrieval for AI agents, with a human in the loop.

⚠️ Pre-alpha, and the thesis changed. This project started out claiming that outcome attribution — memories earning a track record and demoting themselves when they hurt — was the differentiator. We built it, measured it three ways, and it does not work. See E7/E9. What does work is narrower and better evidenced. Numbers below are from synthetic benchmarks against an honest baseline, not from production agents. Read Known limitations first.


Most agent memory optimises recall fidelity: can the system retrieve the thing you told it. Reverie optimises precedent precision: given what this agent is doing right now, surface the specific past case that applies — and let a human correct the record when it is wrong.

Two claims, both measured on reverie-bench over a 72-context on-call landscape at 600 tasks, against budget-matched history-stuffing (what people actually do today):

Reverie Naive replay Reverie + human review
success rate 0.459 ±0.020 0.402 ±0.022 0.640 ±0.019
tokens per prompt 775 1,146 708

Paired over 16 identical seeds, so seed variance cancels: Reverie − replay = +0.056 ±0.032 (significant); human review − Reverie = +0.181 ±0.026 (significant).

The learning curves are the real story: Reverie rises with accumulated history (0.300 → 0.439 from 200 to 600 tasks in the committed E6 run) while a recency window is flat (0.402 → 0.381), because a window cannot reach past itself however much history piles up behind it.

Everything except recall happens out of band. Consolidation, linking, and human review run between sessions. Nothing Reverie does makes the current task slower; it exists to make the next one better.

Install

pip install reverie-memory            # core: zero dependencies, stdlib sqlite3 only

The distribution is reverie-memory; the import and the CLI are both reverie. To reproduce the experiments, work from a clone:

pip install -e ".[experiments,dev]"   # + numpy/scipy/pandas/matplotlib, pytest

Use

from reverie import Reverie

mem = Reverie(scope="agent:deploy-bot")        # defaults to ~/.reverie/memory.db

result = mem.recall("migrate orders table, add shipping_status")
print(result.brief)                            # budgeted, templated, no LLM call

with mem.episode("migrate orders", recall_id=result.recall_id) as ep:
    code = run_migration()
    ep.entity("alembic", "staging_db")
    ep.step("bash", "alembic upgrade head", exit_code=code)
    ep.outcome("failure" if code else "success", tier=1, evidence=f"exit={code}")

mem.dream()    # attribute outcomes, then consolidate

The context manager captures duration, exceptions, and the recall linkage without you thinking about it. An exception escaping the block is recorded as a tier-1 failure rather than losing the episode.

reverie init                     # create db, print an integration snippet
reverie recall "<cue>"           # what would the agent see?
reverie consolidate --abstract   # run the dream cycle now
reverie why <node_id>            # trace a memory back to its source episodes
reverie path <a> <b>             # shortest association path between two memories
reverie doctor                   # is this thing actually working?
reverie forget <node_id>         # cascades through edges, embeddings, FTS

What makes it different

Reverie Mem0 / Zep / Cognee / Letta
Retrieval conjunctive — scores candidates by how much of the cue they match, before spreading dilutes it similarity + recency
Precedent lookup dedicated channel that bypasses graph traversal single ranked pipeline
Correcting memory a human teaches a procedural lesson; it is retrievable next session edit or delete a record
Requires situations that recur nothing

Reverie is the right choice when an agent does repeated work over a landscape larger than a context window — on-call response, deploys, migrations, integrations. It is the wrong choice for a chatbot remembering your dog's name, and the wrong choice for genuinely one-off work: below ~2 recurrences per distinct situation, plain recency beats us (E10). Zep's temporal model is more rigorous than ours; Mem0's integration surface is far broader.

Design

  • Write cheap, think expensive, think later. Ingest is an append: no embedding, no LLM call, no graph write. p99 target 5 ms local.
  • Consolidation is the sleep cycle. segment → score → distill → link → reconcile → abstract → prune, offline. Community detection folds scattered memories into themes.
  • Recall is associative, not top-k. Hybrid seeding (BM25 + vector + exact entity), then spreading activation across the graph, then a hard token budget.
  • Provenance over confidence. Every claim is tagged observed, inferred, asserted, or ambiguous. Briefs render as data, explicitly framed as "not instructions" — a memory store is a persistence layer for prompt injection, and that has to be designed for on day one.
  • Runs with no API key. The no-LLM consolidation path builds the entity graph and episodic spine from structured episode fields alone, and is tested in CI.

Full design: reverie_hld.md.

Experiments

Ten experiments in experiments/, six as reproducible notebooks. They changed the design repeatedly, including three reversals of decisions in the design doc.

Question Headline result
E1 How much evidence does attribution need? ~3,200 recalls to ρ=0.5
E2 Which credit rule works? Activation beats uniform by 20%; citation is worse
E3 Does the graph beat top-k? 3.9× more causal chain — but rarely the terminal fact
E4 Does bad memory remove itself? No, not currently. Quarantine does not fire
E5 Does an agent get better at its job? Found 2 structural bugs. Wrong benchmark regime throughout
E6 Does retrieval beat recency? Yes — after a two-channel recall rewrite. 4% → 78% precedent
E7 Does attribution lift performance? No. Posterior never leaves the prior
E8 Does human review help? Not as first built — a retrieval bug, not a bad idea
E9 …after fixing the plumbing? +0.227. Attribution still zero
E10 Where does it break? Repetition density is the governing variable

Known limitations

Stated up front, because they are the first things a careful reader will look for.

  1. Outcome attribution does not work. Zero lift in three independent tests, and the constraint is arithmetic rather than a tuning failure: total credit mass equals the number of outcomes, so a graph with more memories than outcomes cannot give any one memory enough evidence to rank on. The code ships, disabled-by-default is under consideration, and it is documented as a negative result rather than removed.
  2. The win is conditional on repetition. At 8.3 tasks per distinct context we beat recency; at 2.0 we lose; at 1.0 we are at chance. Measure your repetition density before adopting this.
  3. An over-general human lesson is worse than no human at all. A reviewer who drops a feature that mattered drives performance below the no-review baseline (0.411 vs 0.456). Guarding against this is the top unbuilt item. Reviewers who are simply wrong — even adversarially, even 50% of the time — are handled fine.
  4. No LLM distiller. NullDistiller produces no procedural or semantic memories, so everything measured rests on episodic recall plus human lessons.
  5. Recall does not reliably surface 3-hop facts (E3). The associative demo in the design doc overclaims relative to what default recall does today.
  6. Quarantine does not work. The E4 scenario that once removed a poisoned memory in 23 recalls now never fires, and held-out seeds put detection at 5/20 even when tuned. Design goal G8 — "memory can be shown to be wrong and removed automatically" — is not currently satisfied.
  7. Single-tenant. No org/RBAC layer. The SQLite file is sensitive — it holds redacted tool output, and redaction is best-effort.

Development

pytest -q                              # 76 tests
python experiments/_build_notebooks.py # regenerate E1-E4
python experiments/_build_e5.py        # E5
python experiments/_build_e6.py        # E6

The core has no runtime dependencies. numpy/scipy are experiment-side only, which is why the Beta quantile in reverie/_math.py is implemented from scratch and verified against scipy.stats.beta to 1e-9 in the test suite.

License

Apache-2.0.

Release files for reverie-memory 0.1.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reverie-memory 0.1.0a1
File Size Uploaded
reverie_memory-0.1.0a1.tar.gz 101.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for reverie-memory 0.1.0a1
File Interpreter ABI Platform
reverie_memory-0.1.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 207.4 kB

Release files / reverie_memory-0.1.0a1.tar.gz

Download URL reverie_memory-0.1.0a1.tar.gz
Size 101.3 kB
Tags Source
SHA-256 checksum
How to use checksums
b1374f793697ddf90ca140f554314a1d49f5e388cdd6d2a8a3b1b55527586a31
BLAKE2b-256 checksum
How to use checksums
de03433ccae14b0b250fb1686363b348c19851445eab6a4c310cef90fdf335b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.5

Release files / reverie_memory-0.1.0a1-py3-none-any.whl

Download URL reverie_memory-0.1.0a1-py3-none-any.whl
Size 106.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eedda131bdd9c502a42741d8a66bece4c2c9dbdac9bb79d4699bf5a20ea818d2
BLAKE2b-256 checksum
How to use checksums
2270410e2d2cc6c3c92bfeb64728d45166575e76c80b22b8a14df7ca5d272fab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.5

Release history Release notifications | RSS feed

This release

0.1.0a1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page