longform-memory
Chapter 1000 gets the same token budget as chapter 10.
Documentation and the numbers behind it →
Zero dependencies — standard library only. Python 3.9+.
The problem
Write anything long with an LLM — a novel, a manual, a course, a screenplay — and you hit the same wall twice.
Withhold the past and it contradicts itself. A character who died in chapter 12 speaks in chapter 30. The key planted in chapter 3 is never mentioned again.
Send the past and you can't afford it. Concatenating every prior summary is O(n): by chapter 100 that's tens of thousands of tokens, by chapter 1000 it doesn't fit at all. And stuffing the window makes things worse — models skip the middle.
These are the same problem wearing two hats: there is no budget ceiling. Better retrieval doesn't fix it. A budget does.
What this does
Given whatever you know about the document so far, it assembles a fixed-size memory block and keeps it that size forever.
assembled block
|
61k -| ,---- no budget: O(n)
| ,-----'
30k -| ,-------------'
| ,-------------'
4k -|============+=============+=============+==== longform-memory: O(1)
+------+-----+------+------+------+------+----
10 100 500 1000 chapter
Measured on a real 1000-chapter book in production:
| before | after | |
|---|---|---|
| assembled memory block at chapter 1000 | 61,331 tokens | 4,396 tokens |
| growth from chapter 500 → 1000 | grows with chapter count | 2 tokens |
Install
pip install longform-memory
Quick start
from longform_memory import SectionInput, fit_sections, skeleton_chapters
chapter, recent_from = 1000, 980
block = fit_sections(
# Every section takes candidates sorted by DESCENDING importance —
# overflow is cut from the tail.
entity=SectionInput(cast, lambda e: f"{e.name} [{e.state}]"),
recent=SectionInput(summaries[recent_from:], lambda s: s.text),
skeleton=SectionInput(
[load(n) for n in skeleton_chapters(chapter, recent_from)],
lambda s: s.text,
),
retrieval=SectionInput(semantic_hits, lambda h: h.text),
total=6000, # total token budget — this is the whole point
)
block.used_tokens # <= 6000, at chapter 10 or chapter 10,000
block.recent.dropped # diagnostics: what got cut, and from where
The four sections split the budget 21.7 / 33.7 / 14.5 / 30.1 %. Whatever a section doesn't claim is redistributed to the sections that were trimmed, continuity first.
skeleton_chapters is where the O(1) guarantee lives. A fixed stride returns ~50 entries at chapter 1000 and keeps growing; here the entry count is capped first and the stride derived from it, so it never exceeds 12 entries whether the document is 100 chapters or 100,000. It samples summaries you already have — no model call, no recurring cost.
What's inside
| module | exports |
|---|---|
| budget | estimate_tokens · allocate · skeleton_chapters · fit_items · fit_sections |
| vector | encode_vector · decode_vector · search_top_k |
| language | resolve_language · dominates · is_char_counted_language · language_mismatch |
| threads | UNRESOLVED_THREAD_STATUSES · resolve_thread_deadline · plan_thread_ops · plan_owed_payoff · compare_thread_urgency · render_known_threads |
vector stores Float32 vectors as base64 in any TEXT column and scores them with brute-force cosine in pure Python — built for SQLite, LibSQL, Turso and Cloudflare D1, none of which can load the compiled extensions sqlite-vec and sqlite-vss require.
Interoperable with the TypeScript package
There is a TypeScript package of the same name — npm · source — with identical behaviour, and the vector wire format is byte-compatible in both directions: little-endian float32, normalised at write time, base64 encoded. Write vectors from a Node ingest job, read them from a Python worker, or the reverse.
That contract is pinned twice:
tests/test_cross_language.pyholds fixtures generated from the TypeScript implementation, so a drift in byte order or float width fails loudly instead of silently corrupting a store.- A CI job installs the published npm package and makes both implementations answer the same inputs, comparing byte for byte. Each test suite alone only proves self-consistency — that job is the one that catches "both sides self-consistent, mutually unreadable".
Three bugs that cost us weeks
1. One predicate, seven copies, two of them wrong. "Which loops are still open?" was answered in seven hand-written places; two forgot the progressing status. That created a ratchet — once a loop was marked as having advanced, it vanished from the list the model could cite, so resolving it became physically impossible. Payoff rate over 20 chapters: 0% (0/25). After collapsing the predicate to one definition: 63% (17/27).
2. A deadline of 0 is a dead value. Models answered 0 for "only the ending can resolve this", but overdue checks skip <= 0 — those loops were never overdue and nobody ever closed them. The fix isn't forcing the model to invent a number; it's giving 0 a meaning that can be judged: the ending means the final chapter.
3. The guard shared the bug, so it failed at the same moment. Script detection treated a single Han character anywhere as "this is Chinese". On a real English book, five chapters in, all 9 newly extracted records came back in Chinese — and the write-time guard, reading the same classification, considered that correct and let all of it through. dominates() is a majority test now.
What it does not do
- It stores nothing. No database, no file format, no server.
- It calls no model. Producing summaries and embeddings is your pipeline's job.
- Brute-force cosine is finite. Sized for one document — hundreds to a few thousand vectors.
- The token estimate is a heuristic, not a tokenizer, deliberately biased high.
- The similarity floor is calibrated per embedding model. Unrelated queries score 0.12–0.19 under
text-embedding-3-smallbut 0.275–0.348 underbge-m3; reusing one threshold across models silently disables the filter.
License
MIT © Emberspun
Extracted from Emberspun, where it runs in production on every chapter.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file longform_memory-0.1.1.tar.gz.
File metadata
- Download URL: longform_memory-0.1.1.tar.gz
- Upload date:
- Size: 23.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
76b7a65daa2a31743b2e67ebb048635b655d1bdd7429a62bb20cc3ac8fb1084a
|
|
| MD5 |
d2d6300621d8e85ff5e215b994cdfda4
|
|
| BLAKE2b-256 |
25dc49272a86e3c97c84b8f80aa6565dd8f4190460b03d2e9aac448b6b30fe2c
|
Provenance
The following attestation bundles were made for longform_memory-0.1.1.tar.gz:
Publisher:
release.yml on emberspun/longform-memory-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
longform_memory-0.1.1.tar.gz -
Subject digest:
76b7a65daa2a31743b2e67ebb048635b655d1bdd7429a62bb20cc3ac8fb1084a - Sigstore transparency entry: 2475338008
- Sigstore integration time:
-
Permalink:
emberspun/longform-memory-python@d19774a0ed79a3b759b850d02f03d042f899f5c2 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/emberspun
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d19774a0ed79a3b759b850d02f03d042f899f5c2 -
Trigger Event:
push
-
Statement type:
File details
Details for the file longform_memory-0.1.1-py3-none-any.whl.
File metadata
- Download URL: longform_memory-0.1.1-py3-none-any.whl
- Upload date:
- Size: 21.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4bc6514ebc6de740d822e6da59be8529d6376b0a863925fd266cabc1d3bc61b3
|
|
| MD5 |
0ec2f074d7495510334f04fe3ecf3b17
|
|
| BLAKE2b-256 |
0c42777c1836b9be00126e5bcb7e549f599a70e90786ef3c408f5e029e6db12e
|
Provenance
The following attestation bundles were made for longform_memory-0.1.1-py3-none-any.whl:
Publisher:
release.yml on emberspun/longform-memory-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
longform_memory-0.1.1-py3-none-any.whl -
Subject digest:
4bc6514ebc6de740d822e6da59be8529d6376b0a863925fd266cabc1d3bc61b3 - Sigstore transparency entry: 2475338060
- Sigstore integration time:
-
Permalink:
emberspun/longform-memory-python@d19774a0ed79a3b759b850d02f03d042f899f5c2 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/emberspun
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d19774a0ed79a3b759b850d02f03d042f899f5c2 -
Trigger Event:
push
-
Statement type: