slim-llm-memory
Local memory and retrieval for LLM apps. Put your notes or docs into a store, ask a question, and get back the passages to paste into a prompt, or a cited answer from a local model.
It needs only numpy and httpx. Embeddings come from Ollama on your own machine. It is meant for personal projects and research code, up to about 50,000 passages per store, where one numpy array is fast enough and a vector database is extra work.
Install
pip install slim-llm-memory
ollama pull nomic-embed-text # embeddings
ollama pull llama3.2:3b # only needed for answer()
Python 3.10 or newer. Optional extras:
| Extra | Adds |
|---|---|
slim-llm-memory[rerank] |
cross-encoder reranking (sentence-transformers) |
slim-llm-memory[graph] |
links between documents (NetworkX) |
slim-llm-memory[mcp] |
the slim-memory-mcp server for MCP clients |
slim-llm-memory[obsidian] |
experimental Obsidian vault ingest; Python API only, no command yet |
The [gemini] and [anthropic] extras are placeholders. No code uses them yet.
Quickstart
from slim_llm_memory import topic
t = topic("nginx") # a store in ~/.slim-llm-memory/topics/nginx
t.add("docs/nginx/") # every .md, .txt and .rst file in the folder
r = t.ask("how do I enable TLS?") # the best matching passages
print(r) # hits, scores and timings
print(r.context) # numbered passages, ready to paste into a prompt
print(t.answer("how do I enable TLS?")) # a local model answers from those passages and cites them
add also takes a single file, raw text, or a {name: text} dict. Running it again only
re-embeds what changed, and t.forget("old.md") removes a document. To try the API
without Ollama, pass embedder="noop": every call works, but the similarity scores mean
nothing.
Many topics
A library is a folder of topics that you can search together.
from slim_llm_memory import library
db = library() # ~/.slim-llm-memory/topics
db.topic("nginx").add("docs/nginx/")
db.topic("cooking").add({"pasta.md": "Boil 100 g pasta per person in salted water."})
print(db) # a table of topics
db.ask("how do I enable TLS?") # searches every topic; each hit names its topic
db.ask("...", topics=["nginx"]) # search only some topics
db.archive("cooking") # hide from ask(); db.restore("cooking") brings it back
Once a library holds more than 50,000 passages, ask first picks the topics closest to
the question and searches only those. db.route(question) shows that choice on its own.
Better results
By default ask combines embedding search with keyword search (BM25), so exact names,
numbers and file names still match. You can change that per call:
t.ask(q, mode="dense") # embeddings only
t.ask(q, mode="keyword") # keywords only
t.ask(q, rerank=True) # a cross-encoder re-reads the top candidates (needs [rerank])
t.ask(q, rerank="auto") # the same, but only when the top results are close together
t.answer(q, refuse_below=0.4) # refuse instead of answering when no passage is close enough
t.answer(q, rewrite=True) # the model rewrites the question as a search query first
Measure on your own questions before you tune anything:
from slim_llm_memory import evaluate
# each case: a question, and a word from the right passage (or its document name)
evaluate(t, [("which file is the commit point?", "manifest")], k=5) # hit@1, hit@5, MRR
On eight questions about this repo's docs, the default search had the right passage in
its top 5 for 7 of them, and adding the reranker found all 8. rerank="auto" gave the same
answers as always reranking and was 4.8× faster. Tables and setup:
docs/BENCHMARKS.md.
Links, entities and chat history
t.link("nginx.md", "certbot.md", relation="uses") # needs [graph]
t.related("nginx.md") # similar and linked documents
t.add("docs/", enrich=True) # a local model extracts names and relations (slow, needs [graph])
t.ask(q, entity="Postgres") # only passages that mention Postgres
s = db.session("2026-09-13 refactor") # a conversation you can search later
s.turn("user", "the flaky test was the shared tmp dir")
s.recall("why were tests flaky?")
s.history(5) # the last 5 turns, in order
With [graph] installed, [[wikilinks]] in your documents become links when you add them.
For agents: five tools, an MCP server, and an API sheet
An agent needs five verbs: remember, recall, forget, answer, topics. MemoryTools gives them
JSON in and JSON out, and the two helpers emit the tool definitions in the shape each API expects.
from slim_llm_memory import MemoryTools, anthropic_tools, openai_tools
mem = MemoryTools() # one library, ~/.slim-llm-memory/topics
mem.remember("prefs", "The user prefers dark mode.", name="theme")
mem.recall("what theme does the user like?", k=2) # {"hits": [...], "context": "...", "ms": ...}
mem.dispatch("forget", {"topic": "prefs", "doc": "theme"})
tools = anthropic_tools() # or openai_tools(); pass as tools=...
The same five verbs as an MCP server, for Claude Code, Claude Desktop, Cursor and friends:
pip install "slim-llm-memory[mcp]"
claude mcp add memory -- slim-memory-mcp # Claude Code
slim-memory-mcp --path ~/notes-memory --model qwen2.5:7b-instruct # or configure by hand
Every result type (Result, Hit, Answer, Report, Route, Added) has to_dict().
docs/llms.txt is the whole API
on one page, written for a coding agent: paste it into the context of whatever is writing code
against this package.
Low-level API
topic() is built on Memory, a vector index for when you want to manage ids and
splitting yourself.
from slim_llm_memory import Embedder, Memory
with Memory("./mymemory", Embedder.ollama("nomic-embed-text")) as mem:
mem.upsert([
{"id": "doc1", "text": "how to set up nginx", "meta": {"kind": "note"}},
{"id": "doc2", "text": "buy milk", "meta": {"kind": "shopping"}},
]) # embeds only new or changed text
hits = mem.search("nginx tutorial", k=5, kinds={"note"}, min_score=0.55)
groups = mem.find_duplicates(threshold=0.86)
Memory also has neighbours(id), search_vector(vec), update_text(id, text),
remove(id), stats() and flush(). Embedder.noop() gives offline vectors for tests.
How your data is stored
Each store is a plain folder:
items.vN.jsonl one line per item: id, text, content hash, metadata
vectors.vN.npy the embeddings, one row per item
manifest.json points at the current version
.lock only one process writes to a folder at a time
A save writes new versioned files first and switches manifest.json over only once they
are complete. If the process crashes mid-save, the previous version still loads.
Speed and limits
Search is one numpy scan: 0.3 ms over 1,000 passages, 2.4 ms over 10,000 and 12 ms over 50,000 (median, 8-core CPU). The slow step is Ollama embedding the question, about 1.2 to 1.5 s on that CPU without a GPU.
It is built for one process on one machine. When that stops being enough:
| When | Replace |
|---|---|
| search takes over 100 ms at your size | index.py with faiss (HNSW) |
| a second process needs to write | store.py with SQLite + sqlite-vss |
| more than 1M items | both, with a vector database such as Qdrant or Weaviate |
Notebooks and examples
Start with the four hello notebooks. Each is about ten lines and needs Ollama running.
| Notebook | Shows |
|---|---|
| 00_hello_topic | topic(), .add(), .ask() |
| 01_hello_library | many topics behind one handle, and route() |
| 02_hello_memory | the low-level Memory API |
| 03_hello_answer | a cited answer, and refusal |
Then the five use-case notebooks. Each builds one small application end to end and measures it.
| Notebook | Use case |
|---|---|
| 10_usecase_codebase_qa | ask a repo questions; dense vs keyword vs hybrid on identifiers, cited answer |
| 11_usecase_ticket_triage | route support tickets to an area, draft a reply, escalate with refuse_below |
| 12_usecase_agent_memory | an agent that recalls facts and earlier turns, and survives a restart |
| 13_usecase_semantic_cache | cache expensive calls by meaning; choose the threshold from measured scores |
| 14_usecase_notes_housekeeping | re-index a notes folder cheaply, find duplicates, related() over [[wikilinks]], forget() |
The longer notebooks in notebooks/
measure things. use_cases_demo tests cited answers, paraphrased questions, other
languages and chat memory, and ends with a list of what is missing compared with a full
RAG stack.
Scripts in examples/:
01_minimal.py runs offline, 02_topic_context.py is the full question-to-answer flow,
and 03_routing_bench.py and 04_rerank_bench.py produce the benchmark numbers.
Development
git clone https://github.com/trbck/slim-llm-memory && cd slim-llm-memory
pip install -e ".[test]"
pytest # no network needed; tests for extras you have not installed are skipped
Install the package instead of running with PYTHONPATH=., which hides packaging bugs.
- docs/IMPLEMENTATION.md: the original design plan
- RELEASING.md: how to publish a new version
- CHANGELOG.md: what changed
License
MIT
Metadata
Release files for slim-llm-memory 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| slim_llm_memory-0.2.1.tar.gz | 117.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| slim_llm_memory-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 189.1 kB
Release files / slim_llm_memory-0.2.1.tar.gz
| Download URL | slim_llm_memory-0.2.1.tar.gz |
|---|---|
| Size | 117.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
71fe6a78f4f3096981a93e3490f9398b6ffc8c5a6f29beb74fbd46fbe60063ce
|
|
BLAKE2b-256 checksum How to use checksums |
6031a3e31be8b65bd4ee5f84bc8d92e90615728c12a9f4136d0f781df21b1edd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency logRelease files / slim_llm_memory-0.2.1-py3-none-any.whl
| Download URL | slim_llm_memory-0.2.1-py3-none-any.whl |
|---|---|
| Size | 71.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4640970290c3e4d705767708b4042e92d597f29c5c64e3358c9650bb17bba2a6
|
|
BLAKE2b-256 checksum How to use checksums |
dac010bf39f56ccf7a6abb2eb91871bc8b3a42ef30c0da86fe8cfdbc54089153
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency log