Skip to main content

slim-llm-memory

Local memory and retrieval for LLM apps. Put your notes or docs into a store, ask a question, and get back the passages to paste into a prompt, or a cited answer from a local model.

It needs only numpy and httpx. Embeddings come from Ollama on your own machine. It is meant for personal projects and research code, up to about 50,000 passages per store, where one numpy array is fast enough and a vector database is extra work.

Install

pip install slim-llm-memory
ollama pull nomic-embed-text      # embeddings
ollama pull llama3.2:3b           # only needed for answer()

Python 3.10 or newer. Optional extras:

Extra Adds
slim-llm-memory[rerank] cross-encoder reranking (sentence-transformers)
slim-llm-memory[graph] links between documents (NetworkX)
slim-llm-memory[mcp] the slim-memory-mcp server for MCP clients
slim-llm-memory[obsidian] experimental Obsidian vault ingest; Python API only, no command yet

The [gemini] and [anthropic] extras are placeholders. No code uses them yet.

Quickstart

from slim_llm_memory import topic

t = topic("nginx")                  # a store in ~/.slim-llm-memory/topics/nginx
t.add("docs/nginx/")                # every .md, .txt and .rst file in the folder
r = t.ask("how do I enable TLS?")   # the best matching passages
print(r)                            # hits, scores and timings
print(r.context)                    # numbered passages, ready to paste into a prompt

print(t.answer("how do I enable TLS?"))   # a local model answers from those passages and cites them

add also takes a single file, raw text, or a {name: text} dict. Running it again only re-embeds what changed, and t.forget("old.md") removes a document. To try the API without Ollama, pass embedder="noop": every call works, but the similarity scores mean nothing.

Many topics

A library is a folder of topics that you can search together.

from slim_llm_memory import library

db = library()                          # ~/.slim-llm-memory/topics
db.topic("nginx").add("docs/nginx/")
db.topic("cooking").add({"pasta.md": "Boil 100 g pasta per person in salted water."})

print(db)                               # a table of topics
db.ask("how do I enable TLS?")          # searches every topic; each hit names its topic
db.ask("...", topics=["nginx"])         # search only some topics
db.archive("cooking")                   # hide from ask(); db.restore("cooking") brings it back

Once a library holds more than 50,000 passages, ask first picks the topics closest to the question and searches only those. db.route(question) shows that choice on its own.

Better results

By default ask combines embedding search with keyword search (BM25), so exact names, numbers and file names still match. You can change that per call:

t.ask(q, mode="dense")          # embeddings only
t.ask(q, mode="keyword")        # keywords only
t.ask(q, rerank=True)           # a cross-encoder re-reads the top candidates (needs [rerank])
t.ask(q, rerank="auto")         # the same, but only when the top results are close together
t.answer(q, refuse_below=0.4)   # refuse instead of answering when no passage is close enough
t.answer(q, rewrite=True)       # the model rewrites the question as a search query first

Measure on your own questions before you tune anything:

from slim_llm_memory import evaluate

# each case: a question, and a word from the right passage (or its document name)
evaluate(t, [("which file is the commit point?", "manifest")], k=5)   # hit@1, hit@5, MRR

On eight questions about this repo's docs, the default search had the right passage in its top 5 for 7 of them, and adding the reranker found all 8. rerank="auto" gave the same answers as always reranking and was 4.8× faster. Tables and setup: docs/BENCHMARKS.md.

t.link("nginx.md", "certbot.md", relation="uses")   # needs [graph]
t.related("nginx.md")                               # similar and linked documents
t.add("docs/", enrich=True)                         # a local model extracts names and relations (slow, needs [graph])
t.ask(q, entity="Postgres")                         # only passages that mention Postgres

s = db.session("2026-09-13 refactor")               # a conversation you can search later
s.turn("user", "the flaky test was the shared tmp dir")
s.recall("why were tests flaky?")
s.history(5)                                        # the last 5 turns, in order

With [graph] installed, [[wikilinks]] in your documents become links when you add them.

For agents: five tools, an MCP server, and an API sheet

An agent needs five verbs: remember, recall, forget, answer, topics. MemoryTools gives them JSON in and JSON out, and the two helpers emit the tool definitions in the shape each API expects.

from slim_llm_memory import MemoryTools, anthropic_tools, openai_tools

mem = MemoryTools()                                  # one library, ~/.slim-llm-memory/topics
mem.remember("prefs", "The user prefers dark mode.", name="theme")
mem.recall("what theme does the user like?", k=2)   # {"hits": [...], "context": "...", "ms": ...}
mem.dispatch("forget", {"topic": "prefs", "doc": "theme"})

tools = anthropic_tools()                            # or openai_tools(); pass as tools=...

The same five verbs as an MCP server, for Claude Code, Claude Desktop, Cursor and friends:

pip install "slim-llm-memory[mcp]"
claude mcp add memory -- slim-memory-mcp            # Claude Code
slim-memory-mcp --path ~/notes-memory --model qwen2.5:7b-instruct   # or configure by hand

Every result type (Result, Hit, Answer, Report, Route, Added) has to_dict(). docs/llms.txt is the whole API on one page, written for a coding agent: paste it into the context of whatever is writing code against this package.

Low-level API

topic() is built on Memory, a vector index for when you want to manage ids and splitting yourself.

from slim_llm_memory import Embedder, Memory

with Memory("./mymemory", Embedder.ollama("nomic-embed-text")) as mem:
    mem.upsert([
        {"id": "doc1", "text": "how to set up nginx", "meta": {"kind": "note"}},
        {"id": "doc2", "text": "buy milk",            "meta": {"kind": "shopping"}},
    ])                                          # embeds only new or changed text
    hits = mem.search("nginx tutorial", k=5, kinds={"note"}, min_score=0.55)
    groups = mem.find_duplicates(threshold=0.86)

Memory also has neighbours(id), search_vector(vec), update_text(id, text), remove(id), stats() and flush(). Embedder.noop() gives offline vectors for tests.

How your data is stored

Each store is a plain folder:

items.vN.jsonl    one line per item: id, text, content hash, metadata
vectors.vN.npy    the embeddings, one row per item
manifest.json     points at the current version
.lock             only one process writes to a folder at a time

A save writes new versioned files first and switches manifest.json over only once they are complete. If the process crashes mid-save, the previous version still loads.

Speed and limits

Search is one numpy scan: 0.3 ms over 1,000 passages, 2.4 ms over 10,000 and 12 ms over 50,000 (median, 8-core CPU). The slow step is Ollama embedding the question, about 1.2 to 1.5 s on that CPU without a GPU.

It is built for one process on one machine. When that stops being enough:

When Replace
search takes over 100 ms at your size index.py with faiss (HNSW)
a second process needs to write store.py with SQLite + sqlite-vss
more than 1M items both, with a vector database such as Qdrant or Weaviate

Notebooks and examples

Start with the four hello notebooks. Each is about ten lines and needs Ollama running.

Notebook Shows
00_hello_topic topic(), .add(), .ask()
01_hello_library many topics behind one handle, and route()
02_hello_memory the low-level Memory API
03_hello_answer a cited answer, and refusal

Then the five use-case notebooks. Each builds one small application end to end and measures it.

Notebook Use case
10_usecase_codebase_qa ask a repo questions; dense vs keyword vs hybrid on identifiers, cited answer
11_usecase_ticket_triage route support tickets to an area, draft a reply, escalate with refuse_below
12_usecase_agent_memory an agent that recalls facts and earlier turns, and survives a restart
13_usecase_semantic_cache cache expensive calls by meaning; choose the threshold from measured scores
14_usecase_notes_housekeeping re-index a notes folder cheaply, find duplicates, related() over [[wikilinks]], forget()

The longer notebooks in notebooks/ measure things. use_cases_demo tests cited answers, paraphrased questions, other languages and chat memory, and ends with a list of what is missing compared with a full RAG stack.

Scripts in examples/: 01_minimal.py runs offline, 02_topic_context.py is the full question-to-answer flow, and 03_routing_bench.py and 04_rerank_bench.py produce the benchmark numbers.

Development

git clone https://github.com/trbck/slim-llm-memory && cd slim-llm-memory
pip install -e ".[test]"
pytest          # no network needed; tests for extras you have not installed are skipped

Install the package instead of running with PYTHONPATH=., which hides packaging bugs.

License

MIT

Metadata

Release files for slim-llm-memory 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for slim-llm-memory 0.2.1
File Size Uploaded
slim_llm_memory-0.2.1.tar.gz 117.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for slim-llm-memory 0.2.1
File Interpreter ABI Platform
slim_llm_memory-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 189.1 kB

Release files / slim_llm_memory-0.2.1.tar.gz

Download URL slim_llm_memory-0.2.1.tar.gz
Size 117.5 kB
Tags Source
SHA-256 checksum
How to use checksums
71fe6a78f4f3096981a93e3490f9398b6ffc8c5a6f29beb74fbd46fbe60063ce
BLAKE2b-256 checksum
How to use checksums
6031a3e31be8b65bd4ee5f84bc8d92e90615728c12a9f4136d0f781df21b1edd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / slim_llm_memory-0.2.1-py3-none-any.whl

Download URL slim_llm_memory-0.2.1-py3-none-any.whl
Size 71.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4640970290c3e4d705767708b4042e92d597f29c5c64e3358c9650bb17bba2a6
BLAKE2b-256 checksum
How to use checksums
dac010bf39f56ccf7a6abb2eb91871bc8b3a42ef30c0da86fe8cfdbc54089153
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page