Skip to main content

slim-llm-memory

Local memory and retrieval for LLM apps. Put your notes or docs into a store, ask a question, and get back the passages to paste into a prompt, or a cited answer from a local model.

It needs only numpy and httpx. Embeddings come from Ollama on your own machine. It is meant for personal projects and research code, up to about 50,000 passages per store, where one numpy array is fast enough and a vector database is extra work.

Install

pip install slim-llm-memory
ollama pull nomic-embed-text      # embeddings
ollama pull llama3.2:3b           # only needed for answer()

Python 3.10 or newer. Optional extras:

Extra Adds
slim-llm-memory[rerank] cross-encoder reranking (sentence-transformers)
slim-llm-memory[graph] links between documents (NetworkX)
slim-llm-memory[obsidian] experimental Obsidian vault ingest; Python API only, no command yet

The [gemini] and [anthropic] extras are placeholders. No code uses them yet.

Quickstart

from slim_llm_memory import topic

t = topic("nginx")                  # a store in ~/.slim-llm-memory/topics/nginx
t.add("docs/nginx/")                # every .md, .txt and .rst file in the folder
r = t.ask("how do I enable TLS?")   # the best matching passages
print(r)                            # hits, scores and timings
print(r.context)                    # numbered passages, ready to paste into a prompt

print(t.answer("how do I enable TLS?"))   # a local model answers from those passages and cites them

add also takes a single file, raw text, or a {name: text} dict. Running it again only re-embeds what changed, and t.forget("old.md") removes a document. To try the API without Ollama, pass embedder="noop": every call works, but the similarity scores mean nothing.

Many topics

A library is a folder of topics that you can search together.

from slim_llm_memory import library

db = library()                          # ~/.slim-llm-memory/topics
db.topic("nginx").add("docs/nginx/")
db.topic("cooking").add({"pasta.md": "Boil 100 g pasta per person in salted water."})

print(db)                               # a table of topics
db.ask("how do I enable TLS?")          # searches every topic; each hit names its topic
db.ask("...", topics=["nginx"])         # search only some topics
db.archive("cooking")                   # hide from ask(); db.restore("cooking") brings it back

Once a library holds more than 50,000 passages, ask first picks the topics closest to the question and searches only those. db.route(question) shows that choice on its own.

Better results

By default ask combines embedding search with keyword search (BM25), so exact names, numbers and file names still match. You can change that per call:

t.ask(q, mode="dense")          # embeddings only
t.ask(q, mode="keyword")        # keywords only
t.ask(q, rerank=True)           # a cross-encoder re-reads the top candidates (needs [rerank])
t.ask(q, rerank="auto")         # the same, but only when the top results are close together
t.answer(q, refuse_below=0.4)   # refuse instead of answering when no passage is close enough
t.answer(q, rewrite=True)       # the model rewrites the question as a search query first

Measure on your own questions before you tune anything:

from slim_llm_memory import evaluate

# each case: a question, and a word from the right passage (or its document name)
evaluate(t, [("which file is the commit point?", "manifest")], k=5)   # hit@1, hit@5, MRR

On eight questions about this repo's docs, the default search had the right passage in its top 5 for 7 of them, and adding the reranker found all 8. rerank="auto" gave the same answers as always reranking and was 4.8× faster. Tables and setup: docs/BENCHMARKS.md.

Links, entities and chat history

t.link("nginx.md", "certbot.md", relation="uses")   # needs [graph]
t.related("nginx.md")                               # similar and linked documents
t.add("docs/", enrich=True)                         # a local model extracts names and relations (slow, needs [graph])
t.ask(q, entity="Postgres")                         # only passages that mention Postgres

s = db.session("2026-09-13 refactor")               # a conversation you can search later
s.turn("user", "the flaky test was the shared tmp dir")
s.recall("why were tests flaky?")
s.history(5)                                        # the last 5 turns, in order

With [graph] installed, [[wikilinks]] in your documents become links when you add them.

Low-level API

topic() is built on Memory, a vector index for when you want to manage ids and splitting yourself.

from slim_llm_memory import Embedder, Memory

with Memory("./mymemory", Embedder.ollama("nomic-embed-text")) as mem:
    mem.upsert([
        {"id": "doc1", "text": "how to set up nginx", "meta": {"kind": "note"}},
        {"id": "doc2", "text": "buy milk",            "meta": {"kind": "shopping"}},
    ])                                          # embeds only new or changed text
    hits = mem.search("nginx tutorial", k=5, kinds={"note"}, min_score=0.55)
    groups = mem.find_duplicates(threshold=0.86)

Memory also has neighbours(id), search_vector(vec), update_text(id, text), remove(id), stats() and flush(). Embedder.noop() gives offline vectors for tests.

How your data is stored

Each store is a plain folder:

items.vN.jsonl    one line per item: id, text, content hash, metadata
vectors.vN.npy    the embeddings, one row per item
manifest.json     points at the current version
.lock             only one process writes to a folder at a time

A save writes new versioned files first and switches manifest.json over only once they are complete. If the process crashes mid-save, the previous version still loads.

Speed and limits

Search is one numpy scan: 0.3 ms over 1,000 passages, 2.4 ms over 10,000 and 12 ms over 50,000 (median, 8-core CPU). The slow step is Ollama embedding the question, about 1.2 to 1.5 s on that CPU without a GPU.

It is built for one process on one machine. When that stops being enough:

When Replace
search takes over 100 ms at your size index.py with faiss (HNSW)
a second process needs to write store.py with SQLite + sqlite-vss
more than 1M items both, with a vector database such as Qdrant or Weaviate

Notebooks and examples

Start with the four hello notebooks. Each is about ten lines and needs Ollama running.

Notebook Shows
00_hello_topic topic(), .add(), .ask()
01_hello_library many topics behind one handle, and route()
02_hello_memory the low-level Memory API
03_hello_answer a cited answer, and refusal

The longer notebooks in notebooks/ measure things. use_cases_demo tests cited answers, paraphrased questions, other languages and chat memory, and ends with a list of what is missing compared with a full RAG stack.

Scripts in examples/: 01_minimal.py runs offline, 02_topic_context.py is the full question-to-answer flow, and 03_routing_bench.py and 04_rerank_bench.py produce the benchmark numbers.

Development

git clone https://github.com/trbck/slim-llm-memory && cd slim-llm-memory
pip install -e ".[test]"
pytest          # no network needed; tests for extras you have not installed are skipped

Install the package instead of running with PYTHONPATH=., which hides packaging bugs.

License

MIT

Metadata

Release files for slim-llm-memory 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for slim-llm-memory 0.1.1
File Size Uploaded
slim_llm_memory-0.1.1.tar.gz 102.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for slim-llm-memory 0.1.1
File Interpreter ABI Platform
slim_llm_memory-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 165.1 kB

Release files / slim_llm_memory-0.1.1.tar.gz

Download URL slim_llm_memory-0.1.1.tar.gz
Size 102.1 kB
Tags Source
SHA-256 checksum
How to use checksums
14740aa3ff0c38c777f733a886767af02d96d86bec4655e84dcda3cd6324b62f
BLAKE2b-256 checksum
How to use checksums
272bb37343c38a18db12053256a839a96ba238913f593b4f67ac48cab7705179
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release files / slim_llm_memory-0.1.1-py3-none-any.whl

Download URL slim_llm_memory-0.1.1-py3-none-any.whl
Size 63.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d77ead6474df8ded9bd316e4d282a68f55b9af5d35a3143bed5ed430fdf75fb9
BLAKE2b-256 checksum
How to use checksums
be505c49a0c7c67c245b1336f021b70c568a8bcc5f82c58ce5e72692205109f6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.1

2 release files

0.2.0

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page