Cuimhin
Is cuimhin liom — "I remember."
Cuimhin is a hybrid-retrieval memory layer for AI agents and applications. It stores text chunks with metadata, indexes them twice — dense vectors in Qdrant and keywords in SQLite FTS5 — and fuses the two rankings with Reciprocal Rank Fusion, so proper nouns and project names are found as reliably as paraphrases. It exposes the result as a Python library, a REST API, and an MCP server that plugs straight into Claude, Cursor, or any MCP-capable agent.
It is the open-source core extracted from Mnemos, a private memory system in daily production use since early 2026 at ~130k chunks. The retrieval design, the schema discipline and the failure modes documented here were all learned there — including the one that returned HTTP 200 for two days while the keyword leg was silently dead (docs/SCOPE-SPLIT.md §4.1).
Status
0.1.0. The library, CLI, REST and MCP are here and tested; the cross-encoder reranker and the retrieval eval are 0.2.0. Each milestone's acceptance was run against a real Qdrant and written down as it happened — docs/M0-RESULT.md and docs/M1-RESULT.md, including the two runs that failed and what was wrong.
What it does
- Ingest — chunk text, embed it, write vectors to Qdrant and keywords to FTS5 in one idempotent call. Re-ingesting the same document is a no-op.
- Query — hybrid search with metadata filters, date windows, and source scoping; returns ranked hits with score, source, date, and a reconstruction handle.
- Serve — FastAPI REST (
/query,/ingest,/document/{id},/export) and FastMCP tools (query_memory,ingest_document,get_document,list_sources,get_stats) from one process. - Export — cursor-paged JSONL export so Cuimhin can be a source of record, not a dead end.
What it deliberately does not do
Cuimhin is single-tenant and connector-free by design. Multi-tenant isolation, live connectors (email, drives, chat exports), dedup pipelines, temporal encoding, and enrichment layers are out of scope for this repo. See docs/SCOPE-SPLIT.md for the boundary and why it sits where it does.
Quick start
pip install 'cuimhin[sentence-transformers]' # the library and the local embedder
docker run -p 6333:6333 qdrant/qdrant # or the Linux binary from Qdrant's releases
Two install sizes, and the difference is one dependency:
| site-packages | what you get | |
|---|---|---|
pip install cuimhin |
195 MB | library, CLI, REST, MCP — you supply an Embedder |
pip install 'cuimhin[sentence-transformers]' |
1.5 GB | the above plus the default local model |
The extra brings torch (769 MB) and transformers (115 MB), and on a machine without CUDA a plain resolve pulls another ~2 GB of nvidia-* wheels it can never use — install the CPU build first if that matters:
pip install --index-url https://download.pytorch.org/whl/cpu torch
pip install 'cuimhin[sentence-transformers]'
If you embed through an API, or already have a model loaded, skip the extra: Memory(embedder=...) takes anything with eight methods (cuimhin/embed.py), and without it Memory() fails at startup with the command to run rather than at your first query. The model itself (all-MiniLM-L6-v2, ~90 MB) downloads once from the Hugging Face hub.
cuimhin ingest ./notes/ # any folder of .md/.txt; re-ingest is a no-op
cuimhin query "what did I decide about the auth redesign"
cuimhin query "Ballyvaughan" --leg sparse # ablate: keyword leg only
cuimhin export > notes.jsonl # the corpus, replayable, in (ingested_at, id) order
cuimhin import notes.jsonl # restore it exactly — ids, stamps, order; idempotent
Serve REST and MCP from one process behind one key:
export CUIMHIN_API_KEY=$(openssl rand -hex 32)
cuimhin serve # REST on 127.0.0.1:8000, MCP at /mcp
curl -s localhost:8000/health # no key needed
curl -s -XPOST localhost:8000/query -H "X-API-Key: $CUIMHIN_API_KEY" \
-H "Content-Type: application/json" -d '{"q": "auth redesign", "top_k": 5}'
POST /query, POST /ingest, GET/DELETE /document/{id}, GET /export?after=&source=&limit=, GET /stats, GET /health. Every error is {"error": ..., "detail": ...}; a drifted keyword index is a 500 that says run cuimhin rebuild-fts, never an empty 200. Rate limit is CUIMHIN_RATE_LIMIT per client per minute (default 60, 0 off); /health and /export are exempt.
MCP
Five tools: query_memory, ingest_document, get_document, list_sources, get_stats. Tools return data, not prose, and raise on error rather than returning nothing.
Locally (Claude Code, Claude Desktop, Cursor): stdio, no key — the client spawns the process under your own account.
claude mcp add cuimhin -e CUIMHIN_QDRANT_URL=http://localhost:6333 -- cuimhin serve --stdio
{
"mcpServers": {
"cuimhin": {
"command": "cuimhin",
"args": ["serve", "--stdio"],
"env": { "CUIMHIN_QDRANT_URL": "http://localhost:6333" }
}
}
}
Hosted: streamable HTTP at https://your-host/mcp with Authorization: Bearer $CUIMHIN_API_KEY — the same key as REST, the same rate limit. Set CUIMHIN_ALLOWED_HOST=your-host so the SDK's Host-header check accepts it (localhost and 127.0.0.1 are always accepted). OAuth is not in this layer.
Python
from cuimhin import Memory
mem = Memory() # reads CUIMHIN_* env, or pass Config()
mem.ingest("The auth redesign ships in September.", source="notes", date="2026-08-25")
for hit in mem.query("auth redesign", top_k=5):
print(hit.score, hit.source, hit.text[:80])
page = mem.export(limit=500) # records, cursor, has_more — replay with ingest_chunk(ingested_at=...)
Configuration
All CUIMHIN_*; nothing else is read. CUIMHIN_QDRANT_URL, CUIMHIN_QDRANT_API_KEY, CUIMHIN_COLLECTION, CUIMHIN_DB_PATH, CUIMHIN_EMBED_MODEL, CUIMHIN_CHUNK_TOKENS, CUIMHIN_CHUNK_OVERLAP, CUIMHIN_API_KEY, CUIMHIN_RATE_LIMIT, CUIMHIN_ALLOWED_HOST, CUIMHIN_RERANKER, CUIMHIN_LOG_LEVEL — defaults and meanings in docs/ARCHITECTURE.md §6.
Why hybrid, and why RRF
Dense retrieval alone misses the things people actually ask agents about: the name of a project, a person, a file, a variable. FTS5 catches those exactly. RRF merges the two rankings without a tunable weight, which means there is nothing to mis-tune when the corpus changes. A cross-encoder reranker is available as an opt-in hook; on the corpus this was developed against it regressed quality and cost ~90× latency on CPU, so it is off by default and stays that way until you measure otherwise.
Documents
| File | What it is |
|---|---|
docs/PRD-cuimhin.md |
Product requirements: positioning, scope, API surface, non-goals |
docs/SCOPE-SPLIT.md |
Open / licensed / private boundary, and the extraction map from the parent system |
docs/ARCHITECTURE.md |
Module layout, the dual-index schema contract, data flow |
docs/BRIEF-m0-skeleton.md |
M0 build brief — the walking skeleton; M0-ISSUES.md and M0-RESULT.md are its findings and acceptance |
docs/BRIEF-m1-serve.md |
M1 build brief — REST, MCP, auth, export; M1-ISSUES.md and M1-RESULT.md likewise |
docs/BRIEF-m3-release.md |
M3 build brief — packaging and the 0.1.0 release |
docs/BRIEF-m2-eval.md |
M2 build brief — the reranker and the retrieval ablation, for 0.2.0 |
CHANGELOG.md |
What changed in each release, and what is deliberately absent |
CLAUDE.md |
Working rules for Claude Code in this repo |
Licence
Apache-2.0. Copyright 2026 Todd McCaffrey.
Cuimhin is the shell of Mnemos, which stays private. The boundary and the reasoning behind it are in docs/SCOPE-SPLIT.md — including the four conditions a feature has to meet to belong here rather than there.
Metadata
Release files for cuimhin 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cuimhin-0.1.0.tar.gz | 161.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cuimhin-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 229.0 kB
Release files / cuimhin-0.1.0.tar.gz
| Download URL | cuimhin-0.1.0.tar.gz |
|---|---|
| Size | 161.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4530595958b27b7221d271d10579b67cc7bf86cb5950ecb3e50b2b508a41ac9f
|
|
BLAKE2b-256 checksum How to use checksums |
ba425c1d31575f98f9e1d58c677124ac8494caccc5f018c4e0808b2d467ea1b2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / cuimhin-0.1.0-py3-none-any.whl
| Download URL | cuimhin-0.1.0-py3-none-any.whl |
|---|---|
| Size | 67.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
99c8641c50774c92afd67e530915b061ad5549d691f119a413852ee2618740f2
|
|
BLAKE2b-256 checksum How to use checksums |
5ad8e6731fbe4bc4ca38068dddda0540c46b119f4eb77f54053a4235601e4af0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log