Skip to main content

loci 🧠

English 简体中文 繁體中文 日本語 한국어

CI License Python loci MCP server — quality and maintenance score on Glama ModelScope MCP Square

Two thousand years ago, orators stored their speeches in the rooms of a palace and walked through them to remember. loci does the same for your files.

Loci is the method behind every memory palace: place knowledge in locations, recall it by walking the path.

A queryable "second brain" for the project docs, notes, and chat logs scattered across a dozen directories — and an MCP server so your AI agents can use it too.

Local files → heading-aware chunking → embeddings → hybrid retrieval (vector + BM25) → LLM answer with section-level citations. The index lives entirely on your machine; only embedding/chat calls go out, to any OpenAI-compatible API (Zhipu / DeepSeek / Kimi / OpenAI / …).

The thesis (from studying the 90k-star platforms and the graveyard of dead lightweight tools — see our competitive landscape study): don't build another chat app. Build the memory layer that every chat app can mount. Claude Desktop, Cursor, Cline, or any MCP host becomes this project's UI, for free.

Demo

Real session, indexed against the docs of minimax-h3-turing (paths shortened for display):

$ python main.py search "what the 22G card can and cannot do" -k 3

[1] minimax-h3-turing/docs/en/01-hardware-limits.md > 01 · What a 2080Ti 22G Can and Cannot Do    (similarity 0.562)
[2] minimax-h3-turing/docs/en/02-w4a8-vs-w4a4.md > 02 · Quantization Measured > You Can Try Without 22G  (similarity 0.446)
[3] minimax-h3-turing/docs/en/01-hardware-limits.md > ... > 3. VRAM is just barely enough — manage it  (similarity 0.504)

$ python main.py ask "How should I choose between T8 aggressive mode and the final-render mode, and why?"

Answer:
* Drafts / preview / shot selection: use T8 aggressive mode — a 43% speedup
  (2.7 min/clip), and "a different picture of equal quality" is fine for picking shots.
* Final shots: use final-render mode (no T8). T8 makes the numerical trajectory
  fork, so re-running with the same seed produces a different clip — which breaks
  the reproducibility final outputs need.

[source: docs/en/08-t8-blockcache-4step.md > Practical Advice (4-step Turbo route)]
[source: docs/en/06-faq.md > 12. Cache-style accelerators break "same-seed re-runs"]

Hybrid retrieval means a Chinese query still finds the English doc (and vice versa) — keyword evidence (BM25) catches what embeddings miss, and every citation points at a section, not just a file.

Does hybrid actually help? (mini-eval, 10 bilingual queries)

$ python scripts/eval_retrieval.py scripts/eval_cases.example.jsonl
vector-only: 9/10  →  hybrid: 10/10

Hybrid also fixed the #1 ranking on keyword-ish queries (e.g. "T8 block cache threshold speedup": vector put an FAQ first, hybrid puts the actual T8 writeup first). Run it against your own corpus with your own cases file.

Reranking: two providers

--rerank reorders the fused candidates for precision:

Provider How Cost
llm (default) pointwise 0–3 relevance scoring by your chat model one extra LLM call
local cross-encoder, via pip install 'loci[rerank]' ~30–70 ms for 5 pairs on GPU — offline, free
python main.py search "T8 speedup" --rerank          # provider from config
python main.py search "T8 speedup" --rerank local    # cross-encoder (BAAI/bge-reranker-base)

The local model downloads on first use (~1.1 GB; set HF_ENDPOINT=https://hf-mirror.com if HuggingFace is slow in your region). Measured on a 2080 Ti, bilingual query.

Office documents, PDF tables, chat logs

  • PDFs: with the [pdf] extra, PyMuPDF4LLM extracts pages as markdown — tables come through as pipe rows (plain pypdf text is the fallback)
  • Word: with the [docx] extra, .docx paragraphs and table rows are indexed
  • Chat exports: drop a ChatGPT or Claude conversations.json into any source directory — it becomes one searchable document per conversation, tagged chatlog (search --tag chatlog scopes to chat history)

How it relates to Obsidian / your note app

It doesn't compete — the two layer up. Obsidian (or any editor) is the note-taking frontend; this is the cross-vault search engine: point sources at any directories (Obsidian vaults, project docs, chat exports) and query all of them at once — from your terminal, your scripts, or your AI agent via MCP. Obsidian-native details are understood: frontmatter tags: (filter with search --tag), [[wikilinks]] (walk the graph with links), code blocks are never cut mid-block, and one-line notes stay searchable.

How it works

%%{init: {'theme':'base','themeVariables':{'background':'#000000','primaryColor':'#000000','primaryTextColor':'#00FF41','primaryBorderColor':'#00FF41','lineColor':'#00FF41','secondaryColor':'#001a00','tertiaryColor':'#000000','clusterBkg':'#000000','clusterBorder':'#00FF41','edgeLabelBackground':'#000000','fontSize':'14px','fontFamily':'trebuchet ms, verdana, arial, sans-serif'},'themeCSS':'.nodeLabel { color: #00FF41 !important; } .edgeLabel { background: #000 !important; color: #00FF41 !important; } .cluster-label { color: #00FF41 !important; }'}}%%
flowchart LR
    subgraph sources["📥 Your machine"]
        notes["Obsidian / markdown notes"]
        docs["PDF tables · docx · project docs"]
        chats["ChatGPT / Claude exports"]
        mem["memories/ — agent-written notes"]
        wikidir["wiki/ — consolidated pages"]
    end

    subgraph loci["🧠 loci — local index, nothing leaves the machine"]
        ingest["ingest / watch<br>loaders → chunker → embedder"]
        store[("ChromaDB<br>hybrid index")]
        retrieve["hybrid retrieval<br>vector + BM25 → RRF"]
        mcp["loci-mcp<br>8 tools · resources · prompts"]
    end

    subgraph hosts["🖥️ Your AI hosts"]
        ide["Claude Code · Qoder · Trae<br>Cursor · Cline"]
        desktop["Claude Desktop"]
        term["Terminal<br>search / ask / chat / wiki"]
    end

    api["☁️ OpenAI-compatible API<br>Zhipu / DeepSeek / Kimi / OpenAI<br>or 100% offline via Ollama"]

    sources --> ingest --> store
    mem -. auto-indexed .-> store
    wikidir -. auto-indexed .-> store
    store --> retrieve
    retrieve --> term
    retrieve --> mcp
    mcp <--> ide
    mcp <-.-> desktop
    retrieve -. "embedding + chat calls only" .-> api

The write path in one line: loaders → chunker (heading-aware split) → embedder → store (ChromaDB, persistent) — incremental, deduplicated by content hash.

Install & quick start

Requires Python 3.11+ (uses the stdlib tomllib).

# option A: install as a package (adds `loci` and `loci-mcp` commands)
pip install -e ".[pdf,docx]"   # optional extras: PDF w/ tables, Word documents

# option B: zero-install quickstart
pip install -r requirements.txt

# 1. Configure: copy the example and fill in your values
cp config.example.toml config.toml

# 2. Ingest (incremental — deduplicated by content hash, safe to re-run)
loci ingest            # or: python main.py ingest

# 3. Ask
loci ask "what did I write about X?"

The workflow

%%{init: {'theme':'base','themeVariables':{'background':'#000000','primaryColor':'#000000','primaryTextColor':'#00FF41','primaryBorderColor':'#00FF41','lineColor':'#00FF41','secondaryColor':'#001a00','tertiaryColor':'#000000','clusterBkg':'#000000','clusterBorder':'#00FF41','edgeLabelBackground':'#000000','fontSize':'14px','fontFamily':'trebuchet ms, verdana, arial, sans-serif'},'themeCSS':'.nodeLabel { color: #00FF41 !important; } .edgeLabel { background: #000 !important; color: #00FF41 !important; } .cluster-label { color: #00FF41 !important; }'}}%%
flowchart TD
    A["pip install loci-rag"] --> B["cp config.example.toml config.toml<br>fill API keys + source dirs"]
    B --> C["loci ingest — hybrid index built"]
    C --> D["loci watch — index stays fresh (optional)"]
    C --> E{"What do you need?"}
    E -->|"a synthesized answer"| F["loci ask --verify<br>claim-by-claim audit"]
    E -->|"raw excerpts to quote"| G["loci search --tag memory"]
    E -->|"back-and-forth"| H["loci chat"]
    E -->|"scattered notes on a topic"| I["loci wiki topic<br>consolidate into a wiki page"]
    F --> J["loci remember —<br>keep what you learned"]
    I --> J

Commands

Command What it does
ingest scan sources, index new/changed files, prune deleted ones (--force re-embeds everything)
search "query" retrieval only — ranked excerpts with path > section breadcrumbs
ask "question" retrieval + LLM answer with [source: path > section] citations
ask "…" --verify additionally audit the answer claim-by-claim against the sources (✓ supported, ~ partial, ✗ unsupported)

Filter operators (combine freely, on search and ask):

Flag Filters to
--tag foo files whose frontmatter tags contain foo
--in docs/en files whose path contains the substring
--since 2026-08 / --since 2026-08-15 files modified on/after that date
-e "exact phrase" chunks containing the exact phrase
-k N return N hits (default 5)
links "note" show the [[wikilink]] graph around a note — outbound and inbound
chat multi-turn Q&A loop with conversation memory (/clear, /exit)
watch keep the index current by polling sources (interval in [watch])
ask "…" --rewrite LLM-rewrite the query (keyword + cross-language variants) before retrieval
feedback good|bad rate the chunks used in the last ask; bad-rated chunks sink in future results
wiki --suggest suggest wiki-worthy topics that don't have a page yet
bench cases.jsonl retrieval benchmark: hit@k, vector-only vs hybrid
sync push|pull sync memories/wiki across machines via git ([sync] remote)
serve-http HTTP REST API (search/ask/remember/stats) with Bearer auth
stats what's in the index: chunks per source, models, retrieval settings
doctor health check: config, source dirs, embed/LLM endpoints, store (exit code 1 on failure — CI-friendly)
python mcp_server.py MCP server over stdio (see below)

One memory, every IDE

Because every MCP host mounts the same loci server (same config.toml, same index), memory written from one tool is recalled from every other:

# Claude Code
claude mcp add loci -- loci-mcp
// Cursor / Cline / Qoder / Trae (mcpServers JSON — same shape everywhere)
{ "mcpServers": { "loci": { "command": "loci-mcp" } } }

Then, from any of them: "remember that the staging password rotates on Mondays"brain_remember → later, from a different IDE: "when does the staging password rotate?" → answered, with the memory cited. Memories live as plain markdown in the memories directory (git-friendly, no lock-in) and are tagged memory, so loci search --tag memory scopes to them.

Cross-IDE tip: the default store / memories paths are relative to the directory loci is launched from. If your IDEs start in different project folders, point both at one absolute location in config.toml — e.g. store.path = "~/.loci/store" and memories.path = "~/.loci/memories" — and every IDE shares the exact same memory store.

%%{init: {'theme':'base','themeVariables':{'background':'#000000','primaryColor':'#000000','primaryTextColor':'#00FF41','primaryBorderColor':'#00FF41','lineColor':'#00FF41','actorBkg':'#000000','actorBorder':'#00FF41','actorTextColor':'#00FF41','signalColor':'#00FF41','signalTextColor':'#00FF41','noteBkgColor':'#001a00','noteBorderColor':'#00FF41','activationBkgColor':'#001a00','edgeLabelBackground':'#000000','fontSize':'14px','fontFamily':'trebuchet ms, verdana, arial, sans-serif'},'themeCSS':'.messageText { fill: #00FF41 !important; } .actor { fill: #000 !important; stroke: #00FF41 !important; } text.actor { fill: #00FF41 !important; }'}}%%
sequenceDiagram
    participant CC as Claude Code
    participant L as loci-mcp
    participant S as ChromaDB (local)
    participant T as Trae / Qoder / any IDE
    CC->>L: brain_remember("deploy rotates Mondays")
    L->>S: write memory.md + embed + index
    Note over S: persists across sessions and IDEs
    T->>L: brain_search("password rotation")
    L->>S: hybrid retrieval
    L-->>T: cited answer — the memory is recalled

Mount it in any MCP host

Add to claude_desktop_config.json (Claude Desktop) or your MCP client's config:

{
  "mcpServers": {
    "loci": {
      "command": "python",
      "args": ["/path/to/loci/mcp_server.py"]
    }
  }
}

The server exposes three tools (zero dependencies beyond the core):

Tool Purpose
brain_search(query, k?, tag?, in?) ranked excerpts with breadcrumbs
brain_ask(question, verify?) grounded answer with citations; verify=true adds a claim-by-claim audit
brain_links(note) outbound/inbound [[wikilink]] graph around a note
brain_stats() index overview (chunks per source)
brain_remember(text, title?, tags?) write a memory — durable, shared across sessions and IDEs
brain_forget(query) soft-delete matching memories (they go to a .trash folder)
brain_wiki(topic) memory consolidation — distill the index into a curated wiki page about a topic
brain_ingest(force?) incremental re-index

Beyond tools, the server speaks the full protocol:

  • Resourcesresources/list exposes brain://stats plus one brain://note/… resource per indexed file (raw markdown via resources/read)
  • Prompts — three ready-made templates: brain-briefing, study-plan, contradiction-check; hosts render them with your topic pre-filled

Fully offline with Ollama

The index is local by design — and the embedding/chat calls can be too. Any OpenAI-compatible server works; Ollama is verified end-to-end:

[llm]
base_url = "http://localhost:11434/v1"
api_key = "ollama"          # any non-empty placeholder
model = "qwen2.5:0.5b"

[embed]
base_url = "http://localhost:11434/v1"
api_key = "ollama"
model = "all-minilm"

With this config, ingest / search / ask make zero cloud calls. Swap in a bigger local chat model for better answers — the pipeline is model-agnostic.

Configuration

Key Meaning
[llm] base_url / api_key / model — any OpenAI-compatible endpoint
[embed] same; the model must be an embedding model (e.g. embedding-3)
[[sources]] document directories, scanned recursively for .md / .txt (plus .pdf/.docx with the matching extras)
[[sources]] chunk_size / chunk_overlap optional per-directory chunking override — wins over the global [chunk] block
[chunk] chunking params (default 800 chars / 100 overlap)
[top_k] number of hits per search (default 5)
[retrieval] hybrid (vector+BM25 fusion, default on), rrf_k, rerank (LLM reranking, default off)
[watch] poll interval seconds

API keys can also come from the environment variables BRAIN_LLM_API_KEY / BRAIN_EMBED_API_KEY (these override the config file).

Design decisions

  • ~300 lines of core, no LangChain — every stage is readable, hackable, and learnable. The whole engine fits in one sitting.
  • MCP-first — the agent ecosystem is the UI layer. No web app to maintain.
  • Hybrid retrieval on by default — vector search fused with a native ~60-line BM25 (CJK-aware tokenizer) via Reciprocal Rank Fusion.
  • Citations always, with breadcrumbspath > section, so claims are verifiable at a glance.
  • Robust, inspectable indexing — defensive loaders (skip what can't be parsed, never hang), content-hash incrementality, real pruning, stats and doctor so the index is never a black box.
  • Tiny notes stay searchable — no minimum-chunk filter; a one-line note is still indexed (a lesson from watching other tools drop or choke on them).
  • Keys never in codeconfig.toml (gitignored) or env vars.

Where it sits

loci AnythingLLM (65k★) Khoj (37k★) RAGFlow (90k★)
Positioning personal retrieval backend + MCP all-in-one chat platform self-hosted AI assistant enterprise RAG engine
Footprint 2 runtime deps, no Docker desktop app / Docker Django server + workers Docker, DeepDoc models
UI your terminal & your agents built-in web/desktop web + Obsidian/Emacs web
MCP server ✅ native consumer
Hackable core ✅ ~300 lines
Multi-user by design, no

(Full data and reasoning: competitive landscape study.)

Roadmap

See docs/roadmap.md — reranking, GraphRAG experiments, more loaders.

License

MIT

Release files for loci-rag 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loci-rag 0.5.0
File Size Uploaded
loci_rag-0.5.0.tar.gz 70.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for loci-rag 0.5.0
File Interpreter ABI Platform
loci_rag-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 118.4 kB

Release files / loci_rag-0.5.0.tar.gz

Download URL loci_rag-0.5.0.tar.gz
Size 70.6 kB
Tags Source
SHA-256 checksum
How to use checksums
1f72dbe76016739d6cb2d2535e9fc009e02fc45fe9cd6c8786038ae88cb5eba8
BLAKE2b-256 checksum
How to use checksums
48be48de0c9eac36ed0cecadb64f7abb833bd769d4ae8a4078d865cbf34c5f5d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release files / loci_rag-0.5.0-py3-none-any.whl

Download URL loci_rag-0.5.0-py3-none-any.whl
Size 47.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a1460869da30b6c882c13703e93924fda02eb82457d01008037843bc05c1bf9b
BLAKE2b-256 checksum
How to use checksums
564172b1660ce87f1d17e6d4eee6f3a7ba813b02b878ddb45b5058abc2d0355a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release history Release notifications | RSS feed

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page