Skip to main content

mdsearch

License: MIT

Semantic search over a folder of markdown files — built for Obsidian vaults, but works on any directory of .md notes.

Instead of grepping for exact words, mdsearch finds notes that are about what you're asking, even if they never use your exact phrasing. It chunks your notes, embeds each chunk with a local sentence-transformers model, stores the vectors in a FAISS index, and searches by cosine similarity.

This was originally built so Claude could search an Obsidian vault intelligently via MCP — see docs/MCP_SETUP.md for wiring it up with Claude Desktop. It works standalone as a CLI too.

Why

Filename and grep search only find notes that contain your literal words. If you wrote "used a dictionary to cache API responses" six months ago and search for "memoization", grep finds nothing — mdsearch finds it, because the embedding model understands the two phrases are related.

Install

pip install mdsearch-cli

The package is named mdsearch-cli on PyPI (mdsearch was already taken by an unrelated project), but it installs the same mdsearch command and mdsearch.* Python package.

To install from source instead:

git clone https://github.com/Nakshh/mdsearch.git
cd mdsearch
pip install -e .

Requires Python 3.11+. The first run downloads a small embedding model (all-MiniLM-L6-v2, ~90MB) from Hugging Face and caches it locally.

Usage

Build (or incrementally update) the index for a vault:

$ mdsearch index ~/ObsidianVault
Indexed. added=142 updated=0 removed=0 unchanged=0 chunks=891 (38.42s)

Search it:

$ mdsearch search "notes on memoization" --vault-path ~/ObsidianVault --top-k 3
┏━━━━━━━━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ File                Chunk  Score  Snippet                                 ┃
┡━━━━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ caching-notes.md        2  0.612  # Caching strategies Used a dict as a  │                                   simple memo table to avoid recomputing │
│                                   expensive recursive calls...           │
│ algorithms.md           0  0.401  # Dynamic programming DP is really     │                                   just recursion plus caching...         │
│ interview-prep.md       5  0.318  # Common gotchas Forgetting to cache   │                                   results leads to exponential blowup... │
└────────────────────┴───────┴───────┴─────────────────────────────────────────┘

Re-running index on an unchanged vault is a near-instant no-op — every file's content hash is checked against a manifest, and only changed or new files get re-embedded:

$ mdsearch index ~/ObsidianVault
Indexed. added=0 updated=0 removed=0 unchanged=142 chunks=891 (0.01s)

Commands

Command Description
mdsearch index <vault_path> Build or incrementally update the index for a vault.
mdsearch search <query> Search an already-indexed vault.
mdsearch --version Print the installed version.

index options

Flag Description
--force, -f Re-embed every file regardless of content hash.

search options

Flag Description
--top-k, -k Number of results to return (default 5, must be ≥ 1).
--vault-path Vault to search; defaults to the current directory. Must already be indexed.

The index lives at <vault_path>/.mdsearch — inside the vault itself, so it travels with the vault and doesn't depend on where you run the command from.

Architecture

markdown files -> chunk -> embed -> FAISS index -> search
  • Chunking: each .md file is split on heading boundaries, with a hard fallback split for oversized sections, so chunks stay small and topical.
  • Embedding: chunks are encoded with a local sentence-transformers model (default all-MiniLM-L6-v2) — no API calls, no data leaves your machine.
  • Indexing: vectors are L2-normalized and stored in a FAISS IndexFlatIP, so inner product search is equivalent to cosine similarity. It's an exact, brute-force index rather than an approximate one (HNSW, IVF, ...) — at the scale of a personal vault (thousands, not millions, of chunks) exact search is fast enough and there's no accuracy tradeoff to make.
  • Incremental updates: a per-file sha256 manifest is stored alongside the index. On re-index, unchanged files reuse their cached vectors instead of being re-embedded, so re-running index on a mostly-unchanged vault is near-instant.
  • Metadata: chunk text and file/chunk IDs are stored as human-readable JSONL, not pickled, so the index directory is easy to inspect or diff.

MCP server

mdsearch also ships an MCP server so an AI assistant (e.g. Claude Desktop) can search your vault directly. See docs/MCP_SETUP.md for setup instructions.

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mdsearch_cli-0.1.0.tar.gz (15.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mdsearch_cli-0.1.0-py3-none-any.whl (13.4 kB view details)

Uploaded Python 3

File details

Details for the file mdsearch_cli-0.1.0.tar.gz.

File metadata

  • Download URL: mdsearch_cli-0.1.0.tar.gz
  • Upload date:
  • Size: 15.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mdsearch_cli-0.1.0.tar.gz
Algorithm Hash digest
SHA256 f3d5144b6f4d1759803d817a43b691e0b436a632db2f9778dc101e3bf200692e
MD5 c5a91476532a894e115f949b20694a98
BLAKE2b-256 8651aec82fa516847f1a4ec90138f00ce249905d32d47d06424efc44ef8a32d3

See more details on using hashes here.

Provenance

The following attestation bundles were made for mdsearch_cli-0.1.0.tar.gz:

Publisher: publish.yml on Nakshh/mdsearch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mdsearch_cli-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: mdsearch_cli-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 13.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mdsearch_cli-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8757b02af6f0d5387c748c57e89e5d5aff878138b06bc2a1aef8682de4d8d129
MD5 e2353cb9bfa05efc398c40cd75a44584
BLAKE2b-256 537a5081db453bf28175eead29bc971900e1a148dc985e12e9cb641603ef5b37

See more details on using hashes here.

Provenance

The following attestation bundles were made for mdsearch_cli-0.1.0-py3-none-any.whl:

Publisher: publish.yml on Nakshh/mdsearch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page