Skip to main content

Doc Index MCP

What is This For?

A local-first semantic search server for your documents. Index PDFs, Word docs, PowerPoints, Excel files, and text/markdown, then search them using natural language via the Model Context Protocol (MCP).

  • Semantic search - Find relevant content using natural language queries
  • Boundary-aware chunking - Respects document structure (chapters, sections, headers)
  • Table extraction - Extract tables from documents as CSV
  • Fully local - No external APIs, no cloud services, no Docker containers, no PyTorch
  • Lightweight - ONNX-based embeddings (~50MB vs ~2GB for PyTorch)

Quick Start

1. Add to your MCP config

Requires uv. If you don't have uv, see Alternative Installation below.

Add to .mcp.json in your project root (for Claude Code) or your Claude Desktop config:

{
  "mcpServers": {
    "doc-index": {
      "command": "uvx",
      "args": ["doc-index-mcp"]
    }
  }
}

2. Install the skill (optional)

The skill teaches the agent how to use the search tools effectively (token budgets, boundary expansion, structure-first retrieval):

uvx --from doc-index-mcp doc-index-install-skill

That's it — start asking Claude to index and search your documents.

Other agents

Hermes Agent reads skills from ~/.hermes/skills/<category>/ rather than per-project, and configures MCP servers in ~/.hermes/config.yaml:

uvx --from doc-index-mcp doc-index-install-skill --target hermes
mcp_servers:
  doc-index:
    command: uvx
    args:
      - doc-index-mcp

The installer honours HERMES_HOME, so non-default profiles install to the right place. Restart the Hermes session afterwards so the skill is picked up.

Supported Formats

Format Extensions Notes
Text .txt Plain text
Markdown .md, .markdown Preserves headers for boundaries
PDF .pdf Text extraction with page markers
Word .docx Paragraphs, headings, tables
PowerPoint .pptx Slides, notes, tables
Excel .xlsx, .xls Sheets as tables

Why No External Services?

Component Traditional RAG This Server
Embeddings OpenAI API / hosted model Local ONNX model (fastembed)
Vector DB Pinecone / Weaviate / Qdrant Local file (usearch)
Storage Cloud / managed DB Local .docindex/ directory
Dependencies PyTorch (~2GB) ONNX Runtime (~50MB)

Tools

doc_index

Index a document for semantic search.

{
  "file_path": "docs/manual.pdf",
  "source_name": "manual"
}

doc_search

Search indexed documents using natural language.

{
  "query": "how to configure authentication",
  "top_k": 5,
  "expand_to_boundary": "section",
  "max_return_tokens": 4096
}

Parameters:

  • query - Search query
  • sources - Filter to specific sources (optional)
  • top_k - Number of results (default: 5)
  • expand_to_boundary - Expand results to full "chapter", "section", "subsection", or "page"
  • max_return_tokens - Token budget for results (default: 4096)
  • include_siblings - Include sibling sections when expanding

doc_list

List all indexed sources.

doc_chunk

Retrieve a specific chunk by ID with optional neighbors.

{
  "chunk_id": "manual:42",
  "neighbors": 2
}

doc_toc

Get the table of contents (chapters, sections, subsections) for an indexed document. Use this to understand document structure before retrieving specific content.

{
  "source_name": "manual",
  "max_depth": 3
}

doc_get_content

Retrieve document content by structural location. Provide exactly one locator: boundary_id, chapter, section, or pages.

{
  "source_name": "manual",
  "chapter": "3",
  "max_return_tokens": 8192
}

read_document

Read a document without indexing. Returns formatted text.

{
  "file_path": "report.pdf",
  "max_chars": 100000
}

list_tables

List all tables in a document.

{
  "file_path": "data.xlsx"
}

extract_table

Extract a specific table as CSV.

{
  "file_path": "data.xlsx",
  "table_index": 0,
  "max_rows": 100
}

Environment Variables

Variable Description Default
MCP_WORKING_DIR Base directory for resolving file paths Current working directory
DOC_INDEX_DIR Directory for storing vector indices .docindex in working dir

Alternative Installation

Install globally with pip

pip install doc-index-mcp

Then in your .mcp.json:

{
  "mcpServers": {
    "doc-index": {
      "command": "doc-index-mcp"
    }
  }
}

Install from source

Clone the repo and install dependencies:

git clone https://github.com/mike-anderson/doc-index-mcp.git
cd doc-index-mcp
pip install -e .

Then point your .mcp.json at the server entrypoint:

{
  "mcpServers": {
    "doc-index": {
      "command": "python",
      "args": ["/path/to/doc-index-mcp/src/server.py"]
    }
  }
}

Architecture

Everything runs locally - no external APIs, databases, or embedding servers required.

flowchart TB
    subgraph Client["MCP Client (Claude Desktop, etc.)"]
        LLM[LLM]
    end

    subgraph MCP["Doc Index MCP Server"]
        Server[server.py]

        subgraph Services["Local Services"]
            Loader[Document Loader<br/>PDF, DOCX, PPTX, XLSX]
            Chunker[Boundary-Aware<br/>Chunker]
            Embedder[Embedder<br/>ONNX Runtime]
            VectorStore[Vector Store<br/>usearch]
        end
    end

    subgraph Storage["Local Filesystem"]
        Docs[(Source<br/>Documents)]
        Index[(".docindex/<br/>├── manifest.json<br/>└── vectors/<br/>    ├── index.usearch<br/>    ├── chunks.jsonl<br/>    └── boundaries.json")]
    end

    subgraph Models["Embedded Model (downloaded once)"]
        ONNX[BAAI/bge-small-en-v1.5<br/>ONNX format ~50MB]
    end

    LLM <-->|MCP Protocol| Server
    Server --> Loader
    Server --> Chunker
    Server --> Embedder
    Server --> VectorStore

    Loader -->|read| Docs
    VectorStore <-->|read/write| Index
    Embedder -->|load once| ONNX

    style Client fill:#e1f5fe
    style Storage fill:#fff3e0
    style Models fill:#f3e5f5
    style MCP fill:#e8f5e9

Data Flow

flowchart LR
    subgraph Index["Indexing"]
        direction TB
        A[Document] --> B[Load & Extract Text]
        B --> C[Detect Boundaries]
        C --> D[Chunk ~256 tokens]
        D --> E[Generate Embeddings]
        E --> F[Save to Disk]
    end

    subgraph Search["Searching"]
        direction TB
        G[Query] --> H[Embed Query]
        H --> I[Vector Similarity Search]
        I --> J[Expand to Boundaries]
        J --> K[Return Results]
    end

    Index -.->|stored in .docindex/| Search

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

doc_index_mcp-0.2.0.tar.gz (199.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

doc_index_mcp-0.2.0-py3-none-any.whl (74.6 kB view details)

Uploaded Python 3

File details

Details for the file doc_index_mcp-0.2.0.tar.gz.

File metadata

  • Download URL: doc_index_mcp-0.2.0.tar.gz
  • Upload date:
  • Size: 199.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for doc_index_mcp-0.2.0.tar.gz
Algorithm Hash digest
SHA256 0731a34d50c91e24a99ccff361f1b3807c463353ef88c1b4523b79d8e8ff2b73
MD5 ac657b96a61b9aa8550330acbc8db32d
BLAKE2b-256 89a1d6a20fdaf64addde939181bc17c4c42320ef31bd9f903d46deada266a865

See more details on using hashes here.

Provenance

The following attestation bundles were made for doc_index_mcp-0.2.0.tar.gz:

Publisher: publish.yml on mike-anderson/doc-index-mcp

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file doc_index_mcp-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: doc_index_mcp-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 74.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for doc_index_mcp-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 be041b2d4c580c9183c7a6f517407bb9d1ba7210fd26a8c4cede90185d5d9a64
MD5 3788b944d1097694126671b0f1544e75
BLAKE2b-256 0873fc23d880aa7fd070ba2732c055dba2f040ca9fa92c57d0628afdde2b6140

See more details on using hashes here.

Provenance

The following attestation bundles were made for doc_index_mcp-0.2.0-py3-none-any.whl:

Publisher: publish.yml on mike-anderson/doc-index-mcp

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page