Skip to main content

TokenNuke

Intelligent code indexing MCP server. 15 tools, 10 languages, tree-sitter AST extraction, hybrid search (FTS5 + vector), call graphs, remote repo indexing, incremental indexing.

Save 99% of tokens — get exact function source via byte-offset seek instead of reading entire files.

Formerly codemunch-pro. Same code, better name.

Install

pip install tokennuke

Quick Start

Claude Desktop / Cline

Add to your MCP client config:

{
  "mcpServers": {
    "tokennuke": {
      "command": "tokennuke"
    }
  }
}

HTTP Server

tokennuke --transport streamable-http --port 5002

15 MCP Tools

Tool Description
index_folder Index a local directory (incremental, SHA-256 based)
index_repo Index a GitHub/GitLab repo (tarball download, no git needed)
list_repos List all indexed repositories with stats
invalidate_cache Force re-index a repository
file_tree Get directory tree with file counts
file_outline List symbols in a single file
repo_outline List all symbols in repo (summary)
get_symbol Get full source of one symbol (O(1) byte seek)
get_symbols Batch get multiple symbols
search_symbols Hybrid search (FTS5 + vector RRF)
search_text Full-text search in file contents
get_callees What does this function call?
get_callers Who calls this function?
diff_symbols What changed since last index? (PR review)
dependency_map What does this file depend on? What depends on it?

10 Languages

Python, JavaScript, TypeScript, Go, Rust, Java, C, C++, C#, Ruby

All via tree-sitter-language-pack — zero compilation, pre-built binaries.

Key Features

O(1) Symbol Retrieval

Every symbol stores its byte offset and length. get_symbol seeks directly to the function source — no reading entire files. A 200-byte function from a 40KB file = 99.5% token savings.

Incremental Indexing

Files are hashed (SHA-256). Only changed files are re-parsed. Re-indexing a 10K file repo after changing one file takes milliseconds.

Hybrid Search (FTS5 + Vector)

Combines BM25 keyword matching with semantic vector similarity using Reciprocal Rank Fusion. Search "authentication middleware" and find auth_middleware, verify_token, and login_handler.

Call Graphs

Traces function calls through the AST. get_callees("main") shows what main calls. get_callers("authenticate") shows who calls authenticate. Supports depth traversal.

Remote Repo Indexing

Index any public GitHub or GitLab repo by URL — no git binary needed. Downloads the tarball via API, extracts, and indexes. Cached locally with SHA-based freshness checks. Supports private repos with auth tokens and sparse paths.

Full-Text Content Search

Search raw file contents — string literals, TODO comments, config values, error messages. Not just symbol names.

How It Works

  1. Parse — tree-sitter builds an AST for each source file
  2. Extract — Walk AST to find functions, classes, methods, types, interfaces
  3. Store — SQLite database per repo with FTS5 virtual tables
  4. Embed — FastEmbed (ONNX, CPU-only) generates 384-dim vectors for semantic search
  5. Graph — Call expressions extracted from function bodies, edges stored and resolved
  6. Serve — FastMCP exposes 15 tools via stdio or HTTP

Architecture

~/.tokennuke/
├── myproject_a1b2c3d4e5f6.db    # Per-repo SQLite database
├── otherproject_7890abcdef.db
└── ...

Each DB contains:
├── files          # Indexed files with SHA-256 hashes
├── symbols        # Functions, classes, methods, types
├── symbols_fts    # FTS5 full-text search index
├── symbols_vec    # sqlite-vec 384-dim vector index
├── call_edges     # Call graph (caller → callee)
└── file_content_fts  # Raw file content search

Use Cases

  • AI Coding Agents: Give your agent surgical access to codebases without burning context
  • Code Review: Find all callers of a function before changing its signature
  • Onboarding: Search symbols semantically — "where is error handling?" finds relevant code
  • Refactoring: Map call graphs before moving functions between modules
  • Documentation: Extract all public APIs with signatures and docstrings

License

MIT

Release files for tokennuke 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokennuke 1.3.0
File Size Uploaded
tokennuke-1.3.0.tar.gz 33.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokennuke 1.3.0
File Interpreter ABI Platform
tokennuke-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size:64.6 kB

Release files / tokennuke-1.3.0.tar.gz

Download URL tokennuke-1.3.0.tar.gz
Size 33.6 kB
Tags Source
SHA-256 checksum
How to use checksums
b9912867d028ff0bc533ff65cae29decf4ea9cb5bb76d1e9ae78e87437b2929c
BLAKE2b-256 checksum
How to use checksums
ab9e76994e5271c71f8212093780d79c88d01618536b6ecba17c449dbf278fe7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release files / tokennuke-1.3.0-py3-none-any.whl

Download URL tokennuke-1.3.0-py3-none-any.whl
Size 31.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b6ea30a6ca1ebb162b5d32224058f21320f3ca198cd089c3f5dc060a7e159fa5
BLAKE2b-256 checksum
How to use checksums
19f9d46a260232364902184c70e38f857e63cca9e0b551be667990e7047b42a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page