MCP server for GitHub code retrieval - semantic search for Claude Desktop & Claude Code CLI
Project description
GitHub Code Retrieval MCP Server
An MCP (Model Context Protocol) server that gives Claude the ability to search and analyze any GitHub repository using semantic code search.
Works with: Claude Desktop, Claude Code CLI, Cursor IDE
Fully local - No API keys required. All embeddings run locally.
Quick Start (2 Steps)
Step 1: Install
pip install github-code-retrieval
Step 2: Add to Claude Config
For Claude Desktop - Add to your config file:
| OS | Config File Location |
|---|---|
| macOS | ~/Library/Application Support/Claude/claude_desktop_config.json |
| Windows | %APPDATA%\Claude\claude_desktop_config.json |
| Linux | ~/.config/claude/claude_desktop_config.json |
{
"mcpServers": {
"github-code-retrieval": {
"command": "github-code-retrieval"
}
}
}
For Claude Code CLI - Add to ~/.claude.json or .claude.json in your project:
{
"mcpServers": {
"github-code-retrieval": {
"command": "github-code-retrieval"
}
}
}
Restart Claude after adding the config. Done!
Usage
Once configured, just ask Claude to analyze any GitHub repo:
"Analyze https://github.com/pallets/flask and explain how routing works"
"How does authentication work in https://github.com/tiangolo/fastapi?"
"Find the database connection code in https://github.com/django/django"
Claude will automatically use the tool to:
- Clone the repository
- Index all code files locally
- Search for relevant code snippets
- Return the most relevant code to answer your question
Features
- No API Keys - Everything runs locally on your machine
- Semantic Search - Finds code by meaning, not just keywords
- Fast - First query indexes the repo (~30s), subsequent queries are instant (~0.01s)
- Private Repos - Works with private repos if you have git credentials set up
- 20+ Languages - Python, JavaScript, TypeScript, Go, Rust, Java, and more
How It Works
┌─────────────────────────────────────────────────────────────────┐
│ You: "How does routing work in this Flask app?" │
│ │ │
│ ▼ │
│ Claude calls: analyze_github_repo( │
│ repo_url="https://github.com/pallets/flask", │
│ question="How does routing work?" │
│ ) │
│ │ │
│ ▼ │
│ MCP Server (runs locally on your machine): │
│ ┌───────────────────────────────────────────────────────────┐ │
│ │ 1. Clone repo (cached for 24h) │ │
│ │ 2. Index code with local embeddings (sentence-transformers)│ │
│ │ 3. Semantic search (ChromaDB) │ │
│ │ 4. Return top relevant code snippets │ │
│ └───────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Claude receives code snippets and explains them to you │
└─────────────────────────────────────────────────────────────────┘
Tool Schema
The MCP server exposes one tool: analyze_github_repo
Input:
{
"repo_url": "https://github.com/owner/repo",
"question": "How does authentication work?",
"top_k": 10
}
Output:
{
"success": true,
"repository": {
"url": "https://github.com/owner/repo",
"owner": "owner",
"name": "repo",
"total_files_indexed": 45,
"total_chunks": 230
},
"query": "How does authentication work?",
"code_snippets": [
{
"file_path": "src/auth/handler.py",
"content": "def authenticate(request):\n ...",
"start_line": 45,
"end_line": 78,
"language": "python",
"relevance_score": 0.89
}
],
"total_results": 10
}
Configuration (Optional)
Create a .env file to customize settings:
# Embedding model (default: all-MiniLM-L6-v2)
EMBEDDING_MODEL=all-MiniLM-L6-v2
# Device for embeddings: cpu, cuda (NVIDIA), mps (Apple Silicon)
EMBEDDING_DEVICE=cpu
# Storage paths
VECTOR_STORE_PATH=./data/vector_db
REPO_STORAGE_PATH=./data/repos
# Indexing settings
CHUNK_SIZE=1500
CHUNK_OVERLAP=200
GPU Acceleration
For faster embeddings:
NVIDIA GPU:
EMBEDDING_DEVICE=cuda
Apple Silicon:
EMBEDDING_DEVICE=mps
Private Repositories
The tool uses your system's git credentials. To access private repos:
# Option 1: GitHub CLI (recommended)
gh auth login
# Option 2: SSH key
# Just make sure your SSH key is set up for GitHub
# Option 3: Credential helper
git config --global credential.helper store
Alternative: HTTP Server Mode
If you want to run the server on a network (e.g., shared team server):
# Start HTTP server
github-code-retrieval-http
# Server runs on http://0.0.0.0:8000
# API docs at http://localhost:8000/docs
Then configure clients to connect:
{
"mcpServers": {
"github-code-retrieval": {
"command": "http",
"args": ["http://SERVER_IP:8000"]
}
}
}
Supported Languages
Python, JavaScript, TypeScript, JSX, TSX, Java, Kotlin, Go, Rust, C, C++, C#, Ruby, PHP, Swift, Scala, SQL, GraphQL, YAML, JSON, TOML, Markdown, and more.
Troubleshooting
"Command not found: github-code-retrieval"
Make sure pip's bin directory is in your PATH:
# Find where pip installs scripts
python -m site --user-base
# Add that path + /bin to your PATH
Or use the full path:
{
"mcpServers": {
"github-code-retrieval": {
"command": "python",
"args": ["-m", "mcp_stdio_server"]
}
}
}
Slow first query
The first query for a repo takes ~30-60 seconds to:
- Clone the repository
- Load the embedding model (~400MB)
- Index all code files
Subsequent queries are fast (~0.01s).
Out of memory
Large repos may need more RAM. Try:
- Reducing
CHUNK_SIZEin.env - Using a smaller embedding model
Development
# Clone the repo
git clone https://github.com/yourusername/github-code-retrieval
cd github-code-retrieval
# Create virtual environment
python -m venv venv
source venv/bin/activate # or venv\Scripts\activate on Windows
# Install in development mode
pip install -e ".[dev]"
# Run tests
pytest
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file github_code_retrieval-2.0.0.tar.gz.
File metadata
- Download URL: github_code_retrieval-2.0.0.tar.gz
- Upload date:
- Size: 46.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0rc1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
946de9f40d91b76325cb7d2d8296f6def76a1007386732dc7c36bb3a5df1f31b
|
|
| MD5 |
6693cdc9068797bad6ebf860c7d76148
|
|
| BLAKE2b-256 |
6d7f41f380a275f6dfe4f7fbf1494abd7afead4135d883f4b0ea5ed5e0816977
|
File details
Details for the file github_code_retrieval-2.0.0-py3-none-any.whl.
File metadata
- Download URL: github_code_retrieval-2.0.0-py3-none-any.whl
- Upload date:
- Size: 57.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0rc1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1ce276ad10f0415f6cd8e0763db714727903f6ba083bc8b7219b6fb40e9d716c
|
|
| MD5 |
5c6478d0083e491831af5e917f905863
|
|
| BLAKE2b-256 |
a6ac5c29488f39579ef3edcc7b75f28370c5a2ce2780dc877e80a3da4eac14cb
|