Skip to main content
Archived

This project has been archived by its maintainers, and is no longer receiving any updates.

MCP Nexus

The discovery and routing layer that keeps MCP servers out of your context window until you actually need them.

CI PyPI version Python 3.11+ Coverage License: MIT


The problem

The Model Context Protocol lets an LLM talk to any number of servers — GitHub, Slack, Postgres, your internal tools. The catch: most clients load every tool definition from every configured server at startup. Six servers can mean 70+ tool schemas and thousands of tokens spent before the conversation even starts, most of which the model never touches in a given session.

What MCP Nexus does

MCP Nexus sits in front of your MCP servers as a thin discovery layer. Instead of loading everything up front, the LLM asks for what it needs — by server or by tool — and MCP Nexus resolves the request and connects on demand. You keep your existing MCP servers unmodified; MCP Nexus only changes how (and when) their tools reach the model's context.

Two modes cover the two ways teams actually want this to work:

Discovery Mode Dynamic Mode
Granularity Whole server Individual tool
Tools always in context 4 (mcpd_find, mcpd_list, mcpd_connect, mcpd_get_schema) 1 (find_tools)
Best for "Connect me to GitHub" style workflows Cherry-picking one tool from many servers
After resolution LLM talks to the server directly — MCP Nexus exits the data path MCP Nexus lazy-connects and stays in the loop per tool call

v1.0.0 adds Lazy Schema Loading: Discovery Mode can hand back a stub tool list (names only, no schemas) and fetch a single tool's full schema only when it's about to be called. Real, reproducible numbers (not rough estimates) are in Benchmarks below.

How it works

Discovery Mode — server-level selection

Step 1  LLM -> mcpd_find("github issues")
               MCP Nexus searches the registry
               returns: { id: "github", tools: ["create_issue", "search_repos", ...] }

Step 2  LLM -> mcpd_connect("github")
               MCP Nexus starts the GitHub MCP server
               returns: 20 tools now available as github__create_issue, etc.

Step 3  LLM -> github__create_issue({ title: "...", body: "..." })
               MCP Nexus proxies to the GitHub MCP server, returns the result

Real token counts for this flow are in Benchmarks — see "Discovery Mode (before connect)" and "(after connect)".

Lazy Schema Loading

No proxy, no changes to the target MCP server required:

  1. mcpd_find("github") → choose a server
  2. mcpd_connect("github", lazy_mode=true) → get a stub list (tool names only, no schemas)
  3. mcpd_get_schema("github", "create_issue") → fetch one full schema
  4. github__create_issue(...) → direct call, as always

Pass --sync-on-start so the registry has schemas cached ahead of time via nexus-sync.

Dynamic Mode — tool-level selection

Step 1  LLM -> find_tools("create issue, post slack message")
               MCP Nexus searches the tool index across all servers
               returns: create_issue (github), post_message (slack)
               both tools added to tools/list

Step 2  LLM -> create_issue({ title: "Bug #42" })
               MCP Nexus lazy-connects to the GitHub MCP server
               executes create_issue, returns the result

Step 3  LLM -> post_message({ channel: "#eng", text: "Done" })
               MCP Nexus lazy-connects to the Slack MCP server
               executes post_message, returns the result

Real token counts for this flow are in Benchmarks — see "Dynamic Mode (before find)" and "(after find_tools)".

Installation

# Core — keyword search, stdio transport
pip install mcpnexus

# With HTTP and SSE transport (remote MCP servers)
pip install mcpnexus[http]

# With semantic search (sentence-transformers)
pip install mcpnexus[embeddings]

# With exact token counting for nexus-benchmark (tiktoken)
pip install mcpnexus[benchmark]

# Full installation
pip install mcpnexus[all]

# Development
pip install mcpnexus[dev]

Quick start

1. Build your registry

The registry is a lightweight JSON catalog of your MCP servers and their tool summaries. Build it from your MCP client's config (e.g. Cursor: ~/.cursor/mcp.json, Claude Desktop: ~/Library/Application Support/Claude/claude_desktop_config.json):

nexus-sync --config /path/to/your/mcp-config.json --output registry/mcpd-registry.json

2. Point your MCP client at MCP Nexus

Discovery Mode (4 tools, connect to one server at a time):

{
  "mcpServers": {
    "nexus-server": {
      "command": "nexus-server",
      "args": ["--registry", "/path/to/mcpd-registry.json", "--sync-on-start"]
    }
  }
}

Dynamic Mode (1 tool, cherry-pick tools across all servers):

{
  "mcpServers": {
    "nexus-gateway": {
      "command": "nexus-gateway",
      "args": ["--registry", "/path/to/mcpd-registry.json"]
    }
  }
}

Or invoke the Python module directly (avoids PATH issues):

{
  "mcpServers": {
    "nexus-gateway": {
      "command": "python",
      "args": ["-m", "mcpnexus.dynamic.server", "--registry", "/path/to/mcpd-registry.json"]
    }
  }
}

3. Measure the token savings

nexus-benchmark --registry registry/mcpd-registry.json

(Generate your registry first with nexus-sync.) See Benchmarks for real, reproducible numbers and what they mean.

Architecture

┌─────────────────────────────────────────────────────────────┐
│                        LLM / AI Client                       │
└──────────────────────┬──────────────────────────────────────┘
                       │ MCP (stdio / JSON-RPC 2.0)
          ┌────────────┴─────────────┐
          │                          │
   ┌──────▼──────┐           ┌───────▼──────┐
   │  Discovery  │           │   Dynamic    │
   │    Mode     │           │    Mode      │
   │             │           │              │
   │ mcpd_find   │           │ find_tools   │
   │ mcpd_list   │           │              │
   │ mcpd_connect│           │ LazyPool     │
   │ mcpd_get_schema│        │              │
   └──────┬──────┘           └───────┬──────┘
          │                          │
          └────────────┬─────────────┘
                       │
          ┌────────────▼─────────────┐
          │        Shared Core        │
          │                           │
          │  Registry (mcpd-registry) │
          │  KeywordSearchEngine      │
          │  ToolSearchEngine         │
          │  HybridSearch (TF-IDF +   │
          │    sentence-transformers) │
          │  NexusConnector           │
          │   ├─ stdio transport      │
          │   ├─ streamable-http      │
          │   └─ SSE transport        │
          └───────────────────────────┘

Design principles

A discovery layer, not a permanent proxy. In Discovery Mode, once mcpd_connect resolves, the LLM gets direct tool access to the connected server. In Dynamic Mode, server connections stay lazy — a server process starts only when one of its tools is actually called.

Offline-first registry. Tool summaries (name, description, tags) are captured at sync time. Searches run against the cached registry with zero network traffic; full tool schemas load only on connection.

Search degrades gracefully. Keyword search (TF-IDF with synonyms) is the default — always available, no extra dependencies. Semantic search is optional (pip install mcpnexus[embeddings]): when installed, sentence-transformers embeddings blend with keyword results, which helps for loosely-phrased natural-language queries like "a tool for reading web pages" → Playwright. The keyword synonym table also understands multilingual input (e.g. Polish query terms resolve to the right English tool concepts). Keyword-first keeps installs frictionless when you don't need semantic search.

Benchmarks

Methodology. Every number below comes from nexus-benchmark, which instantiates the real DiscoveryServer/DynamicServer classes and measures their actual tools/list JSON-RPC output — not hardcoded stand-ins that can drift out of sync with the real code. Tokens are counted with tiktoken's o200k_base encoding (the GPT-4o tokenizer) when installed; without it, the CLI clearly labels its output as a rougher char/4 estimate rather than presenting both with false equal precision. Reproduce any number here yourself:

pip install mcpnexus[benchmark]
nexus-benchmark --registry registry/benchmark-registry.json --query "create github issue" --quality

At realistic scale

6 servers, 61 tools, every tool has a complete real schema — registry/benchmark-registry.json, not a partial catalog:

Scenario Tools Tokens vs. direct load
Direct (all servers, no MCP Nexus) 61 3,786
Discovery Mode (before connect) 4 480 87% fewer
Discovery Mode (after connect: github) 18 955 75% fewer
Dynamic Mode (before find) 2 201 95% fewer
Dynamic Mode (after find_tools("create github issue")) 7 390 90% fewer

recall@k on this registry: 100% (8/8) — these savings aren't bought with worse search accuracy (quality.py verifies it).

The honest part: this only pays off at scale

The same benchmark against a deliberately tiny registry (3 servers, 7 tools — the fixture in tests/conftest.py) tells a different story:

Scenario Tools Tokens vs. direct load
Direct (all servers, no MCP Nexus) 7 252
Discovery Mode (before connect) 4 480 90% more
Discovery Mode (after connect) 7 576 129% more
Dynamic Mode (before find) 2 201 20% fewer
Dynamic Mode (after find_tools) 3 288 14% more

Discovery Mode's own 4-tool interface has real, fixed token overhead. Below a certain number of configured servers, that overhead costs more than just loading everything directly — Dynamic Mode's leaner 2-tool interface still wins here, but barely. If you have 2-3 MCP servers you always use, you don't need MCP Nexus. It exists for the case competitor benchmarks don't show: many configured servers, only a few actually used per session.

Versus other tools on the market

One real side-by-side, run ourselves: NCP Orchestrator v2.3.1, installed fresh (npx -y @portel/ncp@latest), zero backend servers configured — its own static meta-tool interface, tokenized the exact same way:

Tool Meta-tools exposed Tokens
NCP v2.3.1 (find + code) 2 903
MCP Nexus Discovery Mode (before connect) 4 480
MCP Nexus Dynamic Mode (before find) 2 201

Measured 2026-07-29, tiktoken o200k_base. This is the one comparison in this section we actually ran ourselves — same tokenizer, same "before connecting anything" scenario, reproducible by anyone with Node.js installed.

For everyone else below, we're citing published numbers, not our own measurements — different registries, different tokenizers, different baselines. Treat these as directional, not as line-by-line comparable to the numbers above:

Tool Claimed reduction Source
Anthropic native Tool Search (Claude Code) ~85–96% (reported 134k→5k tokens internally) community writeup
Speakeasy Dynamic Toolsets ~99% ("100x") speakeasy.com
NCP Orchestrator (vendor-claimed) 83–97%, varies by source arul.sg/ncp, mcp.directory

The most important line in this table isn't a percentage: Anthropic shipped this exact pattern natively into Claude Code. If you're specifically on Claude Code, check whether you need any third-party discovery layer — this one included — before reaching for one.

Registry format

The registry file (mcpd-registry.json) is a JSON catalog of MCP servers:

{
  "mcpd_version": "1.0",
  "metadata": {
    "name": "My MCP Registry",
    "description": "Personal registry of MCP servers"
  },
  "servers": [
    {
      "id": "github",
      "name": "GitHub MCP Server",
      "description": "Official GitHub MCP server (remote). Repositories, issues, pull requests, and code search",
      "version": "remote-2025-11",
      "transport": {
        "type": "streamable-http",
        "url": "https://api.githubcopilot.com/mcp/",
        "headers": { "Authorization": "Bearer ${GITHUB_MCP_PAT}" }
      },
      "tags": ["github", "git", "code", "issues"],
      "tools_summary": [
        {
          "name": "issue_write",
          "description": "Create or update an issue or pull request",
          "tags": ["issues", "create"]
        }
      ],
      "estimated_tools_count": 90,
      "enabled": true,
      "last_synced": "2026-06-11T00:00:00Z"
    }
  ]
}

Full schema: registry/schemas/mcpd-schema.json

Project structure

mcpnexus/
├── mcpnexus/                    # Python package
│   ├── __init__.py              # Public API and version
│   ├── models.py                # Shared dataclasses
│   ├── registry.py               # Registry loader (mcpd-registry.json)
│   ├── connector.py              # MCP connector — stdio, HTTP, SSE transports
│   ├── sync.py                   # Registry builder (sync from mcp.json)
│   ├── benchmark.py              # Token savings measurement
│   ├── search/
│   │   ├── keyword_search.py    # Server-level TF-IDF search
│   │   ├── tool_search.py       # Tool-level TF-IDF search
│   │   ├── embeddings.py        # Sentence-transformer embedding engine
│   │   └── hybrid.py            # Hybrid keyword + semantic search
│   ├── discovery/
│   │   └── server.py            # Discovery Mode MCP server
│   └── dynamic/
│       ├── server.py            # Dynamic Mode MCP server
│       ├── tool_index.py        # O(1) tool lookup index
│       └── lazy_pool.py         # On-demand connection pool
├── registry/
│   ├── mcpd-registry.example.json  # Example registry
│   ├── benchmark-registry.json  # Fully-specified registry used by the Benchmarks section
│   └── schemas/
│       └── mcpd-schema.json     # JSON Schema for registry validation
├── docs/
│   ├── specification.md         # Protocol specification
│   ├── architecture.md          # Architecture deep-dive
│   ├── registry-format.md       # Registry format reference
│   └── dynamic-mcp.md           # Dynamic Mode guide
├── examples/
│   ├── cursor-config-discovery.json
│   ├── cursor-config-dynamic.json
│   └── README.md
├── tests/                       # 509 tests, 100% coverage
└── pyproject.toml

Development

git clone https://github.com/KrzysztofAugiewicz/MCPNexus.git
cd MCPNexus
pip install -e ".[dev]"

# Run tests
pytest

# Run tests with coverage
pytest --cov=mcpnexus --cov-report=term-missing

# Run end-to-end integration test
python test_e2e.py

# Benchmark token savings (generate registry first with nexus-sync)
nexus-benchmark --registry registry/mcpd-registry.json

CLI reference

Command Description
nexus-server Start the Discovery Mode MCP server
nexus-gateway Start the Dynamic Mode MCP server
nexus-sync Build or update the registry from an mcp.json config
nexus-benchmark Measure token savings for a given registry

All commands accept --help for the full option reference.

Transport support

Transport Install extra Use case
stdio (core) Local process-based MCP servers
Streamable HTTP mcpnexus[http] Remote HTTP MCP servers
SSE mcpnexus[http] Legacy remote servers (Server-Sent Events)

Transport type is resolved automatically from the registry entry's transport.type field.

Documentation

Publishing to PyPI

Releases are published automatically when a GitHub Release is created. Prerequisites:

  1. Add PYPI_API_TOKEN to repository secrets (create at pypi.org/manage/account/token)
  2. Create a release with a tag (e.g. v1.0.1)

The publish workflow builds and uploads to PyPI.

Contributing

Contributions are welcome. Please read CONTRIBUTING.md before opening a pull request. For bug reports and feature requests, use GitHub Issues.

Authors

  • Krzysztof Augiewicz — Lead Architect & Creator — LinkedIn · GitHub
  • Kacper Pisarczyk — Core Contributor, Discovery & Registry Systems — LinkedIn
  • Sebastian Pawłowski — Advisory & QA Support (testing, hardware/software provisioning) — LinkedIn
  • Mateusz Wiszniowski — Core Contributor

Full details in AUTHORS.md.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mcpnexus-1.0.5.tar.gz (118.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mcpnexus-1.0.5-py3-none-any.whl (67.4 kB view details)

Uploaded Python 3

File details

Details for the file mcpnexus-1.0.5.tar.gz.

File metadata

  • Download URL: mcpnexus-1.0.5.tar.gz
  • Upload date:
  • Size: 118.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mcpnexus-1.0.5.tar.gz
Algorithm Hash digest
SHA256 72cd45f8e726fdac6414ec25b3bd7d7c8bef61acdc522cd9fc3a2cdebc099a09
MD5 d0a03bc8aa4c7a21e2fe9a6d94218b02
BLAKE2b-256 0e95fe1d8ed8002be6564e3143984bdcb99a0ed8c261af8a4b24f041e12e8994

See more details on using hashes here.

File details

Details for the file mcpnexus-1.0.5-py3-none-any.whl.

File metadata

  • Download URL: mcpnexus-1.0.5-py3-none-any.whl
  • Upload date:
  • Size: 67.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mcpnexus-1.0.5-py3-none-any.whl
Algorithm Hash digest
SHA256 271bb2b86fe79d0c853dacf5b92d9c432fb43b197e93529a1692dc6ac07de66a
MD5 c521948ccb521a385e96d6b37839ed0b
BLAKE2b-256 a9c10ce6933a71f82c924c185df89c1e5fe2dda4d2c3889d8922c7df0c4e5001

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page