Skip to main content

MCP Rancher — the capability-aware control plane for AI-assisted Rancher operations

CI MCP server 206 tools Python 3.12+ Pyright strict MIT license

Operate Rancher-managed Kubernetes through any MCP client — discovery, generic resource access, and curated operator workflows, wrapped in an audit-logged, rate-limited, confirmation-guarded safety model.

Quick start · Tool surface · Architecture · Safety model · Compatibility · Development


Why this exists

Rancher is how real fleets run Kubernetes — and it speaks two APIs (the legacy Norman /v3 plane and the modern Steve /v1 plane), varies by version, and wraps every cluster behind its own proxy. Pointing a generic Kubernetes MCP server at it misses everything Rancher-specific; pointing an agent at raw kubectl gives up auditability, guardrails, and the management-plane view entirely.

MCP Rancher is built for that reality:

  • Capability-aware, not version-naive. It detects what each connected Rancher actually supports instead of assuming. One binary spans 2.6.5 → 2.9.3 with the same tool surface.
  • Multi-instance first. Lab, staging, prod — configure them all; mark prod read_only: true and every mutation is refused at the config layer, before any guard even has to fire.
  • Nothing is out of reach. Curated tools cover the common 95%; the generic engine reaches every resource either API plane exposes — even types nobody wrote a tool for yet.

The tool surface

206 tools: 178 read-only · 28 writes · 5 destructive — counted from the registry itself, not by hand. docs/tool-manifest.json is generated from the live FastMCP registry (make tool-manifest) and a CI gate fails the build if it ever drifts from the code. Per-tool descriptions, safety annotations, and parameters all live there; the narrative registry with slice tracking is docs/tool-catalog.md.

Layer What it does Examples
Discovery & schema Explore what any instance can do rancher_server_version, rancher_norman_schema_list, rancher_capability_domain_list
Generic engine CRUD + actions + links + watch on any resource, both planes rancher_steve_resource_list, rancher_norman_resource_action_invoke, rancher_steve_resource_watch
Curated reads Typed, shaped responses across ~25 domains rancher_pods_list, rancher_deployments_list, rancher_longhorn_volumes_list, rancher_policy_reports_list
Curated writes Guarded mutations rancher_deployment_scale, rancher_deployment_restart, rancher_cron_job_suspend, rancher_node_cordon, rancher_secret_create
Operator rollups One-call triage rancher_cluster_health_check, rancher_find_failing_pods, rancher_find_stalled_rollouts, rancher_project_health_summary

Domains covered: clusters & nodes · projects & namespaces · workloads · pods & services · storage · networking · config & secrets (values masked) · certificates (keys masked) · RBAC · auth & identity · apps & catalogs · logging pipeline · Prometheus monitoring · policy reports · CIS compliance · backup operator · etcd backups · Longhorn · Fleet · provisioning · settings & features · alerts & notifiers.

All 206 stay exposed by default — every tool schema is deferred behind Claude Code's own search, so a small default would only help other hosts at the good host's expense. A constrained host (a small local model, a tight context budget) can opt into a smaller surface via RANCHER_TOOLSETS; see Toolset profiles under Configuration.

Quick start

Requirements

Install & run

# From source
git clone https://github.com/rex/mcp-rancher.git
cd mcp-rancher
make setup                 # deps, .env scaffold, pre-commit hooks
cp .env.example .env       # set RANCHER_URL + RANCHER_TOKEN
make dev                   # run the MCP server (stdio)

Once published to PyPI, it's one line: uvx rancher-mcp.

Claude Code

claude mcp add rancher \
  -e RANCHER_URL=https://rancher.example.com \
  -e RANCHER_TOKEN=token-xxxxx:yyyyyyyyy \
  -- uv run --directory /path/to/mcp-rancher rancher-mcp

Claude Desktop

{
  "mcpServers": {
    "rancher": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/mcp-rancher", "rancher-mcp"],
      "env": {
        "RANCHER_URL": "https://rancher.example.com",
        "RANCHER_TOKEN": "token-xxxxx:yyyyyyyyy"
      }
    }
  }
}

Multiple instances

RANCHER_INSTANCES_JSON='{
  "production": {"url": "https://rancher.prod.example.com",    "token": "token-a:xxx", "verify_ssl": true, "read_only": true},
  "lab":        {"url": "https://rancher.lab.example.com",     "token": "token-b:yyy", "verify_ssl": false, "read_only": false}
}'
RANCHER_DEFAULT_INSTANCE=production

Every tool takes an optional instance argument. Instances flagged read_only: true refuse all mutations at the settings layer.

Architecture

flowchart LR
    A[MCP client<br/>Claude Code · Claude Desktop · any] -- stdio --> S

    subgraph S[rancher-mcp]
        direction TB
        L1[Discovery & schema<br/>planes · schemas · capabilities]
        L2[Generic engine<br/>any resource · both planes<br/>CRUD · actions · links · watch]
        L3[Curated tools<br/>typed models · shaped output<br/>next-step hints]
        G[Safety layer<br/>read-only guard · confirmation phrases<br/>audit log · rate limit · masking]
        L1 --> L2 --> L3
        L3 --> G
        L2 --> G
    end

    G -- Norman /v3 --> R1[(Rancher<br/>instance A)]
    G -- Steve /v1 + k8s proxy --> R1
    G -- Norman + Steve --> R2[(Rancher<br/>instance B)]

Three layers, deliberately separate: discovery tells you what an instance can do, the generic engine can touch anything it exposes, and curated tools make the common paths typed, shaped, and self-describing (every response carries suggested_next_steps). Most curated tools are generated from YAML descriptors (catalog/curated_tools/) with a drift gate — the editorial decisions live in descriptors, not boilerplate.

Safety model

Built for the day an agent is pointed at the cluster that pays your salary:

Guard Behavior
Read-only instances read_only: true refuses every mutation for that instance, before tool logic runs
Destructive confirmation Deletes require an explicit typed phrase (e.g. "delete steve namespace foo") — no phrase, no delete
Tool annotations Every tool declares readOnlyHint / destructiveHint / idempotentHint, so clients can gate UX on them
Audit log Every mutation emits a structured event="audit" record — tool, operation, plane, instance, resource, outcome. Argument names only; values never logged
Rate limiting Token-bucket on writes (default 60/min) — a runaway loop can't machine-gun your API
Secret & key masking Secret values and certificate private keys are structurally absent from curated responses (reveal is an explicit generic-tool opt-in)
Structured errors Guard rejections return typed error_code envelopes agents can branch on — never raw strings

Compatibility

Primary target Rancher 2.9.3 (production-validated)
Compatibility floor Rancher 2.6.5 (kept green via capability detection)
API planes Norman /v3 + Steve /v1 (+ per-cluster Kubernetes proxy)
Transport stdio

Capability detection bridges version differences at runtime — no version-pinned builds, no "works on my Rancher." Both targets are exercised by the same test suite, and read paths have been validated live against both a 2.6.5 lab and a 2.9.3 production fleet (validation report).

Configuration

Variable Default Purpose
RANCHER_URL — Rancher server URL (single-instance mode)
RANCHER_TOKEN — API token (token-xxxxx:yyyyyyyyy)
RANCHER_VERIFY_SSL true TLS verification
RANCHER_INSTANCES_JSON — Multi-instance config (see above)
RANCHER_DEFAULT_INSTANCE first defined Instance used when a tool call names none
RANCHER_MCP_SERVER_NAME rancher-mcp Server identity announced to clients
RANCHER_MCP_SERVER_DESCRIPTION built-in Server description announced to clients
RANCHER_MCP_WRITE_RATE_LIMIT_PER_MIN 60 Write rate limit (0 disables)
RANCHER_TOOLSETS all Comma-separated toolset profile(s) exposed at startup — see below
RANCHER_TOOLS — Comma-separated tool names force-included on top of the selected profile(s)
RANCHER_EXCLUDE_TOOLS — Comma-separated tool names removed, applied last — always wins over RANCHER_TOOLS

Toolset profiles

The default is all: every tool stays exposed. Claude Code, the primary host, defers every tool schema behind its own search, so a small default would only help other hosts at the cost of making the good host worse — this is a deliberate choice, not an oversight.

RANCHER_TOOLSETS opts a constrained host (a small local model, or a context budget) into a smaller surface. Values are either a family name — one per src/rancher_mcp/tools/ module (storage, workloads, pods_services, rbac, …; see rancher_mcp.toolsets.FAMILY_REGISTRARS for the full list) — or the cross-family core profile: a ~32-tool triage/orientation set (rancher_find_*, the health/summary rollups, core list/get pairs, and the generic rancher_{steve,norman}_resource_{list,get} escape hatches). core cuts the tools/list payload from 206 tools / ~401 KB to 32 tools / ~78 KB (~80% smaller). Profiles compose: RANCHER_TOOLSETS=core,storage gives the triage set plus everything storage-related.

RANCHER_TOOLSETS=core                        # small triage surface
RANCHER_TOOLSETS=core,storage,workloads      # triage + two full families
RANCHER_TOOLS=rancher_secret_get             # add one extra tool on top
RANCHER_EXCLUDE_TOOLS=rancher_secret_create  # remove one, wins over the above

An unknown profile name fails loudly at startup rather than silently producing an empty or shrunken surface. Calling a real tool that exists but isn't in the active profile returns a structured TOOLSET_NOT_ENABLED error naming the tool, the toolset that would enable it, and the env var to set — never a bare "unknown tool", which would make a disabled tool indistinguishable from a typo.

Project status

Shipping and stable for read, triage, and guarded write operations. Honest ledger of what's beyond that:

  • Destructive workflows (node drain, etcd/backup restore, cert rotation, cluster upgrade/delete) are roadmap — deliberately staged after real-world read-path mileage. The generic engine + confirmation guard already covers these cases for operators who need them today.
  • The Alertmanager routes/silences surface needs an in-cluster API integration and is deferred.
  • The full per-version compatibility matrix (Track G) is in progress; the first live validation run covers the read matrix on both targets.

Work is tracked to the tool level: docs/tool-catalog.md (every tool has a row, every gap a slice ID) and ROADMAP.md.

Development

make help               # every target, documented
make validate           # codegen drift + manifest drift + architecture + lint + typecheck + tests
make tool-manifest      # regenerate docs/tool-manifest.json from the registry
make lab-up             # local Rancher 2.6.5 lab (kind + helm), fully scripted
make integration-current # isolated Rancher 2.14.3 end-to-end test run
make live-read-matrix   # read-only validation probes against configured instances
make mock-rancher       # fixture-backed mock Rancher for provider-config testing
  • Local lab — a self-contained Rancher 2.6.5 on kind with a simulated downstream cluster; repo-local kubeconfigs, never touches your machine state.
  • Current integration lab — Rancher 2.14.3 on separate Kind clusters and port 9443; run it serially with make integration-current to avoid overlapping Docker resource demand with the legacy lab.
  • Contract fixtures — sanitized captures from live Rancher committed under tests/fixtures/; respx pins the HTTP boundary in tests.
  • Codegen — curated tools are emitted from catalog/curated_tools/*.yml descriptors; make check-codegen fails on drift.
  • Gates — ruff, pyright strict, 624 tests with coverage floor, architecture line-limits, module-shape checks, secret scanning. All fail closed, all wired into pre-commit.

Stack: Python 3.12 · FastMCP · httpx · Pydantic v2 · structlog · uv

Security

See SECURITY.md for the threat model, token guidance, and how to report vulnerabilities.

License

MIT

Metadata

Release files for rancher-mcp 1.59.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rancher-mcp 1.59.0
File Size Uploaded
rancher_mcp-1.59.0.tar.gz 234.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rancher-mcp 1.59.0
File Interpreter ABI Platform
rancher_mcp-1.59.0-py3-none-any.whl Python 3 none any Details

Total release size: 690.0 kB

Release files / rancher_mcp-1.59.0.tar.gz

Download URL rancher_mcp-1.59.0.tar.gz
Size 234.3 kB
Tags Source
SHA-256 checksum
How to use checksums
3832718c624b5b0dd4d2bc4314ec8953564e4f8aaa25b0b99ef73bc0f4ceae01
BLAKE2b-256 checksum
How to use checksums
3762a8d3123699ca51b21278dd8b06d33fc20ef827fe28fbae44a41f59a6a218
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / rancher_mcp-1.59.0-py3-none-any.whl

Download URL rancher_mcp-1.59.0-py3-none-any.whl
Size 455.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6c17629d5646498fcc715130790fde8c20e09a2a63d692bc6d5dcc33503f69b0
BLAKE2b-256 checksum
How to use checksums
30830e7f4e542b454dedca5038f7d29987285b2051e32745b28cac14bde2ca15
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

1.59.0 This release

2 release files

1.26.4

2 release files

1.12.3

2 release files

1.12.2

2 release files

1.12.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page