Skip to main content

matmul-mcp

An MCP server for MatMul, so you can find and launch GPU compute from Claude Code, Claude Desktop, Cursor, VS Code or any other MCP client. Ask things like "find me the cheapest H100 under $3/hr and give me an SSH command", and the assistant drives the MatMul API for you. Every launch needs your explicit OK on a price quote first.

It runs locally over stdio and talks to https://api.matmul.cloud/v1 with your API key. It is built on the stdlib-only matmul-cloud Python SDK (import name matmul) and the official mcp Python SDK (FastMCP).

Install

You need an API key. Mint one in the console at https://app.matmul.cloud/app/cli.

uvx matmul-mcp                 # run without installing (recommended)
pipx install matmul-mcp        # or install the `matmul-mcp` command

pip install matmul-mcp also installs the matmul CLI (from matmul-cloud), so matmul login works too.

Authentication is resolved in this order:

  1. the MATMUL_API_KEY environment variable;
  2. the key saved by matmul login --api-key <key> (~/.config/matmul/config.json, or $MATMUL_CONFIG).
Variable Meaning
MATMUL_API_KEY API key (Authorization: Bearer ...).
MATMUL_API_URL API base URL. Defaults to https://api.matmul.cloud; a trailing /v1 is accepted.
MATMUL_MCP_MAX_HOURLY_USD Optional hard cap: this server refuses to quote or launch any instance above this $/hr, whatever the model asks for. Also caps find_gpus.
MATMUL_MCP_ENABLE_JOBS 1 adds the preview managed-jobs tools (off by default: see below).

The product was previously called Lemnos: the old LEMNOS_* variable names still work when the MATMUL_* one is unset, and a key saved in ~/.config/lemnos/config.json is still read.

Configure your client

In every snippet below, replace lmk_... with your key. You can leave MATMUL_API_KEY out if you've run matmul login on this machine. The cap is optional but recommended.

Claude Code

claude mcp add matmul \
  -e MATMUL_API_KEY=lmk_... \
  -e MATMUL_MCP_MAX_HOURLY_USD=5 \
  -- uvx matmul-mcp

Add --scope user to make it available in every project.

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "matmul": {
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "lmk_...",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}

If Claude Desktop can't find uvx, use its absolute path (which uvx).

Cursor

~/.cursor/mcp.json (global) or .cursor/mcp.json (per project):

{
  "mcpServers": {
    "matmul": {
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "lmk_...",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}

VS Code (GitHub Copilot agent mode)

.vscode/mcp.json. VS Code prompts for the key once and stores it securely:

{
  "inputs": [
    { "type": "promptString", "id": "matmul-key", "description": "MatMul API key", "password": true }
  ],
  "servers": {
    "matmul": {
      "type": "stdio",
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "${input:matmul-key}",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}

What it exposes

Tools

Tool Does Annotations
find_gpus(gpu?, max_hourly_usd?, region?, count?, limit) Live GPU offers, cheapest first read-only
get_balance(include_ledger?, ledger_limit?) Prepaid balance, frozen state, new-account limits, burn rate, runway read-only
add_credit(amount_usd) Returns a Stripe checkout URL for the human to open and pay; charges nothing itself write
list_instances(include_deleted?) Instances with status, price, accrued cost, SSH commands read-only
get_instance(id_or_name) One instance; poll it until running to get ssh_command read-only
launch_instance(name, gpu, max_hourly_usd, image?, keep_alive?, ssh_key?, count?, region?, confirm_token?) Two-step: quote, then launch (see below) write, spends money
terminate_instance(id_or_name) Delete an instance and stop billing destructive
list_ssh_keys() / add_ssh_key(name, public_key) Manage the keys installed on new machines. add_ssh_key refuses anything that looks like a private key. read-only / write

With MATMUL_MCP_ENABLE_JOBS=1: run_job, list_jobs, job_logs, cancel_job (destructive). These are opt-in because the /v1/jobs API is still a preview and has no price guard yet. Turn them on by default once it has one.

Every read tool is marked readOnlyHint, and terminate_instance and cancel_job are marked destructiveHint, so clients that auto-approve reads still ask before destructive calls.

Resources and prompts

  • matmul://instances: your non-deleted instances, as JSON.
  • matmul://balance: balance and frozen state, as JSON.
  • Prompt gpu_dev_box(gpu, max_hourly_usd, image): "launch a GPU dev box". It walks the assistant through find, add key, quote, confirm, wait and SSH.

Money safety

launch_instance rents a real machine that bills every hour until it is terminated. The safeguards:

  1. max_hourly_usd is required. It's sent as the API's max_hourly_price_cents price guard. The API refuses to launch if the live price has risen above it, and the tool reports "Price guard: nothing was launched and nothing was charged".
  2. Two-step confirm (quote → token → launch). The first call launches nothing. It picks the cheapest live offer at or under the ceiling and returns a quote: offer, estimated $/hr, the ceiling, image or plain VM, SSH key, balance and runway, and a line saying it bills until terminated. It also returns a confirm_token. Only a second call with that token and identical arguments launches. The launch result repeats the hourly cost and says to call terminate_instance.
  3. MATMUL_MCP_MAX_HOURLY_USD, if set, is enforced client-side on both the quote and the launch. It lives in the client config, so the model can't change it through the server.
  4. No auto top-up, ever. add_credit only returns a Stripe link for a human to pay. When a launch fails because the account is frozen or short on balance, the error points at get_balance / add_credit and states that the server never tops up on its own.

Why a confirm token, not dry_run=True

A dry_run flag that defaults to true is only one boolean away from a launch. Nothing stops a model from sending dry_run=false on its first call, and then the user never sees a price. The token makes the quote a required step:

  • The token proves a quote happened. It's random (lq_...), so the only way to get one is the quoting call, and that call's result, with the price, lands in the transcript the user sees. The tool description and the server instructions tell the model to get an explicit yes first.
  • It's bound to the request. Name, GPU, count, region, image, keep-alive, SSH key and price ceiling all have to match, and the launch uses the quoted offer. So a model can't quote a cheap box and launch a different one. A mismatch burns the token.
  • It's single-use and expires after 10 minutes. A stale "yes" can't be replayed hours later at a different price.
  • It works in every MCP client. It's plain tool calls. It doesn't depend on elicitation (which most clients don't support yet) or on the client's per-call approval dialog, which people often switch to "always allow". Where clients do show approval prompts, the annotations still apply on top.

Tokens are held in the server process's memory. A stdio server is one process per client session, which is the right lifetime for a quote. A future version could also use MCP elicitation for the confirm step on clients that support it, and keep the token as the fallback.

Development

cd sdk/mcp
uv sync                        # installs mcp + the local ../python SDK (editable path dependency)
uv run pytest                  # unit tests + a real stdio handshake; no network, no real API
uv run matmul-mcp              # run the server on stdio
npx @modelcontextprotocol/inspector uv run matmul-mcp   # poke at it interactively

The tests mock the HTTP layer (they patch urlopen under the SDK), so each tool runs through the real SDK code and a real MCP client session. They cover every tool's happy path; error mapping (401 means run matmul login or set MATMUL_API_KEY, a 402 or frozen account gives the balance message, the price guard, 403/404/502/503); the quote → confirm flow, including single-use, expiry, argument binding and the env cap; and a stdio subprocess initialize + tools/list handshake.

To try it against a local API with a fake GPU supplier (no accounts, no spend), use apps/api/dev/stage1.sh fake. It builds and runs the API from a temp dir; never run the API from apps/api/. Then point the server at it:

PORT=8081 API_KEYS=lmk_dev=dev apps/api/dev/stage1.sh fake       # terminal 1
MATMUL_API_KEY=lmk_dev MATMUL_API_URL=http://localhost:8081 \
  npx @modelcontextprotocol/inspector uv run --project sdk/mcp matmul-mcp   # terminal 2

The matmul-cloud dependency is a path dependency on ../python for development ([tool.uv.sources]). Installs from PyPI resolve the plain matmul-cloud>=0.1.0 requirement from PyPI.

Releases go out through .github/workflows/release-pypi.yml: bump __version__ in matmul_mcp/__init__.py, add a changelog entry below, merge, then push a matmul-mcp-v<version> tag from main.

Changelog

0.0.1

First release on PyPI: the stdio server with offers, quote → confirm launch, instances, SSH keys and billing tools, the optional MATMUL_MCP_MAX_HOURLY_USD cap, the preview jobs tools behind MATMUL_MCP_ENABLE_JOBS=1, and matmul-mcp --version.

Remote MCP (hosted, OAuth): later

These are design notes only; none of this is built. The next step is a hosted MCP endpoint, so people can add MatMul to claude.ai, ChatGPT or Cursor with a URL, with nothing to install and no key to paste.

  • Endpoint: https://mcp.matmul.cloud/mcp, using the MCP streamable HTTP transport. The same tool set, served from FastMCP(...).run("streamable-http") or ported into the Go API. Stateless HTTP mode lets it scale horizontally.
  • Auth: OAuth 2.1 as the MCP authorization spec describes. The MCP server is a resource server and publishes /.well-known/oauth-protected-resource, pointing at WorkOS AuthKit as the authorization server. AuthKit already runs the console login and supports dynamic client registration, so MCP clients can register themselves. Access tokens carry the user and org, and the server maps them to the same X-Org-Id/X-User-Id identity the console uses, so RBAC (owner/admin/member/viewer) applies per person, the same as in the console.
  • Scopes: split instances:read, instances:write (launch/terminate) and billing:read, with checkout links behind billing:write. That way an org can grant read-only access to an assistant.
  • Confirm tokens move from process memory to a shared store (Postgres or KV), keyed per user, with the same TTL and binding. Where the client supports elicitation, use it for the confirm step.
  • Spend limits move server-side: a per-org "max $/hr via MCP" setting in the console replaces the MATMUL_MCP_MAX_HOURLY_USD env var.
  • Audit: log every launch and terminate made over MCP with the OAuth client id, and show the log in the console.

Release files for matmul-mcp 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for matmul-mcp 0.0.1
File Size Uploaded
matmul_mcp-0.0.1.tar.gz 31.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for matmul-mcp 0.0.1
File Interpreter ABI Platform
matmul_mcp-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 53.4 kB

Release files / matmul_mcp-0.0.1.tar.gz

Download URL matmul_mcp-0.0.1.tar.gz
Size 31.0 kB
Tags Source
SHA-256 checksum
How to use checksums
aae091130040efebfcd0e058589be505bc55a9a9c8a8dae25a4e77c8f4202c5a
BLAKE2b-256 checksum
How to use checksums
784bc5018d5bc98b260287615c97c7cfb0f94c9cb50ca214a4541abcba03d84a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / matmul_mcp-0.0.1-py3-none-any.whl

Download URL matmul_mcp-0.0.1-py3-none-any.whl
Size 22.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0324feaf4e79d642a42a58f1237b1aec871258970b9016ff51533d89823c78bd
BLAKE2b-256 checksum
How to use checksums
413269a2350d73b2475071c4ee2858f575794f062b8aab9aab388a8284c5323e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page