matmul-mcp
An MCP server for MatMul, so you can find and launch GPU compute from Claude Code, Claude Desktop, Cursor, VS Code or any other MCP client. Ask things like "find me the cheapest H100 under $3/hr and give me an SSH command", and the assistant drives the MatMul API for you. Every launch needs your explicit OK on a price quote first.
It runs locally over stdio and talks to https://api.matmul.cloud/v1 with your API key. It is
built on the stdlib-only matmul-cloud Python SDK (import name matmul) and the official mcp Python SDK
(FastMCP).
Install
You need an API key. Mint one in the console at https://app.matmul.cloud/app/cli.
uvx matmul-mcp # run without installing (recommended)
pipx install matmul-mcp # or install the `matmul-mcp` command
pip install matmul-mcp also installs the matmul CLI (from matmul-cloud), so matmul login
works too.
Authentication is resolved in this order:
- the
MATMUL_API_KEYenvironment variable; - the key saved by
matmul login --api-key <key>(~/.config/matmul/config.json, or$MATMUL_CONFIG).
| Variable | Meaning |
|---|---|
MATMUL_API_KEY |
API key (Authorization: Bearer ...). |
MATMUL_API_URL |
API base URL. Defaults to https://api.matmul.cloud; a trailing /v1 is accepted. |
MATMUL_MCP_MAX_HOURLY_USD |
Optional hard cap: this server refuses to quote or launch any instance above this $/hr, whatever the model asks for. Also caps find_gpus. |
MATMUL_MCP_ENABLE_JOBS |
1 adds the preview managed-jobs tools (off by default: see below). |
The product was previously called Lemnos: the old LEMNOS_* variable names still work when the
MATMUL_* one is unset, and a key saved in ~/.config/lemnos/config.json is still read.
Configure your client
In every snippet below, replace lmk_... with your key. You can leave MATMUL_API_KEY out if
you've run matmul login on this machine. The cap is optional but recommended.
Claude Code
claude mcp add matmul \
-e MATMUL_API_KEY=lmk_... \
-e MATMUL_MCP_MAX_HOURLY_USD=5 \
-- uvx matmul-mcp
Add --scope user to make it available in every project.
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or
%APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"matmul": {
"command": "uvx",
"args": ["matmul-mcp"],
"env": {
"MATMUL_API_KEY": "lmk_...",
"MATMUL_MCP_MAX_HOURLY_USD": "5"
}
}
}
}
If Claude Desktop can't find uvx, use its absolute path (which uvx).
Cursor
~/.cursor/mcp.json (global) or .cursor/mcp.json (per project):
{
"mcpServers": {
"matmul": {
"command": "uvx",
"args": ["matmul-mcp"],
"env": {
"MATMUL_API_KEY": "lmk_...",
"MATMUL_MCP_MAX_HOURLY_USD": "5"
}
}
}
}
VS Code (GitHub Copilot agent mode)
.vscode/mcp.json. VS Code prompts for the key once and stores it securely:
{
"inputs": [
{ "type": "promptString", "id": "matmul-key", "description": "MatMul API key", "password": true }
],
"servers": {
"matmul": {
"type": "stdio",
"command": "uvx",
"args": ["matmul-mcp"],
"env": {
"MATMUL_API_KEY": "${input:matmul-key}",
"MATMUL_MCP_MAX_HOURLY_USD": "5"
}
}
}
}
What it exposes
Tools
| Tool | Does | Annotations |
|---|---|---|
find_gpus(gpu?, max_hourly_usd?, region?, count?, limit) |
Live GPU offers, cheapest first | read-only |
get_balance(include_ledger?, ledger_limit?) |
Prepaid balance, frozen state, new-account limits, burn rate, runway | read-only |
add_credit(amount_usd) |
Returns a Stripe checkout URL for the human to open and pay; charges nothing itself | write |
list_instances(include_deleted?) |
Instances with status, price, accrued cost, SSH commands | read-only |
get_instance(id_or_name) |
One instance; poll it until running to get ssh_command |
read-only |
launch_instance(name, gpu, max_hourly_usd, image?, keep_alive?, ssh_key?, count?, region?, confirm_token?) |
Two-step: quote, then launch (see below) | write, spends money |
terminate_instance(id_or_name) |
Delete an instance and stop billing | destructive |
list_ssh_keys() / add_ssh_key(name, public_key) |
Manage the keys installed on new machines. add_ssh_key refuses anything that looks like a private key. |
read-only / write |
With MATMUL_MCP_ENABLE_JOBS=1: run_job, list_jobs, job_logs, cancel_job (destructive).
These are opt-in because the /v1/jobs API is still a preview and has no price guard yet. Turn
them on by default once it has one.
Every read tool is marked readOnlyHint, and terminate_instance and cancel_job are marked
destructiveHint, so clients that auto-approve reads still ask before destructive calls.
Resources and prompts
matmul://instances: your non-deleted instances, as JSON.matmul://balance: balance and frozen state, as JSON.- Prompt
gpu_dev_box(gpu, max_hourly_usd, image): "launch a GPU dev box". It walks the assistant through find, add key, quote, confirm, wait and SSH.
Money safety
launch_instance rents a real machine that bills every hour until it is terminated. The
safeguards:
max_hourly_usdis required. It's sent as the API'smax_hourly_price_centsprice guard. The API refuses to launch if the live price has risen above it, and the tool reports "Price guard: nothing was launched and nothing was charged".- Two-step confirm (quote → token → launch). The first call launches nothing. It picks the
cheapest live offer at or under the ceiling and returns a quote: offer, estimated $/hr, the
ceiling, image or plain VM, SSH key, balance and runway, and a line saying it bills until
terminated. It also returns a
confirm_token. Only a second call with that token and identical arguments launches. The launch result repeats the hourly cost and says to callterminate_instance. MATMUL_MCP_MAX_HOURLY_USD, if set, is enforced client-side on both the quote and the launch. It lives in the client config, so the model can't change it through the server.- No auto top-up, ever.
add_creditonly returns a Stripe link for a human to pay. When a launch fails because the account is frozen or short on balance, the error points atget_balance/add_creditand states that the server never tops up on its own.
Why a confirm token, not dry_run=True
A dry_run flag that defaults to true is only one boolean away from a launch. Nothing stops a
model from sending dry_run=false on its first call, and then the user never sees a price. The
token makes the quote a required step:
- The token proves a quote happened. It's random (
lq_...), so the only way to get one is the quoting call, and that call's result, with the price, lands in the transcript the user sees. The tool description and the server instructions tell the model to get an explicit yes first. - It's bound to the request. Name, GPU, count, region, image, keep-alive, SSH key and price ceiling all have to match, and the launch uses the quoted offer. So a model can't quote a cheap box and launch a different one. A mismatch burns the token.
- It's single-use and expires after 10 minutes. A stale "yes" can't be replayed hours later at a different price.
- It works in every MCP client. It's plain tool calls. It doesn't depend on elicitation (which most clients don't support yet) or on the client's per-call approval dialog, which people often switch to "always allow". Where clients do show approval prompts, the annotations still apply on top.
Tokens are held in the server process's memory. A stdio server is one process per client session, which is the right lifetime for a quote. A future version could also use MCP elicitation for the confirm step on clients that support it, and keep the token as the fallback.
Development
cd sdk/mcp
uv sync # installs mcp + the local ../python SDK (editable path dependency)
uv run pytest # unit tests + a real stdio handshake; no network, no real API
uv run matmul-mcp # run the server on stdio
npx @modelcontextprotocol/inspector uv run matmul-mcp # poke at it interactively
The tests mock the HTTP layer (they patch urlopen under the SDK), so each tool runs through
the real SDK code and a real MCP client session. They cover every tool's happy path; error
mapping (401 means run matmul login or set MATMUL_API_KEY, a 402 or frozen account gives the
balance message, the price guard, 403/404/502/503); the quote → confirm flow, including
single-use, expiry, argument binding and the env cap; and a stdio subprocess
initialize + tools/list handshake.
To try it against a local API with a fake GPU supplier (no accounts, no spend), use
apps/api/dev/stage1.sh fake. It builds and runs the API from a temp dir; never run the API from
apps/api/. Then point the server at it:
PORT=8081 API_KEYS=lmk_dev=dev apps/api/dev/stage1.sh fake # terminal 1
MATMUL_API_KEY=lmk_dev MATMUL_API_URL=http://localhost:8081 \
npx @modelcontextprotocol/inspector uv run --project sdk/mcp matmul-mcp # terminal 2
The matmul-cloud dependency is a path dependency on ../python for development ([tool.uv.sources]).
Installs from PyPI resolve the plain matmul-cloud>=0.1.0 requirement from PyPI.
Releases go out through .github/workflows/release-pypi.yml: bump __version__ in
matmul_mcp/__init__.py, add a changelog entry below, merge, then push a matmul-mcp-v<version>
tag from main.
Changelog
0.0.1
First release on PyPI: the stdio server with offers, quote → confirm launch, instances, SSH keys
and billing tools, the optional MATMUL_MCP_MAX_HOURLY_USD cap, the preview jobs tools behind
MATMUL_MCP_ENABLE_JOBS=1, and matmul-mcp --version.
Remote MCP (hosted, OAuth): later
These are design notes only; none of this is built. The next step is a hosted MCP endpoint, so people can add MatMul to claude.ai, ChatGPT or Cursor with a URL, with nothing to install and no key to paste.
- Endpoint:
https://mcp.matmul.cloud/mcp, using the MCP streamable HTTP transport. The same tool set, served fromFastMCP(...).run("streamable-http")or ported into the Go API. Stateless HTTP mode lets it scale horizontally. - Auth: OAuth 2.1 as the MCP authorization spec describes. The MCP server is a resource server
and publishes
/.well-known/oauth-protected-resource, pointing at WorkOS AuthKit as the authorization server. AuthKit already runs the console login and supports dynamic client registration, so MCP clients can register themselves. Access tokens carry the user and org, and the server maps them to the sameX-Org-Id/X-User-Ididentity the console uses, so RBAC (owner/admin/member/viewer) applies per person, the same as in the console. - Scopes: split
instances:read,instances:write(launch/terminate) andbilling:read, with checkout links behindbilling:write. That way an org can grant read-only access to an assistant. - Confirm tokens move from process memory to a shared store (Postgres or KV), keyed per user, with the same TTL and binding. Where the client supports elicitation, use it for the confirm step.
- Spend limits move server-side: a per-org "max $/hr via MCP" setting in the console replaces
the
MATMUL_MCP_MAX_HOURLY_USDenv var. - Audit: log every launch and terminate made over MCP with the OAuth client id, and show the log in the console.
Release files for matmul-mcp 0.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| matmul_mcp-0.0.1.tar.gz | 31.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| matmul_mcp-0.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.4 kB
Release files / matmul_mcp-0.0.1.tar.gz
| Download URL | matmul_mcp-0.0.1.tar.gz |
|---|---|
| Size | 31.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
aae091130040efebfcd0e058589be505bc55a9a9c8a8dae25a4e77c8f4202c5a
|
|
BLAKE2b-256 checksum How to use checksums |
784bc5018d5bc98b260287615c97c7cfb0f94c9cb50ca214a4541abcba03d84a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / matmul_mcp-0.0.1-py3-none-any.whl
| Download URL | matmul_mcp-0.0.1-py3-none-any.whl |
|---|---|
| Size | 22.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0324feaf4e79d642a42a58f1237b1aec871258970b9016ff51533d89823c78bd
|
|
BLAKE2b-256 checksum How to use checksums |
413269a2350d73b2475071c4ee2858f575794f062b8aab9aab388a8284c5323e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|