Skip to main content

krauncher-mcp

An MCP server that gives an agent one tool: a pre-run cost estimate for a GPU task, from static analysis of the code. The code is never executed.

It wraps the Krauncher analyzer (the assay), not the broker — there is no dispatch, no execution, no market lineup. Just: how long will this cost, and what does it need to run.

The seconds are a relative signal for comparing code against code, not an absolute forecast. They are normalized to a fixed reference card (RTX PRO 6000 WS) so two estimates are comparable; the GPU and host your code actually runs on will differ, so never read a second-count as the wall-clock you will get. Compare variant-to-variant. This is an early 0.x release — the model is approximate and evolving.

The estimate_gpu_time_and_cost tool

Input: code — the task's Python source (a self-contained function; a @client.task-decorated function is fine).

Output, on the reference card (RTX PRO 6000 WS, the card CU is normalized to):

{
  "reference_card": "RTX PRO 6000 WS",
  "compute_sec": 20.2,      // the three phases of wall time on the ref card
  "setup_sec": 3.0,
  "io_sec": 2.1,
  "min_vram_gb": 6,         // raw requirement, no headroom margin
  "min_disk_gb": 10,
  "confidence": 1.0,        // 0-1
  "analysis_method": "ast", // "ast" | "llm"
  "cpu_only": false,
  "spread": 1.51,           // slow end of the measured host population
  "spread_reason": "1.51x { cv_training, nw=one } on the compute phase, observed on 42 runs across 16 hosts (worst 2.45x)",
  "calibration_basis": "calibrated",  // | "extrapolated" | "uncalibrated"
  "knobs": [                // the run parameters worth re-estimating
    { "name": "num_workers", "value": null, "same_work": true },
    { "name": "batch_size",  "value": "16", "same_work": true },
    { "name": "num_epochs",  "value": "1",  "same_work": false }
  ],
  "findings": [             // what the analyzer read from the code
    "num_epochs=1", "batch_size=16",
    "Recognized model: BERT Base (0.11B params)",
    "precision=fp16 from fp16=True"
  ]
}

The loop it is built for: edit the run → estimate → keep what's cheaper → repeat, all before spending a GPU-second. The estimate is a static forecast, not a guarantee; confidence and analysis_method say how much to trust it, and a rough estimate never blocks — it returns a best effort.

knobs is what turns that loop from guesswork into a shortlist. The parameters that move the time without changing what the code produces are listed every time, found or not; only the values are meant to change, since restructuring the job (a smaller model, a different architecture, less data) is not what the number is for.

  • value: null — the analyzer did not see that parameter in the source. It arrives as a call argument, from a config, or from the environment, and nothing can price it until the code states it as a literal.
  • same_work: false — epochs, steps, sequence length. These shrink the job itself, so a lower number is a different task rather than a cheaper one. They appear only when the code actually sets them.

spread and calibration_basis answer a different question than confidence. Confidence is about the reading of the code; spread is about the world the code will run in — the same source on the same card lands over a range of hosts, and 1.51 means the slow end of that measured population takes about half again as long as the estimate. Both can be high at once, which is the honest description of a job whose time the GPU does not govern: the card waits on the host, so the run inherits whichever host it lands on. The lever there is GPU utilization — whatever leaves the card idle in this code (data loaded in the main process, per-item preprocessing, synchronous transfers, a batch too small to fill the card). spread_reason names the population the number came from and how many runs it rests on; calibration_basis says whether this shape was measured at all (calibrated), answered by a neighbour (extrapolated), or matched nothing (uncalibrated — read the seconds as an order of magnitude).

What it does not return: the cost model's calibration coefficients or weights. Only what the analyzer detected in the code leaves the server.

Install

pip install -e .        # from this directory; also installs the analyzer client

An API key is optional. Without one the server calls the public analyzer keyless, under a per-IP daily quota (when the quota is reached, the tool returns a short note to register for a larger one). Set a key to use your own account and skip the quota:

export KRAUNCHER_API_KEY=cas_...   # optional

Verify it works without wiring up a client — runs the tool on a sample task and prints the contract:

krauncher-mcp --selftest

Wire it into an MCP client

stdio transport; the console script is krauncher-mcp. No key needed — this runs keyless against the public analyzer:

{
  "mcpServers": {
    "gpu-estimator": {
      "command": "krauncher-mcp"
    }
  }
}

To use your own account (keyed, exempt from the per-IP quota), add the key:

{
  "mcpServers": {
    "gpu-estimator": {
      "command": "krauncher-mcp",
      "env": { "KRAUNCHER_API_KEY": "cas_..." }
    }
  }
}

Self-hosting the analyzer? Override the endpoint with KRAUNCHER_ANALYZER_URL.

Nothing else to configure for discovery: the tool ships marked anthropic/alwaysLoad, so on hosts that defer MCP schemas behind a tool search (Claude Code does this by default) it is in context from the first turn rather than waiting to be searched for. One tool, one small schema — that is the whole budget it spends.

Scope

v1 is deliberately one tool. The per-GPU market lineup is intentionally left out — the agent's job is to improve its code and know the cost before running, and a pre-run estimate (even rough or partial) is the whole point.

Release files for krauncher-mcp 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for krauncher-mcp 0.4.1
File Size Uploaded
krauncher_mcp-0.4.1.tar.gz 9.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for krauncher-mcp 0.4.1
File Interpreter ABI Platform
krauncher_mcp-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 19.5 kB

Release files / krauncher_mcp-0.4.1.tar.gz

Download URL krauncher_mcp-0.4.1.tar.gz
Size 9.2 kB
Tags Source
SHA-256 checksum
How to use checksums
d72d07834f15a1f9eb0a75b51621a92d1097aa42690df3df4aef6689fbc837e7
BLAKE2b-256 checksum
How to use checksums
473557a1677e69b710b3cef7240a0d820b7231a063df579ae9ae3b3f513f7c2b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / krauncher_mcp-0.4.1-py3-none-any.whl

Download URL krauncher_mcp-0.4.1-py3-none-any.whl
Size 10.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cadf081af80fdeccbdadee77b5bd79189ad0daf5e42530eed7a7e3a51806d524
BLAKE2b-256 checksum
How to use checksums
be7aa71633ae96df4e38c205cb24c4e1a67b705bc28fc78f25e519957bafb94a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

0.5.0

2 release files

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page