mechbench-runner
The machine-side process of the mechbench family: it claims queued jobs from mechbench-api, executes them against mechbench-compute, and posts results back. It also exposes those same primitives as Model Context Protocol tools, so an LLM agent can call them directly.
Status: in use. login pairs a machine with an account; the runner then claims and executes jobs, reports progress and preparing steps, holds a live WSS channel for control and telemetry, and installs as a launchd or systemd service so it survives reboots. doctor tells you whether a machine will work before it tries. Three MCP tools (run_protocol, get_result, list_jobs) expose the same primitives to an agent.
What this repo is for
Two adjacent surfaces for different callers:
- MCP server. An LLM agent (Claude, others) connects via MCP stdio and calls mechbench primitives as structured tools. Tool bodies run in-process against
mechbench-compute. - Job-runner. Polls
mechbench-api's/jobs/nextfor UI-queued protocols, runs them, posts results back. Same compute path as the MCPrun_protocoltool; different trigger.
Both modes share one binary (mechbench-runner) with subcommands; they share the loaded model, API client, and protocol executor. Splitting into separate processes is a later operational decision — see "Open design questions" below.
Architectural decisions (task 000185)
- Python.
mechbench-computeis Python; delegating to Python via RPC or subprocess-shell from a TS runner adds a layer that pays no dividends in v0. The MCP Python SDK is mature. - One binary, two subcommands.
mechbench-runner mcplaunches the MCP server over stdio;mechbench-runner runstarts the job-runner loop. They shareExperimentRunner(owns the loaded Gemma model) andApiClient. - Agent authenticates to
mechbench-apiwith a dedicated API key, not a user's personal session. ExportMECHBENCH_API_KEY(mint one at/settings/api-keys, or viaPOST /auth/api-keys). Matches the pattern from the e2e trace. - MCP
run_protocolruns in-process, not queued throughmechbench-api. The MCP caller wants the answer; we are the compute target. Job-queue round-tripping exists for the UI-triggered path (job-runner subcommand). - stdio transport only. SSE / HTTP-SSE transports earn their seat once remote MCP deploy matters (deferred).
Install
uv tool install mechbench-runner # or: pipx install mechbench-runner
mechbench-runner login
login prints a link, takes the registration token from it, stores a
durable key at ~/.mechbench/config.toml (mode 0600), and offers to
start the runner automatically. Say yes and there is nothing further to
do: it starts at login, comes back after a crash, and is controlled from
the website.
mechbench-runner doctor answers "will this actually work here" —
Python, backend, credentials, API, model cache, disk — before you find
out the slow way.
Running a model needs Apple Silicon (the MLX backend from
mechbench-compute). The rest installs anywhere.
Running it yourself
mechbench-runner run # foreground, ^C to stop
mechbench-runner install-agent # or have the OS keep it running
mechbench-runner agent-status
The service is supervised by launchd or systemd rather than by anything
we wrote — see mechbench_runner/exits.py for the contract that makes
that work.
From a checkout
git clone https://github.com/mechbench/mechbench-runner.git
cd mechbench-runner
python3.11 -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
Usage
MCP server
Launch as a stdio MCP server — connect from Claude Desktop via claude_desktop_config.json:
{
"mcpServers": {
"mechbench": {
"command": "/abs/path/to/mechbench-runner/.venv/bin/mechbench-runner",
"args": ["mcp"],
"env": {
"MECHBENCH_API_URL": "http://localhost:3000",
"MECHBENCH_API_KEY": "mbk_..."
}
}
}
}
Three tools appear in Claude:
| tool | description |
|---|---|
run_protocol |
Run a layer-ablation protocol in-process on a prompt; return per-layer damage. |
get_result |
Fetch a cached payload from mechbench-api by MechbenchPath. |
list_jobs |
List the caller's queued / running / completed jobs. |
Job-runner
Polls mechbench-api for UI-queued jobs. Same compute path as run_protocol; different trigger.
export MECHBENCH_API_URL=http://localhost:3000
export MECHBENCH_API_KEY=mbk_...
mechbench-runner run
Ctrl-C exits cleanly. API-unreachable is retried with exponential backoff capped at 30 s.
In-process smoke test
mechbench-runner smoke # quick: list_jobs + get_result
mechbench-runner smoke --full # adds run_protocol (42 forwards, ~1-2 min)
Configuration
All via env vars:
| var | default | purpose |
|---|---|---|
MECHBENCH_API_URL |
http://localhost:3000 |
mechbench-api base URL. Ignored when credentials are stored, which carry their own. |
MECHBENCH_API_KEY |
(from login) |
Overrides the stored credential entirely, URL included. For CI and containers, which have nowhere to put a config file. |
MECHBENCH_POLL_INTERVAL_SECONDS |
2.0 |
Job-runner poll cadence. |
MECHBENCH_WARM_MODEL_ID |
(none) | Optional model to load at startup so the first job skips cold start. There is deliberately no default: a protocol names the model it runs against, and a job that names none is an error. |
MECHBENCH_WATCHDOG_SECONDS |
900 |
How long without progress counts as wedged. 0 disables it. |
Relationship to other mechbench repos
mechbench-compute— imported directly.Model,Ablate, hook-aware forward.mechbench-schema— producesLayerAblationPayloadetc. as typed results.mechbench-api— the runner's only platform dependency. All workspace state (jobs, cache reads) goes through it.mechbench-ui— no coupling. UI queues jobs; the job-runner consumes them.mechbench-experiments— research scripts that usemechbench-computedirectly, without the job machinery.
Open design questions (deferred)
- One binary or two processes? Current answer: one binary, two subcommands. Revisit if MCP-caller frequency vs. job-runner throughput diverges enough to want independent scaling.
- Structured-summary interface. The family's philosophy doc describes a read-side surface where agents consume JSON summaries of findings / experiments. Currently implicit in
list_jobs+get_result. A richer summary layer (GET /summary,POST /query) is still on the table but unbuilt. - MCP-surface observability. Rate limits, per-tool metrics, audit trail for the tool-calling side. Deferred until a second LLM-agent consumer exists.
License
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mechbench_runner-0.1.2.tar.gz.
File metadata
- Download URL: mechbench_runner-0.1.2.tar.gz
- Upload date:
- Size: 58.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e539c9c367043762d2e24a8ad30b285f576f90886b903ea568e4f9ca078b0a02
|
|
| MD5 |
840e836467a3f8f654afbfd233bd8b65
|
|
| BLAKE2b-256 |
f4a0c906a3f8b715f32efdb52f1df34583b9765a59ec7ee17c2eb5dcd03ed48f
|
File details
Details for the file mechbench_runner-0.1.2-py3-none-any.whl.
File metadata
- Download URL: mechbench_runner-0.1.2-py3-none-any.whl
- Upload date:
- Size: 51.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
df76b9c107c538234ae873f4a74e686a090950c33dac0fdcb0fd48748b6a8838
|
|
| MD5 |
d1bf10cc03807eea08e8c36d771c455c
|
|
| BLAKE2b-256 |
b011673400994a09fe96e3cf06c62d75514f55a93a81a7ed15a03d8fc3722de4
|