mechbench (the runner)
The mechbench command: what you install on a machine to connect it to
mechbench.ai. The repository keeps its old name
— the PyPI package and the command are mechbench (task 000307), and
the distribution ships two modules: mechbench, the bare front door,
and mechbench_runner, the engine it dispatches into.
The machine-side process of the mechbench family: it claims queued jobs from mechbench-api, executes them against mechbench-compute, and posts results back. It also exposes those same primitives as Model Context Protocol tools, so an LLM agent can call them directly.
Status: in use. login pairs a machine with an account; the runner then claims and executes jobs, reports progress and preparing steps, holds a live WSS channel for control and telemetry, and installs as a launchd or systemd service so it survives reboots. doctor tells you whether a machine will work before it tries. Three MCP tools (run_protocol, get_result, list_jobs) expose the same primitives to an agent.
What this repo is for
Two adjacent surfaces for different callers:
- MCP server. An LLM agent (Claude, others) connects via MCP stdio and calls mechbench primitives as structured tools. Tool bodies run in-process against
mechbench-compute. - Job-runner. Polls
mechbench-api's/jobs/nextfor UI-queued protocols, runs them, posts results back. Same compute path as the MCPrun_protocoltool; different trigger.
Both modes share one binary (mechbench) with subcommands; they share the loaded model, API client, and protocol executor. Splitting into separate processes is a later operational decision — see "Open design questions" below.
Architectural decisions (task 000185)
- Python.
mechbench-computeis Python; delegating to Python via RPC or subprocess-shell from a TS runner adds a layer that pays no dividends in v0. The MCP Python SDK is mature. - One binary, two subcommands.
mechbench mcplaunches the MCP server over stdio;mechbench runstarts the job-runner loop. They shareExperimentRunner(owns the loaded Gemma model) andApiClient. - Agent authenticates to
mechbench-apiwith a dedicated API key, not a user's personal session. ExportMECHBENCH_API_KEY(mint one at/settings/api-keys, or viaPOST /auth/api-keys). Matches the pattern from the e2e trace. - MCP
run_protocolruns in-process, not queued throughmechbench-api. The MCP caller wants the answer; we are the compute target. Job-queue round-tripping exists for the UI-triggered path (job-runner subcommand). - stdio transport only. SSE / HTTP-SSE transports earn their seat once remote MCP deploy matters (deferred).
Install
uv tool install --managed-python mechbench
mechbench login
--managed-python has uv fetch its own interpreter rather than adopt
whichever python3 the machine happens to have. It costs a one-time
download and buys a version we support (3.11–3.14) on a machine whose
own Python we then never touch. pipx install mechbench works
too, against an interpreter you already have.
login prints a link and waits. Open it, approve the machine — the page
names it, along with its host and platform, before you do — and the
runner collects a credential it writes to ~/.mechbench/config.toml
(mode 0600). Nothing durable passes through your hands: the code in the
URL grants nothing on its own, and the key is minted directly to the
machine that asked.
For a machine with no browser, mechbench login --token mbr_…
takes a single-use token minted at mechbench.ai/download.
login then offers to start the runner automatically. Say yes and there
is nothing further to do: it starts at login, comes back after a crash,
and is controlled from the website.
Updating
mechbench update
Upgrades and restarts the service. Re-running the install command does
not upgrade anything — uv tool install treats an already-installed
tool as nothing to do and reports that in a way that reads like success,
so a machine can sit on an old version while looking freshly installed.
update verifies by reading the installed version back afterwards
rather than trusting an exit code, and rolls back if the new version
cannot start.
mechbench doctor answers "will this actually work here" —
Python, backend, credentials, API, model cache, disk — before you find
out the slow way.
Running a model needs Apple Silicon (the MLX backend from
mechbench-compute). The rest installs anywhere.
Running it yourself
mechbench run # foreground, ^C to stop
mechbench install-service # or have the OS keep it running
mechbench service-status
The service is supervised by launchd or systemd rather than by anything
we wrote — see mechbench_runner/exits.py for the contract that makes
that work.
On macOS you will be told that software from "Ned Deily" can run in
the background. That is this runner. macOS attributes a background
item to whoever code-signed the executable, and the executable is the
Python interpreter, which Ned Deily signs as CPython's macOS release
manager. Turning it off in Login Items & Extensions stops the runner;
mechbench doctor reports it if that happens.
From a checkout
git clone https://github.com/mechbench/mechbench-runner.git
cd mechbench
python3.11 -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
Usage
MCP server
Launch as a stdio MCP server — connect from Claude Desktop via claude_desktop_config.json:
{
"mcpServers": {
"mechbench": {
"command": "/abs/path/to/mechbench/.venv/bin/mechbench",
"args": ["mcp"],
"env": {
"MECHBENCH_API_URL": "http://localhost:3000",
"MECHBENCH_API_KEY": "mbk_..."
}
}
}
}
Three tools appear in Claude:
| tool | description |
|---|---|
run_protocol |
Run a layer-ablation protocol in-process on a prompt; return per-layer damage. |
get_result |
Fetch a cached payload from mechbench-api by MechbenchPath. |
list_jobs |
List the caller's queued / running / completed jobs. |
Job-runner
Polls mechbench-api for UI-queued jobs. Same compute path as run_protocol; different trigger.
export MECHBENCH_API_URL=http://localhost:3000
export MECHBENCH_API_KEY=mbk_...
mechbench run
Ctrl-C exits cleanly. API-unreachable is retried with exponential backoff capped at 30 s.
In-process smoke test
mechbench smoke # quick: list_jobs + get_result
mechbench smoke --full # adds run_protocol (42 forwards, ~1-2 min)
Configuration
All via env vars:
| var | default | purpose |
|---|---|---|
MECHBENCH_API_URL |
https://api.mechbench.ai |
mechbench-api base URL. Set it to http://localhost:3000 to develop against a local API. Ignored when credentials are stored, which carry their own. |
MECHBENCH_API_KEY |
(from login) |
Overrides the stored credential entirely, URL included. For CI and containers, which have nowhere to put a config file. |
MECHBENCH_POLL_INTERVAL_SECONDS |
2.0 |
Job-runner poll cadence. |
MECHBENCH_WARM_MODEL_ID |
(none) | Optional model to load at startup so the first job skips cold start. There is deliberately no default: a protocol names the model it runs against, and a job that names none is an error. |
MECHBENCH_WATCHDOG_SECONDS |
900 |
How long without progress counts as wedged. 0 disables it. |
Relationship to other mechbench repos
mechbench-compute— imported directly.Model,Ablate, hook-aware forward.mechbench-schema— producesLayerAblationPayloadetc. as typed results.mechbench-api— the runner's only platform dependency. All workspace state (jobs, cache reads) goes through it.mechbench-ui— no coupling. UI queues jobs; the job-runner consumes them.mechbench-experiments— research scripts that usemechbench-computedirectly, without the job machinery.
Open design questions (deferred)
- One binary or two processes? Current answer: one binary, two subcommands. Revisit if MCP-caller frequency vs. job-runner throughput diverges enough to want independent scaling.
- Structured-summary interface. The family's philosophy doc describes a read-side surface where agents consume JSON summaries of findings / experiments. Currently implicit in
list_jobs+get_result. A richer summary layer (GET /summary,POST /query) is still on the table but unbuilt. - MCP-surface observability. Rate limits, per-tool metrics, audit trail for the tool-calling side. Deferred until a second LLM-agent consumer exists.
License
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mechbench-0.6.0.tar.gz.
File metadata
- Download URL: mechbench-0.6.0.tar.gz
- Upload date:
- Size: 89.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
93643044eed071cb57b735077e590b32d286e80c1254df2b55a83ff11a213bce
|
|
| MD5 |
16db6c0271d5d03a0326e418ce908f0a
|
|
| BLAKE2b-256 |
7248ae82df025847260b7f13b9589a79acb8781110ada8d737ec162cc47a1acf
|
File details
Details for the file mechbench-0.6.0-py3-none-any.whl.
File metadata
- Download URL: mechbench-0.6.0-py3-none-any.whl
- Upload date:
- Size: 76.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7a70ed66483590660d007ca8e1832aef0a9a1698f3c7c24a9dd8db8dd1723c4c
|
|
| MD5 |
62d8483402f7bd85f00f03fdf046f80a
|
|
| BLAKE2b-256 |
93d773e75ae4b7eef4e3eeb2f8b746412ba29abc5578eca4965051f0f4ebb719
|