icefold-runner
A self-hosted execution runner for IceFold nodes — like a GitHub self-hosted CI runner. You start it on your own machine; it reverse-connects to an IceFold server (so it works behind NAT with no inbound ports, no public IP, no tunnel), receives node-execution jobs, runs them locally, and streams results back.
It is the place where your uploaded node code runs — on your hardware, with
full subprocess / ffmpeg / GPU / any-dependency access. The IceFold server
never executes third-party code; it only renders it into bundles for a runner.
How it works
your machine (private, behind NAT) IceFold server (public)
┌──────────────────────────────────┐ reverse WSS ┌───────────────────────────┐
│ icefold-runner │ ───────────► │ /v1/ws/worker?worker_id=… │
│ • dials out, token auth │ node_exec ◄─│ routes node runs (per user) │
│ • reconnect + keepalive │ node_done ─►│ │
│ • bundle runner: │ │ │
│ GET /v1/bundles/<hash> │ HTTP pull │ /files /scratch │
│ import bundle + preflight deps │ ◄──────────► │ /v1/workers/output │
│ await __icefold_run__ │ │ │
└──────────────────────────────────┘ HTTP push └───────────────────────────┘
- Control plane rides the reverse WebSocket (
node_exec/cancel→node_status/node_done/missing_dep) as plain JSON text frames — TLS (wss) is the confidentiality layer. The runner token travels in theX-Worker-Tokenrequest header; only the non-secretworker_idis in the query string. Eachnode_execframe carries abundle_hash, the selected variant's inputs, and the confirmed variant input storage needed by the bundle's typed runtime — never node source. - Bulk media + bundles ride plain HTTP: the runner GETs inputs from the
server's
/filesand/scratchmounts and node bundles from/v1/bundles/<hash>(sha256-addressed, cached locally asrunner_work_dir/bundles/<hash>.py, re-hashed on every download), runs the bundle, POSTs products back to/v1/workers/output(which returns server-canonical paths and accepts at most 2 GiB per product). - Media inputs are cached across calls. Signed query strings rotate, so the
cache key uses the server origin, authenticated owner, and immutable canonical
/filesor/scratchartifact path. A cold call downloads distinct inputs concurrently and atomically publishes each blob; concurrent calls requesting the same blob share one transfer. Warm calls use the local file without a network request. Active inputs carry leases and cannot be evicted while node code or a subprocess is using them. After a product upload succeeds, the returned canonical path is atomically linked to that local product, so a downstream node on the same runner does not download it back from the server. Age plus LRU byte ceilings bound disk use. - The runner ships no node implementations and never compiles user source.
The IceFold server renders every node (your custom ones and the platform's
built-in ones) into a self-contained
.pybundle, withpython_deps/binary_depsdeclared in the bundle header. The runner imports the bundle, pre-flights the deps (sending back a structuredmissing_depreply with platform-aware install hints if anything is absent), and awaits__icefold_run__(inputs, ctx_dict). So when the server adds or upgrades nodes, the runner does not need an upgrade. Runner protocol or agent changes are released as a new runner version. - Variant planning / dimension & provider resolution all stay on the server; each job is a single already-sliced leaf call.
- Transient WebSocket loss does not restart a node. Recovery-capable server
and runner versions negotiate this in
worker_ready: the runner keeps the task alive, the server reattaches its pendingcall_id, and terminal frames plus host callbacks replay by id. If the runner process itself is lost, the server automatically resubmits the samecall_idto a replacement runner. A stableworker_idowns scheduling/cache affinity, while a random per-processinstance_idensures a restarted process can never impersonate a socket reconnect and inherit coroutines that died with its predecessor. Older peers do not advertise recovery, so rolling upgrades retain the legacy cancel-on-disconnect behaviour instead of guessing across protocol versions. - Pre-execution transfer failures can move once to a different runner. The runner retries transient input and bundle downloads locally first. If those attempts are exhausted before node code starts, runner 0.2.7+ marks the terminal result as safe to retry; a compatible server may then exclude that runner and dispatch one fresh call to a peer. Structured dependency-preflight failures are safe for the same reason. Node, provider, timeout, upload, and other post-start failures never carry this hint.
Install
Requires only Python ≥ 3.11 (it pulls in icefold-sdk).
The runner itself ships no node tooling. A node declares what it needs in its
bundle header — binary_deps (for example, IceFold video nodes use
google-chrome, ffmpeg, and ffprobe on PATH) and python_deps (whatever
your custom nodes import) — and the runner
pre-flights those before each run, replying with a platform-aware install
hint (missing_dep) for anything absent. So you install a node's dependencies
only when you actually run a job that needs them, and the runner tells you
exactly what to install.
pip install icefold-runner # pulls in icefold-sdk
From source:
git clone <this-repo> icefold-runner
cd icefold-runner
python -m venv .venv && . .venv/bin/activate
pip install -e .
Run
Generate a token in the IceFold app (Settings → Runners), then:
install -m 600 /dev/null ~/.icefold-runner-token
read -rsp 'Runner token: ' ICEFOLD_TOKEN_INPUT
printf '%s' "$ICEFOLD_TOKEN_INPUT" > ~/.icefold-runner-token
unset ICEFOLD_TOKEN_INPUT
icefold-runner --token-file ~/.icefold-runner-token
That's it — the token (GitHub-CI style) encodes + signs your IceFold user id, so there's no server URL or user id to pass. The server is built in.
Every flag also reads an env var:
| flag | env | meaning |
|---|---|---|
--token-file |
ICEFOLD_RUNNER_TOKEN_FILE |
path to a mode-0600 runner token file |
| — | ICEFOLD_RUNNER_TOKEN |
runner token via environment (the insecure --token flag is rejected) |
--runner-id |
ICEFOLD_RUNNER_ID |
stable scheduling/cache identity (default: fresh id; set this when the work dir is durable) |
--work-dir |
ICEFOLD_RUNNER_DIR |
persistent input cache and product scratch root |
--input-cache-max-age |
ICEFOLD_RUNNER_INPUT_CACHE_MAX_AGE |
evict inputs unused for this long (default: 7d) |
--input-cache-max-size |
ICEFOLD_RUNNER_INPUT_CACHE_MAX_SIZE |
maximum persistent input-cache bytes (default: 20GiB) |
--concurrency |
ICEFOLD_RUNNER_CONCURRENCY |
CPU-lane slots for ffmpeg/Pillow work (default: detected CPUs, capped at 8) |
--gpu-concurrency |
ICEFOLD_RUNNER_GPU_CONCURRENCY |
GPU-lane slots (default: 1) |
The runner honors standard proxy env vars (HTTPS_PROXY, ALL_PROXY, …) for
reaching the server, including HTTP and SOCKS proxies. It reconnects
automatically with backoff; an auth rejection is fatal.
The runner advertises both effective lane widths in its WebSocket hello frame.
Servers that understand these fields can weight stable machine ownership and
bound each runner's local CPU/GPU queues. Session work stays on its cache-warm
owner when those queues fill; a designated backup is used only when no primary
is healthy. Older servers safely ignore the extra fields. A newer server treats
an older runner that does not advertise capacity as one slot.
Self-hosting / dev: point the runner at a different server with the
ICEFOLD_RUNNER_SERVERenv var (e.g.ws://127.0.0.1:7000).
Layout
icefold_runner/ the runner agent (connection, input cache, bundle exec)
client.py reverse-WS client: dial / auth / reconnect / keepalive
runner.py fetch /v1/bundles/<hash>, preflight deps, await __icefold_run__
__main__.py CLI entrypoint (icefold-runner)
The runner imports the bundle on demand; the bundle is self-contained and
already inlines whatever it needs (the author's function body, the
Inputs / Output dataclasses, and a minimal NodeContext shim). The only
runtime dependency on icefold-sdk is the wire protocol + a small helper kit
(get_file_id / run_blocking / write_text), used by the runner agent
itself, not by node code.
Security model
- Node code runs unsandboxed here — it's your machine, your risk. That's the point: code the server refuses to execute (subprocess/ffmpeg/native deps and anything third-party) runs on the runner instead. The runner downloads each bundle from the server and executes it; it verifies the bundle's sha256 matches the requested hash, but the bundle itself is whatever the server you authenticated to sends. Only point a runner at a server you trust.
- The runner only talks to the one server you point it at, authenticated by its runner token; it pulls input files and pushes products over HTTP to that host.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file icefold_runner-0.3.2.tar.gz.
File metadata
- Download URL: icefold_runner-0.3.2.tar.gz
- Upload date:
- Size: 37.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
615dfc4c33f84470406f2d49edca6f44ecab8f511c0f4b8dd77bb98ec92ba4df
|
|
| MD5 |
dda9376a7cd79a19d3da57bb3e864fc9
|
|
| BLAKE2b-256 |
1075647756bf3f09942021a7d9d3314dbb6765dd1f9c526a481323808989349b
|
File details
Details for the file icefold_runner-0.3.2-py3-none-any.whl.
File metadata
- Download URL: icefold_runner-0.3.2-py3-none-any.whl
- Upload date:
- Size: 36.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
787bca180764b0f843145a6b8bc3f8397152198be47409e8de942f0c2fdcff36
|
|
| MD5 |
1cf7d48e4ad7c57ab51239e6cae72d3a
|
|
| BLAKE2b-256 |
72af46ce38817f88ebe799c68d1ab36c36c37de83c46a73258bef1b3d476fc8e
|