Run training, fine-tuning, batch experiments, and tool-driven agent sessions when your local machine lacks the GPU memory or capacity they need. Provide a container image and resource requirements. Nodus matches the work to available GPU capacity. Add a budget to set a spending limit.
1. Install and sign in
pip install nodus-compute
nodus login
Get nodus-compute on PyPI.
Requires Python 3.10 or newer. Upgrading an existing installation? Use
pip install --upgrade nodus-compute. These docs cover SDK 0.4.x.
Your browser opens Nodus sign-in. Sign in and approve the code matching your terminal. You can then close the tab. The terminal finishes automatically and saves your credentials. Python clients use that login without extra setup.
Running nodus login again reuses a valid login. Use nodus login --force for
a fresh sign-in.
For a machine without a browser, use nodus login --no-browser.
See authentication for API keys and
custom deployments.
Before starting a workload, open Billing and add a payment method. New accounts start with $30 in credits, but a card is required to use them. Adding a card does not purchase credits. If you joined a shared workspace, its administrator manages the payment method.
2. Run your first workload
This GPU smoke test prints the available GPU name. No local script is uploaded. It submits paid compute with a $5 workload budget. Available capacity and account limits still determine admission.
Save this as first_workload.py:
import nodus
with nodus.Client() as client:
workload = client.run(
image="pytorch/pytorch:2.8.0-cuda12.8-cudnn9-runtime",
command=[
"python", "-c",
"import torch\n"
"assert torch.cuda.is_available()\n"
"print(torch.cuda.get_device_name(0))",
],
budget=5,
)
print("Workload:", workload.id)
done = workload.wait()
print(done.status, done.cost_now_usd)
if not done.succeeded:
raise RuntimeError(f"Workload {done.id} ended: {done.status}")
print(done.logs())
Run it with python first_workload.py. It prints the workload ID and shows live logs, lifecycle events, and elapsed
time while waiting. Training workloads also show reported steps or epochs.
The final output includes status, current cost, and GPU name.
run() accepts the workload. wait() waits for a terminal status, so check
succeeded before using results. Ctrl+C while waiting requests cancellation
and remote resource cleanup.
The script prints the GPU name from the workload logs. For files produced by your own program, see logs and results.
Run an agent in a sandbox
A sandbox is a durable execution environment with its own public API. It is separate from a training or batch workload. Name it once, execute multiple commands, stream ordered stdout and stderr frames, and reconnect with the same name from another process.
import nodus
with nodus.Sandbox(
name="research-agent",
image="pytorch/pytorch:2.8.0-cuda12.8-cudnn9-runtime",
requirements={"gpu": "L40S", "peak_memory_gb": 32},
budget=5,
) as sandbox:
print("Sandbox:", sandbox.id)
process = sandbox.exec("python -c \"print('agent tool finished')\"")
for frame in process.iter_output():
print(frame.stream, frame.text, end="")
done = process.wait()
if not done.succeeded:
raise RuntimeError(f"Command ended: {done.state}")
Creating a sandbox can start paid infrastructure. Nodus checks the account
payment method, account headroom, and sandbox budget before billable placement.
Calling nodus.Sandbox(name="research-agent") reconnects to the named sandbox.
See the sandbox guide.
The CLI mirrors the same resource and verbs.
nodus sandbox new pytorch/pytorch:2.8.0-cuda12.8-cudnn9-runtime --name research-agent --budget 5
nodus sandbox exec SANDBOX_ID "python -c 'print(2 + 2)'"
nodus sandbox cost SANDBOX_ID
nodus sandbox rm SANDBOX_ID
Prefer the terminal?
nodus init
nodus run
init creates nodus.toml with the GPU smoke test and a $5 budget. Review the
file, then run submits it and waits for completion. Edit the image, command,
and budget to run your own workload. See workload files.
nodus status WORKLOAD_ID
nodus logs WORKLOAD_ID
nodus cancel WORKLOAD_ID
Automatic capacity selection
Nodus uses one policy for new runs: choose the cheapest compatible on-demand capacity by full hourly price. A lower hourly price does not guarantee a lower total completion cost. Optimization tiers are not supported. Existing optimization arguments remain accepted for backward compatibility but have no preference effect on new runs.
Set gpu="H100" to require a GPU model, or omit it to let Nodus choose.
No runtime estimate is needed. See resource options.
GPU enforcement, live logs, login verification, and spending limits require a compatible Nodus backend. Installing the SDK alone does not enable these server features. See backend compatibility before relying on them with a custom or older deployment.
Run your own code
Upload your script with client.assets.upload() or package it in a container.
Choose an image with your dependencies and pass its command to client.run().
- Run a Python script
- Attach code and datasets
- Train or fine-tune a model
- Read logs and download results
- Use Nodus with a coding agent
- Measure customer-owned GPU hosts
- Run tool-driven agents in sandboxes
For individual options, use the Python reference and parameter reference. See troubleshooting if a run fails.
Contributing
See RELEASING.md for release steps. Licensed under Apache-2.0.
Devbox preview builds also include named development sessions.
Benchmark a workload
client.benchmark() accepts an API workload payload, gpu_families, batch_sizes, regions, repetitions, an explicit budget, and an explicit idempotency_key. Reuse the same key after an uncertain response. Both synchronous and asynchronous clients return the server report.
The server divides one total cap into fixed cell allocations. Unused allocations are not redistributed. Use {{batch_size}} in a command argument when varying batch size. Inspect the returned workload IDs, posted ledger costs and measurements with client.get_benchmark(id).
nodus benchmark run request.json --idempotency-key customer-attempt accepts the API JSON shape with workload, matrix and budget_usd. nodus benchmark get bm_ID prints the report. These commands require a backend with the benchmark API.
See durable steps for serial recorded-result replay on deployments with the capability enabled.
Action policies and shadow readiness
client.pools.action_policies(pool_id) reads all four per-kind settings and the
Act kill switch. Use set_action_policy with explicit kind, level,
window_cron and parallelism_cap to save one policy. The sync and async clients
support the same methods. The CLI provides pools action-policies,
pools action-policy and pools act-kill-switch.
Approve and auto require funded Predict and active Route with current consent. The kill switch remains available after entitlement loss. Enabling it blocks new Act authorization while preserving cleanup. Saving a policy does not execute an action. Predict is needed to produce recommendations.
start_shadow starts a future 168-hour observation cycle for an explicit policy.
shadow_runs and pools shadows expose trusted hours, elapsed gaps and
counterfactual action counts. Follow next_cursor with the same pool and kind.
The response reports which action kinds currently have a trusted shadow
producer. Missing observations remain gaps. Generic completed evidence does not qualify automatic actions.
A qualified cycle is evidence readiness, not permission to execute, and
counterfactual counts are neither measured savings nor completed actions.
UTC maintenance windows use five fields. Day-of-month and month must be *.
Minute, hour and weekday accept integers, lists, inclusive ranges or *.
Weekday 0 means Sunday. Steps, names and macros are unsupported.
Act approvals and observed outcomes
client.pools.action_proposals(pool_id) reads retained Act proposals with optional
kind, limit and cursor. approve_action_proposal and
reject_action_proposal submit only the proposal identity. The CLI equivalents
are pools action-proposals, pools approve-action and pools reject-action.
Burst approvals remain under pools proposals.
Approval records intent. Dispatch checks current permissions, funding, policy, expiry and evidence again. The inbox preserves pending, approved, applying, uncertain, applied, failed, no-op and expired states. Only a server-observed outcome confirms application. Measured savings remain null when unavailable and are distinct from customer-reported savings. Default policies become approve when funded Predict and Route are active, while explicit per-kind overrides stay in force. Auto still requires a recent matching trusted shadow cycle.
The Route cheaper waiting policy needs funded Predict and current forecast
evidence of lower expected market completion cost. Missing evidence keeps the
workload waiting. Wait-tuning advice uses complete Route coverage and settled
execution outcomes to suggest bounded changes for future waits. It makes no
saving or completion-time guarantee.
Freeze and resume saved work
client.freeze(workload_id) requests a freeze of checkpointed batch work with a
useful saved checkpoint and a compatible runner. client.freeze_status reports
whether saving and exact compute cleanup have completed. client.resume starts
resumption only after the workload is frozen. Each method is also available on a
Workload and through the async client. The CLI provides freeze,
freeze-status and resume with a workload ID.
The workload states freezing and frozen are nonterminal. A freeze request does
not immediately stop billing for an unresolved compute resource. The response
reports retained checkpoint bytes. Retained storage is not separately metered,
so storage_charge_micros is null, not an inferred zero.
Resume restarts the same customer command with saved checkpoint files. Your training program must load its model, optimizer and progress from those files. This does not restore arbitrary process memory or add guessed resume flags.
Observed Act monetary outcomes identify their measurement_basis. The basis
observed_platform_fee_reduction_30m_v1 compares Route platform fees over equal
30-minute windows. It is not total infrastructure saving or a causal estimate.
MCP clients
Sign in once, then connect Claude, Cursor, Codex or another MCP client:
uvx --from 'nodus-compute[mcp]==0.4.2' nodus login
{
"mcpServers": {
"nodus": {
"command": "uvx",
"args": ["--from", "nodus-compute[mcp]==0.4.2", "nodus-mcp"]
}
}
}
Install uv if needed. The public package starts the server and reuses your saved login. See MCP setup and the seven tools for Codex setup, pip installation and examples.
Release files for nodus-compute 0.4.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nodus_compute-0.4.2.tar.gz | 228.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nodus_compute-0.4.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 356.8 kB
Release files / nodus_compute-0.4.2.tar.gz
| Download URL | nodus_compute-0.4.2.tar.gz |
|---|---|
| Size | 228.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1b829586ffa1d77a371f35be79e3a3e2c67c657da01e6a1724d12e7f8476d199
|
|
BLAKE2b-256 checksum How to use checksums |
84d9266dc56da38237edfefcc330d91360fa324d7b2bcf456e08af033ebe3ecb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / nodus_compute-0.4.2-py3-none-any.whl
| Download URL | nodus_compute-0.4.2-py3-none-any.whl |
|---|---|
| Size | 128.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9922de5104966969bf57128a6f85a3e04693c7ad016f1dd27ae896e01965650e
|
|
BLAKE2b-256 checksum How to use checksums |
d2fe4542981d9501d0f7bfe2f85ec6a2b5fca258a09d154dd313e36fa0878054
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|