Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Nomad Harness

A Python action runtime connecting model providers to simulated embodiments through typed observations, bounded actions, and execution feedback.

Nomad includes OpenAI, Anthropic, Gemini, and OpenAI-compatible providers (including DeepSeek and Kimi presets), deterministic offline providers, a fake cube-lift world, Robosuite tasks on Panda, UR5e, Kinova3, IIWA, and XArm7 robots, and any MuJoCo scene through a plain MuJoCo adapter. Python 3.10+ is required. Seeded benchmark suites run any of these with per-task results. This is a development release; hardware adapters, and a CLI are planned in the roadmap. Anthropic has offline contract tests; live proposal validation is pending API credits.

Install

Install the published development version explicitly:

python -m pip install 'nomad-harness==0.1.0.dev3'
# For OpenAI and Robosuite:
python -m pip install 'nomad-harness[robosuite,openai]==0.1.0.dev3'

Use an exact version instead of --pre to avoid opting all dependencies into pre-releases. From a source checkout, install the current code with python -m pip install -e '.[dev]', adding robosuite, mujoco, openai, anthropic, or gemini for those integrations (the openai extra also serves DeepSeek, Kimi, and other OpenAI-compatible servers). Robosuite uses the tested pair robosuite==1.5.2, mujoco==3.3.7, and PyOpenGL<4; the plain MuJoCo adapter ([mujoco]) uses mujoco 3.3 with the rendering and video dependencies listed in that extra. Provider and simulator SDKs are loaded only when their adapters need them.

LIBERO integration uses a separate .[libero] environment and discovers suite tasks automatically, with one shared Panda profile and no task-specific manifests. See LIBERO setup, suite discovery, and validation status.

RoboCasa discovers all installed kitchen tasks automatically with a shared mobile Panda profile, native episode instructions and success checks, and bounded arm, gripper, base, and torso actions. It uses a separate .[robocasa] source stack. See RoboCasa setup, benchmark discovery, and validation status.

Isaac Lab provides a provisional single-Franka cube-lift adapter targeting Isaac Sim 6.0.1 and Python 3.12 on an NVIDIA GPU machine. From a Mac, use the Brev setup and test workflow. All seven GPU contract checks passed on a Brev L40S; see the adapter contract and validation status.

Run offline first

This complete example needs no API key, simulator, or GPU:

from nomad import Agent, RunConfig
from nomad.embodiments import FakeEmbodiment, lift_script
from nomad.models import FakeModel

with FakeEmbodiment() as env:
    agent = Agent(FakeModel(lift_script()), env, RunConfig(max_decisions=30))
    result = agent.run("Lift the cube off the table.")

print(result.status, result.outcome.reason)
print("trace:", result.trace_dir)

The fake model follows a fixed script. For a walkthrough that also demonstrates an invalid action and recovery, run python examples/offline_demo.py from the checkout.

Run a model on Robosuite

Set OPENAI_API_KEY and NOMAD_OPENAI_MODEL to a model your account can access. This example sends images to the API and spends API credits:

import os

from nomad import Agent, RunConfig
from nomad.embodiments import Robosuite
from nomad.models import OpenAI

with Robosuite(task="Lift", robot="Panda") as env, OpenAI(
    model=os.environ["NOMAD_OPENAI_MODEL"]
) as model:
    agent = Agent(model, env, RunConfig(max_decisions=30, max_wall_time_s=600))
    result = agent.run("Lift the cube off the table.")

print(result.status, result.outcome.reason)
print("trace:", result.trace_dir)

Agent owns neither adapter. Close models and embodiments with context managers, including when an exception occurs. An environment can serve multiple runs; each run() resets it with the configured seed. Stateful models may retain their state across runs—create a fresh FakeModel to replay its script.

On headless Linux, set MUJOCO_GL=egl or MUJOCO_GL=osmesa for offscreen rendering. The checkout includes examples/quickstart.py, examples/robosuite_model.py, examples/robosuite_scripted.py, examples/robosuite_custom_task.py, examples/mujoco_scripted.py, and examples/benchmark_scripted.py with command-line options.

Other providers are drop-in replacements for OpenAI(...); every adapter offers the same actions as one tool per action and validates every reply locally:

Provider Adapter Extra Key variable Verified
OpenAI OpenAI (Responses API) openai OPENAI_API_KEY Live proposals and Lift episodes
Anthropic Anthropic (Messages API) anthropic ANTHROPIC_API_KEY Offline request and response tests; live validation pending API credits
Gemini Gemini gemini GEMINI_API_KEY Live proposals with gemini-robotics-er-2-preview
DeepSeek DeepSeek openai DEEPSEEK_API_KEY Request shape only; live replies, strict tools, forced calls, and images not yet tested
Kimi Kimi openai MOONSHOT_API_KEY Request shape only; live replies, strict tools, forced calls, and images not yet tested
Any Chat Completions server OpenAICompatible(base_url=...) openai set api_key_env Request shape checked against OpenAI's Chat Completions
python examples/robosuite_model.py --provider gemini --model gemini-robotics-er-2-preview
python examples/robosuite_model.py --provider anthropic --model "$NOMAD_ANTHROPIC_MODEL"

create_provider("gemini") builds an adapter from its NOMAD_GEMINI_MODEL variable, and nomad.models.PROVIDERS lists each provider's key and model variables. Calls made by Agent runs are bounded by the run deadline and are never retried; a failed call becomes feedback and counts toward max_consecutive_failures. For direct propose() calls, the OpenAI-compatible adapters retry rate limits and server errors but fail at once on quota or billing errors, while OpenAI, Anthropic, and Gemini use their SDKs' retries. Free tiers can be smaller than an episode: on 2026-10-06 our Gemini free-tier quota ran out after about a dozen requests, and an episode needs 20–30.

Install Claude support with pip install -e '.[anthropic]'. The adapter reads ANTHROPIC_API_KEY from the environment exported by your config; select a model with --model or NOMAD_ANTHROPIC_MODEL. For keys that require workspace routing, set ANTHROPIC_WORKSPACE_ID or pass workspace_id to the adapter. You can also use it directly:

from nomad.models import Anthropic

model = Anthropic(model="your-claude-model-id", max_output_tokens=4096)

Claude receives the same state and camera images as the other providers. The adapter disables parallel tool calls and defaults to tool_choice="auto" for compatibility across Claude models. On models that support forced tools, pass tool_choice="any". Every returned proposal is still validated locally; truncated replies cannot dispatch actions. Thinking content and image bytes are excluded from provider trace details.

Tasks, robots, and cameras

Task Goal Robots
Lift Lift the cube off the table Panda, UR5e, Kinova3, IIWA, XArm7
Stack Stack the red cube on the green cube Panda, UR5e, Kinova3, IIWA, XArm7
PickPlaceCan Place the red can in its target compartment and release Panda, UR5e, IIWA
PickPlaceMilk Place the milk carton in its target compartment and release Panda, UR5e
PickPlaceBread Place the loaf of bread in its target compartment and release Panda, UR5e
PickPlaceCereal Place the cereal box in its target compartment and release Panda, UR5e
NutAssemblyRound Drop the round nut over the round peg Panda, UR5e
NutAssemblySquare Drop the square nut over the square peg Panda
DoorNoLatch Grasp the handle and pull the door open (latch disabled) Panda

env.default_goal provides each task's instruction. Call nomad.embodiments.robosuite.supported_pairs() to list packaged manifests. Grippers: Panda (its own parallel gripper), UR5e and Kinova3 (Robotiq 2F-85), IIWA (Robotiq 2F-140), XArm7 (its own linkage gripper). Each pair has measured workspace and controller limits; inspect its packaged manifest before changing those limits. robot_base is the robot's fixed root body (robot0_base, at the mount) for every robot. The XArm7 tracks targets less tightly, so its move_ee position tolerance defaults to 8 mm instead of 5 mm. The Kinova3 and XArm7 do not carry the can to its compartment reliably, the UR5e's grip lets the square nut swing and land tilted, and Sawyer and Jaco failed basic grasps, so those pairs have no manifests. For another pair, supply matching task, robot, and manifest="path.yaml"; draft_manifest(task=..., robot=..., path=...) measures a scene and proposes one. Your own tasks and benchmark variants (another robosuite environment, extra robosuite.make arguments, a stricter success predicate) are added with register_task; see custom tasks and examples/robosuite_custom_task.py.

Select camera streams when creating the environment:

with Robosuite(
    task="Lift",
    cameras={"front_rgb": "agentview", "wrist_rgb": "robot0_eye_in_hand"},
) as env:
    print(env.manifest().camera_names)

Every declared stream must have a camera mapping. Added streams default to 640×480; set added_stream_resolution=(width, height) to change that. Cameras are checked against the scene at construction. OpenAI(cameras=("wrist_rgb",), model=...) sends only that stream; traces retain all streams. Video uses its own video_camera.

Scripted baselines use privileged object state and are not model benchmarks. Run one with python examples/robosuite_scripted.py --task Stack --robot UR5e. The packaged pick-and-place baselines succeeded on seeds 0–9 when added (PickPlaceBread on the Panda: 9/10, a finger hits the bin wall on seed 9); this checks the adapter and controller, not visual model performance.

DoorNoLatch is packaged for Panda using Robosuite Door(use_latch=False). Its native success check requires the hinge angle to exceed 0.3 radians. Run the privileged hinge-arc feasibility policy with python examples/door_scripted.py (separate from the pick-and-place baseline factory). examples/door_benchmark.py runs five GPT-6.1 Sol trials with required reasons and reasoning memory, saving Trace Explorer-compatible results to the supplied --output benchmark directory.

Harness experiments

Use RunConfig(require_rationale=True, require_pose_estimate=True, adaptive_views=True, initial_cameras=("front_rgb",)) to require brief action reasons and pose plans and let the model choose its next camera views. Reasons and estimates carry forward with execution feedback in the bounded history; use remember_reasoning=False to test generating them without recalling them. allowed_actions restricts physical tools for capability comparisons.

RunConfig(require_plan=True, require_rationale=True, remember_reasoning=True) requires the model to create an ordered subtask plan before moving. Each action identifies the active subtask; the model must mark it completed with observed evidence before advancing. The persistent checklist and completion updates appear in Trace Explorer. Independent embodiment evaluation still determines task success.

python examples/harness_ablation.py --seeds 3 runs an offline wiring demo across six configurations. Add --provider openai --model "$NOMAD_OPENAI_MODEL" for seeded Robosuite comparisons using API credits. Results include success rates, reported tokens, selected images, inspection calls, and pose position errors. See harness experiments for controls, camera selection, and measurement limits.

Plain MuJoCo scenes

nomad.embodiments.mujoco drives an arm in any MJCF model, without robosuite. Each control tick solves damped least-squares inverse kinematics for the move_ee target and commands the arm's position actuators; set_gripper drives one gripper actuator. The packaged demo is a six-joint arm built from primitive shapes lifting a cube:

from nomad import Agent, ObservationMode, RunConfig
from nomad.embodiments.mujoco import tabletop_lift
from nomad.models import ScriptedLiftPolicy

config = RunConfig(observation_mode=ObservationMode.PRIVILEGED_STATE, max_decisions=40)
with tabletop_lift() as env:
    result = Agent(ScriptedLiftPolicy(), env, config).run(env.default_goal)
print(result.status, result.outcome.reason)

python examples/mujoco_scripted.py runs it with a trace and video; the scripted lift succeeded on seeds 0–19. For your own scene, name the end-effector site, arm joints and actuators, and gripper in MujocoConfig, describe placement and success in a MujocoTask, and write a manifest with backend: mujoco; see plain MuJoCo scenes.

Benchmarks

nomad.benchmark.run_benchmark runs a suite of tasks over seeds through the normal runtime. Each BenchmarkTask builds its embodiment (any backend) and may bring its own model; otherwise the suite's model is used:

from nomad import ObservationMode, RunConfig
from nomad.benchmark import BenchmarkTask, run_benchmark
from nomad.embodiments import Robosuite
from nomad.embodiments.mujoco import tabletop_lift
from nomad.models import ScriptedLiftPolicy, robosuite_baseline

tasks = [
    BenchmarkTask("lift-panda", lambda: Robosuite(task="Lift"),
                  make_model=lambda: robosuite_baseline("Lift", "Panda")),
    BenchmarkTask("mujoco-lift", tabletop_lift, make_model=ScriptedLiftPolicy),
]
config = RunConfig(observation_mode=ObservationMode.PRIVILEGED_STATE, max_decisions=60)
result = run_benchmark(tasks, run_config=config, output_dir="benchmarks")
print(result.table())

The run directory holds benchmark.json (what was run), results.jsonl (one record per episode, written as each finishes), summary.json (per-task success rates with Wilson 95% intervals), and each episode's trace. Success means the embodiment's evaluator confirmed the task. Episodes that fail to run (including a task without a goal) are recorded with their error and count as failures; cancelled episodes are recorded but not scored. Every record and summary names its observation mode, so privileged-state results stay separate from image-only ones, and different tasks are never pooled into one score. examples/benchmark_scripted.py benchmarks the scripted baselines (--pairs all for every packaged robosuite pair).

Configuration

RunConfig controls each episode. Adapter-specific settings live in OpenAIConfig and RobosuiteConfig; pass a config object or typed keyword overrides:

from pathlib import Path
from nomad import RunConfig
from nomad.models import OpenAI, OpenAIConfig

config = RunConfig(
    seed=0,
    max_decisions=30,
    max_wall_time_s=180,
    max_provider_calls=30,
    max_total_tokens=20_000,
    history_window=8,
    trace_root=Path("runs"),
    save_frames=True,
    save_video=False,
)

with OpenAI(OpenAIConfig(model="your-model-id", timeout_s=45), max_retries=1) as model:
    print(model.model_id)
Setting Default Meaning
max_decisions 30 Maximum model decisions
max_wall_time_s 180 Cooperative episode wall-time limit
max_sim_time_s None Optional simulation-time limit, checked between actions
max_provider_calls None Optional proposal-call limit, including failures
max_total_tokens None Stop further decisions once reported tokens reach this limit
max_consecutive_failures 5 Stop after repeated provider, validation, or execution failures
action_freshness_s 60 Reject an action planned from an observation older than this window
history_window 8 Feedback entries sent to the model; 0 disables feedback
observation_mode rgb_proprio Images and robot state; opt into privileged_state for baselines

Limits are cooperative: arbitrary synchronous custom providers cannot be forcibly interrupted. Nomad checks the wall deadline after inference and does not dispatch late actions. OpenAI requests receive the remaining time as a timeout cap, with SDK retries disabled for that bounded request. Direct OpenAI.propose() calls retain the adapter's configured retry policy. Transport timeouts are not hard process-level deadlines, and an in-flight action runs to its own bounded completion or stop request.

Token limits use all reported usage as a lower bound, including partial usage after an error. Exact totals stay None if any usage is missing; unreported consumption cannot be bounded, and one call can exceed the remaining token allowance.

Records reject unknown fields and freeze field assignment. Nested JSON dictionaries remain editable snapshots: env.manifest(), adapter .config, and provider-visible context are detached from the adapter/runtime state. Create a new validated config and adapter to change active settings. Pydantic's model_copy(update=...) does not validate updates; use the config constructor or model_validate() for new settings.

Results, errors, and cancellation

agent.run(goal) returns a RunResult:

  • status: success, failure, budget_exhausted, cancelled, or error.
  • stop_reason: the specific budget, terminal outcome, cancellation, or error.
  • outcome: the independent task evaluator's result and evidence.
  • metrics: decisions, calls, tokens, action results, and elapsed time.
  • trace_dir and error: artifact location and any recorded error.

Task success and run health are separate. A failed stop() produces run status error even if outcome.status is success. A video-finalization error is recorded in error without changing the task's result. Inspect error as well as status. Model refusals, malformed proposals, and execution failures become feedback for the next decision; repeated failures exhaust the configured failure budget. Constructor errors (such as an invalid manifest or missing API key) raise immediately. Runtime exceptions normally become an error result; trace I/O failures can still propagate.

Pass a threading.Event as agent.run(goal, cancel=event) to request cancellation. It is checked between decisions and after inference. To interrupt an action already executing, also call env.stop("cancelled") from the cancelling thread. KeyboardInterrupt finalizes the run as cancelled and is re-raised; a stop failure is recorded as a run error.

Traces

Each run creates a new directory under trace_root containing:

  • config.json: resolved settings, manifest fingerprint, instructions, schema, and provenance.
  • trace.jsonl: observations, proposals, validation, execution, and outcome events.
  • metrics.json: final status, outcome, error, and counters.
  • frames/: captured images when enabled.
  • video.mp4: recording when enabled and supported by the embodiment.
  • explorer.html: a self-contained viewer for the run. Open it in a browser (no server needed) to scrub the video against a step timeline, plot end-effector and object heights, and inspect each decision's model input, output, validation, and execution. Disable with save_explorer=False.

Credential keys and recognized secret environment values are redacted. Adapter metadata must still exclude credentials. Images are stored separately and referenced by the trace; privileged evaluator data is excluded from the default model context.

Extend and contribute

MIT license; see LICENSE.

Metadata

Release files for nomad-harness 0.1.0.dev3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nomad-harness 0.1.0.dev3
File Size Uploaded
nomad_harness-0.1.0.dev3.tar.gz 664.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nomad-harness 0.1.0.dev3
File Interpreter ABI Platform
nomad_harness-0.1.0.dev3-py3-none-any.whl Python 3 none any Details

Total release size: 901.4 kB

Release files / nomad_harness-0.1.0.dev3.tar.gz

Download URL nomad_harness-0.1.0.dev3.tar.gz
Size 664.0 kB
Tags Source
SHA-256 checksum
How to use checksums
b45bae3fa28db73f65999c4137cd5a759fbfd8f90e2f91cfde94764624c0d470
BLAKE2b-256 checksum
How to use checksums
fc15f9013021e3b39797e0df0e324bbb7d5569daccafa9e7f0bfee0d39bf3bc3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.2

Release files / nomad_harness-0.1.0.dev3-py3-none-any.whl

Download URL nomad_harness-0.1.0.dev3-py3-none-any.whl
Size 237.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6ef75b151208b6598b5e0e3aa0d77b38a5329135d44a704bd02a4ebad45bd104
BLAKE2b-256 checksum
How to use checksums
af848ec16d98fa8183492a14aded108552cfef01d986ace01df3d34da13000ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.2
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page