Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Nomad Harness Interface

A proposed Python action runtime for frontier models in physical environments.

Nomad connects a model provider to a simulator or robot adapter through explicit observation, action, and execution-feedback contracts. The first milestone is a working simulation package: change the model or simulation backend without rewriting the agent loop.

Status: implementation brief, Day 3 complete. This README defines the work to build v0.1. The core contracts, offline runtime, and fake backend/model from the first engineering task are implemented and tested, as are the robosuite Lift adapter with its privileged-state scripted baseline (Day 2 of the delivery plan) and the OpenAI provider adapter with a model-driven robosuite rollout (Day 3). Everything else below (Anthropic and Gemini adapters, Isaac Lab, CLI, evaluation runner) is still a proposed contract, not implemented functionality. Development pre-releases are on PyPI as nomad-harness (install with pip install --pre nomad-harness); no benchmark result or hardware support is claimed here.

For the implementation partner: start with the first engineering task, follow the seven-day delivery plan, and use the release acceptance checklist as the definition of done.

Current status

Component State
Core contracts, manifest, validation, context, budgets, tracing, runtime loop Implemented, offline tests
FakeEmbodiment (deterministic cube lift) and FakeModel (scripted) Implemented, offline tests
Robosuite adapter: Panda Lift, absolute base-frame OSC, RGB/proprioception, video Implemented, simulator tests (robosuite 1.5.2, mujoco 3.3.7)
ScriptedLiftPolicy baseline (privileged object state) Implemented; succeeds on robosuite Lift and the fake world
OpenAI adapter: Responses API, strict tool calling, RGB + proprioception context Implemented; offline tests on recorded responses, opt-in live test (openai 3.24)
Isaac Lab, Anthropic, Gemini adapters; nomad CLI; evaluation runner Not started

Run the offline demo (Python 3.10+; no API keys, simulator, or GPU):

python -m pip install -e ".[dev]"
python examples/offline_demo.py   # writes artifacts under runs/
pytest                            # simulator tests are skipped without robosuite

Run the scripted robosuite baseline (MuJoCo on CPU; no API keys or GPU). The robosuite extra pins robosuite 1.5.2 with mujoco 3.3.7: robosuite does not bound its mujoco requirement, and newer mujoco releases break it.

python -m pip install -e ".[dev,robosuite]"
python examples/robosuite_scripted_lift.py --seed 0   # config, trace, frames, metrics, video.mp4
pytest -m robosuite                                    # adapter acceptance tests

The baseline reads the simulator's cube position, so its runs are labeled privileged_state and are not comparable with model runs that see only images and proprioception. On headless Linux, set MUJOCO_GL=egl or MUJOCO_GL=osmesa for offscreen rendering.

Run an OpenAI model on robosuite Lift (spends API credits). The model sees the front camera image and robot state only (rgb_proprio); success comes from robosuite's own predicate. The key is read from OPENAI_API_KEY and never written to traces; the model must be one your account can access.

python -m pip install -e ".[dev,robosuite,openai]"
export OPENAI_API_KEY=...                  # better: source it from a private file outside the repo
export NOMAD_OPENAI_MODEL=your-model-id
python examples/quickstart.py                      # the minimal script shown in section 1
python examples/robosuite_openai_lift.py --seed 0   # with options; add --reasoning-effort high
NOMAD_OPENAI_LIVE=1 pytest -m openai_live          # live multimodal proposal test (two calls)

Plain pytest and CI never call the API: the adapter is tested offline against responses recorded from the live API, and the live test runs only with NOMAD_OPENAI_LIVE=1.

Pilot, not a benchmark. On 2026-10-06, gpt-6.1-sol (default reasoning effort, history window 8, 30-decision budget) lifted the cube on seeds 0–3 in 18, 9, 8, and 14 decisions; seed 4 exhausted the budget after re-grasping six times at the same position. Rerun with --reasoning-effort high, seed 4 succeeded in 8 decisions. That is five seeds plus one rerun: too few for a success rate or for any claim about reasoning effort. The experiment matrix calls for at least 10 seeds per configuration.

1. Product goal and first release

The developer experience is this complete script (examples/quickstart.py), implemented for OpenAI and robosuite:

import os

from nomad import Agent, RunConfig
from nomad.embodiments import Robosuite
from nomad.models import OpenAI

with Robosuite(task="Lift", robot="Panda") as env:
    agent = Agent(
        model=OpenAI(model=os.environ["NOMAD_OPENAI_MODEL"]),
        embodiment=env,
        config=RunConfig(max_decisions=30, max_wall_time_s=600),
    )
    result = agent.run("Lift the cube off the table.")

print(result.status, result.outcome.reason)
print("trace:", result.trace_dir)

The provider reads its API key from the environment. Model identifiers are supplied by the user and checked against actual account access; do not hard-code an assumed future model name.

The launch target is three provider adapters, two simulation backends, one action contract, and reproducible evaluation artifacts. This is an aggressive target. Build and validate one complete path first; only list integrations as supported after they pass their acceptance checks.

Release scope Concrete target Evidence required
Core Typed contracts, bounded agent loop, validation, trace writer Offline tests and deterministic fake-backend demo
First simulator MuJoCo through robosuite; Panda Lift Scripted controller baseline, model rollout, video, trace
Second simulator Isaac Lab; Franka cube-lift task with relative IK Same runtime and canonical actions; separate adapter and pinned environment
Providers OpenAI, Anthropic, Gemini Each produces a locally validated proposal from a real image and robot-state context
Evaluation Seeded task configurations, outcome predicates, metrics Machine-readable results, failures included
Future integration ROS 2 and LeRobot Documented extension points only until implemented and tested

Both initial backends use a Franka/Panda arm. This establishes cross-simulator integration, not cross-robot generalization. A second robot is a later milestone.

2. Architecture: what we are building

flowchart TD
    P["Model providers: OpenAI, Anthropic, Gemini"] --> M["Model adapter"]
    M --> R["Nomad runtime"]
    C["Context builder"] --> M
    R --> V["Action validation"]
    V --> E["Embodiment adapter"]
    E --> MU["MuJoCo / robosuite"]
    E --> IS["Isaac Lab"]
    E -. "future adapter" .-> HW["ROS 2 / LeRobot"]
    E --> F["Observation and execution feedback"]
    F --> C
    R --> T["Trace and evaluation artifacts"]

Nomad owns context construction, proposal parsing, validation, bounded execution, feedback, and tracing. Simulator adapters own sensor extraction, coordinate conversion, controller integration, task predicates, and stopping behavior. Model adapters own provider-specific requests and response normalization.

Keep model-provider code and simulator code independent. A provider must not import robosuite or Isaac Lab. The runtime must not contain branches such as if backend == "isaaclab" to translate actions.

Execution loop

flowchart TD
    O["Read observation and task outcome"] --> D{"Terminal or budget exhausted?"}
    D -- "yes" --> S["Stop and finalize trace"]
    D -- "no" --> C["Build model context"]
    C --> P["Request structured proposal"]
    P --> V{"Valid and executable?"}
    V -- "no" --> F["Record rejection or failure"]
    V -- "yes" --> E["Execute one bounded action"]
    E --> R["Read actual outcome and feedback"]
    R --> O
    F --> O

The default loop executes one action, observes again, then asks the model again. Later, allow a short sequence with checks between actions. Never execute an unlimited model-generated program.

Separate model decisions from control

A model API call may take seconds. It cannot be assumed to run at a robot servo frequency. The model supplies a bounded motion target; the adapter and existing local controller perform control updates and physics stepping.

For the first simulation demo, advance simulation during action execution and pause it during model inference. Record this timing policy. A future real-time or hardware adapter needs continuous local control and a watchdog while inference runs.

Do not repeat a Cartesian delta on every control tick: a requested 2 cm displacement is one total displacement, not 2 cm per tick. Resolve it to a fixed target once, then track that target until convergence or timeout.

3. Core contracts

Implement these six contracts before adding more integrations. Keep them small and versioned. Use typed Python records and local JSON validation; Pydantic is a reasonable initial choice.

Contract Required contents or methods Responsibility
Action Version, name, typed arguments; runtime execution envelope Describe one bounded command
Observation ID, episode ID, simulation time, capture time, RGB frames, robot state, frame metadata Describe what the agent can see now
ActionResult Action ID, status, elapsed time, feedback code, final state reference Describe what actually executed
Embodiment reset, observe, manifest, execute, evaluate, stop, close Implement the world-facing contract
ModelProvider propose(context, schema) -> ModelResponse Return typed proposals and normalized usage/latency
Runtime run(goal, config) -> RunResult Coordinate the loop and enforce budgets

Supporting records should include Proposal, AgentContext, ModelResponse, TaskOutcome, RunConfig, and RunResult. They do not need a separate framework or plugin system.

Proposed backend protocol:

from typing import Protocol

# Names below refer to the typed records to implement in nomad.core.
class Embodiment(Protocol):
    def reset(self, seed: int) -> Observation: ...
    def manifest(self) -> EmbodimentManifest: ...
    def observe(self) -> Observation: ...
    def execute(self, action: Action) -> ActionResult: ...
    def evaluate(self) -> TaskOutcome: ...
    def stop(self, reason: str) -> None: ...
    def close(self) -> None: ...

class ModelProvider(Protocol):
    def propose(self, context: AgentContext, schema: dict) -> ModelResponse: ...

evaluate() returns success, failure, ongoing, or unknown with evidence. Completing a motion is different from completing the task. An environment termination signal is also not automatically task success.

Observation and context requirements

  • RGB images: camera name, resolution, encoding, capture timestamp, and image orientation. Choose one canonical orientation and normalize backend output.
  • Robot state: joint positions/velocities, end-effector pose, measured gripper opening, and any available execution diagnostics. Declare units and coordinate frames.
  • Pose convention: right-handed frames, meters, radians, and quaternion order xyzw. Adapters explicitly convert backend conventions.
  • Camera calibration: include intrinsics and camera-to-world transforms when available. Do not imply metric grounding from an uncalibrated image.
  • History: a bounded window of previous actions, results, and selected frames. Include the most recent error and progress summary.
  • Available actions: only capabilities supported by the active adapter, including argument schemas and limits.
  • Evaluation state: maintain a separate evaluator channel. Do not expose hidden task-success flags, object poses, or reward terms to the agent unless running a labeled privileged-state experiment.

Start with two labeled observation modes: rgb_proprio and privileged_state. The second may include simulator object state for debugging and baselines. Keep their results separate.

4. Action protocol

Use capability negotiation rather than assuming every robot can execute the same command. Nomad standardizes how capabilities are described and invoked; each embodiment implements the subset it supports.

Level Example v0.1 treatment
Cartesian move_ee, set_gripper Required manipulation interface
Semantic skill pick(object_id), place(object_id, target_id) Optional, separately labeled and evaluated
Native Backend-specific action vector Explicit opt-in research escape hatch

The primary launch demo should use Cartesian actions. A semantic pick requires an actual grounding and control implementation; adding its name to a manifest does not implement grasping. If skills use simulator object poses or a pretrained policy, disclose this and evaluate them separately.

Canonical Cartesian commands

move_ee arguments:

Argument Meaning
frame A declared frame such as robot_base
translation_m Total end-effector displacement [x, y, z] in meters
rotation_vector_rad Axis-angle rotation vector [rx, ry, rz]; magnitude is the rotation angle
duration_s Maximum execution duration in simulation time

For the v0.1 robot_base frame, resolve the command once at execution start: p_target = p_start + translation; R_target = Exp(rotation_vector) @ R_start. Convert that target into the backend's required representation. Do not treat the rotation vector as Euler roll/pitch/yaw.

set_gripper arguments: closure in [0, 1], where 0 means fully open and 1 means fully closed, plus duration_s. This is a desired closure setting, not measured grip force. Adapters map it to their own position or actuator conventions and report the measured result.

Example model proposal:

{
  "schema_version": "0.1",
  "actions": [
    {
      "name": "move_ee",
      "arguments": {
        "frame": "robot_base",
        "translation_m": [0.0, 0.0, -0.02],
        "rotation_vector_rad": [0.0, 0.0, 0.0],
        "duration_s": 0.5
      }
    }
  ],
  "request_finish": false
}

For v0.1, require exactly one action when request_finish is false; allow an empty action list only with a finish request. A finish request triggers an independent outcome check. It cannot mark the task successful.

The runtime attaches episode_id, action_id, observation_id, manifest version, and a freshness deadline after parsing. The model does not supply or override those execution identifiers.

Validation and execution rules

  1. Validate proposal structure and reject unknown fields, actions, or frames. Reject nonfinite values and incorrect vector dimensions.
  2. Validate capability support, translation/rotation norms, durations, workspace bounds, and episode/observation freshness. Check candidate targets rather than only checking individual delta components.
  3. Reject invalid requests with machine-readable feedback; do not silently change their meaning. If a backend clips a command internally, record the effective command and clipping reason.
  4. Execute accepted commands through local controllers. Inspect progress and stop conditions between local updates, not only after a long action ends.
  5. Return completed, rejected, failed, timed_out, or cancelled, plus an explicit code such as unsupported_frame, workspace_violation, or target_not_reached.
  6. On ambiguous execution failure, observe the world before deciding what to do. Do not blindly retry a physical action; it may already have partly executed.
  7. Enforce maximum decisions, wall time, simulation time, provider calls, token usage where available, and consecutive failures. Finalize traces in a finally block and call stop before closing the backend.

These are runtime constraints, not a guarantee of collision avoidance or hardware safety. Hardware support requires backend-specific stopping and control validation.

Proposed embodiment manifest

schema_version: "0.1"
embodiment_id: panda_robosuite_lift
backend: robosuite
robot: Panda
task: Lift
frames:
  robot_base:
    handedness: right
    parent: world
    transform_source: adapter
observations:
  front_rgb:
    type: rgb
    resolution: [640, 480]
    orientation: top_left_origin
  proprioception:
    type: robot_state
    pose_frame: robot_base
    quaternion_order: xyzw
actions:
  move_ee:
    type: cartesian_delta
    frames: [robot_base]
    max_translation_norm_m: 0.05
    max_rotation_norm_rad: 0.20
    max_duration_s: 1.0
  set_gripper:
    type: gripper_closure
    range: [0.0, 1.0]
    max_duration_s: 1.0
skills: []
timing:
  inference_policy: pause_simulation
  control_hz: 20

The numeric limits are initial configuration proposals. Validate and tune them on the selected task before release. Read or configure controller rates explicitly; do not assume this manifest overrides the simulator's physics settings. Include workspace bounds, target tolerances, physics timestep, and controller configuration in the resolved run configuration.

5. Integration work

MuJoCo through robosuite: first working path

  • Use robosuite Lift with Panda; explicitly configure and record the Cartesian controller. MuJoCo is the physics engine, and robosuite supplies the robot, task, controller, and sensors.
  • Extract RGB and proprioception into canonical observations; normalize camera orientation and pose conventions.
  • Convert canonical targets to the controller's expected action scale, reference frame, rotation convention, and gripper convention. Read the configured controller settings; avoid assuming a fixed seven-number vector for every configuration.
  • Implement bounded motion execution with convergence feedback and a task-outcome adapter around the selected environment's success predicate.
  • Build a scripted reach/grasp/lift baseline using privileged state before calling a model. This isolates controller/adapter bugs from model reasoning failures.
  • Add Stack or PickPlace only after Lift works. Do not start by constructing new scenes or installing a large benchmark suite.

robosuite documents composite controllers and OSC_POSE; consult the pinned version's controller configuration rather than forwarding canonical commands unchanged. See the official controller documentation.

Isaac Lab: second backend

  • Select a supported stable Isaac Lab/Isaac Sim combination and pin it. Verify the task registry for that version. The documented Isaac-Lift-Cube-Franka-IK-Rel-v0 is the initial candidate.
  • Launch through the chosen version's supported application entry point. Keep Isaac initialization inside the adapter/runner; importing the Nomad core must not launch a simulator.
  • Run one environment initially. Add a camera observation configuration and the same canonical robot-state fields; do not assume a stock RL observation vector contains RGB.
  • Convert target poses into the configured IK action terms, including frame, scaling, quaternion, and gripper conventions. Inspect the environment's action configuration and map it explicitly.
  • Define and record the task predicate. If the two backends use different lift thresholds or task semantics, report them separately; align a common cube-height predicate for a focused transfer comparison.
  • Run the same scripted canonical actions and then the same agent loop. Only the backend configuration and adapter should change.
  • Keep this backend in a separate dependency environment if its Python or simulator requirements differ. Shared protocol does not require both simulators in one virtual environment.

See the official environment registry. Task IDs and installation requirements are version-dependent.

Model providers

Implement OpenAI first, then Anthropic and Gemini. OpenAI is implemented (nomad.models.OpenAI); its shared pieces are meant to be reused by the others:

  • nomad.models._tools turns the proposal schema into one strict tool per action plus finish, and turns tool calls back into an unvalidated raw proposal. Flat per-action schemas fit every provider's strict-schema subset better than one schema with actions in an anyOf.
  • nomad.models._encoding renders the context as system text, user text, and images, with camera poses in the robot-state frame and an image-free summary for traces.
  • Each provider adds only a schema converter for its JSON Schema dialect (nomad.models.openai.schema) and a thin adapter. The runtime keeps the loop: provider agent loops and automatic tool runners are not used.
  • An OpenAICompatible adapter over Chat Completions (vLLM, Ollama, OpenRouter) can reuse the same pieces later.

Each adapter must:

  • Accept a user-provided model identifier, key through environment configuration, a timeout, and an explicit call budget.
  • Encode the same canonical context into the provider's image/text request format.
  • Use supported structured-output or tool-call facilities for the chosen model. Handle provider-specific JSON Schema subsets; always validate locally afterward.
  • Normalize output into ModelResponse: proposal, usage if available, latency, provider/model ID, and error/refusal status. Preserve missing usage as unknown, not zero.
  • Bound retries for request failures. A provider retry must never implicitly re-execute an already dispatched action.
  • Offer fake/cassette responses for offline tests and an explicit opt-in live smoke test. Default CI must not spend API credits.

Use a short action rationale only if useful for debugging; private reasoning traces are not required. Keep image resolution, recent-history window, and prompt template configurable and recorded.

Provider references: Anthropic structured outputs and Gemini structured outputs. Check model-specific support during implementation.

6. Package and developer experience

Use one distribution initially, nomad-harness, with import namespace nomad. The distribution name was unclaimed on PyPI as of 2026-10-05. An unrelated PyPI package named nomad (a SQL migration tool) may also install a top-level nomad module, so the two should not be installed in the same environment. Split integrations into separate distributions only when dependency or maintenance needs justify it.

Releasing

Releases use PyPI Trusted Publishing from .github/workflows/release.yml, so no upload tokens exist in the repository or in GitHub secrets.

  • Dry run: run the Release workflow manually from the Actions tab. It tests, builds, and checks the distributions and keeps them as a downloadable workflow artifact; nothing is uploaded.
  • Publish: bump version in pyproject.toml and __version__ in src/nomad/__init__.py together, merge, and publish a GitHub release. The workflow builds and, after approval on the pypi environment, uploads to PyPI.

A published version can never be re-uploaded, so install the dry-run wheel in a fresh environment before publishing a release:

python -m pip install ./nomad_harness-*.whl   # from the dry run's "dist" artifact
python -c "import nomad; print(nomad.__version__)"

Proposed installation commands, to make real during implementation:

# Run from the repository after pyproject.toml and these extras exist.
python -m pip install -e ".[dev,robosuite,openai]"

# Example target CLI; NOMAD_MODEL is a model accessible to your API account.
nomad run --provider openai --model "$NOMAD_MODEL" \
  --backend robosuite --task Lift --robot Panda \
  --goal "Lift the cube off the table." --seed 0

nomad eval --config benchmarks/robosuite_lift.yaml

Proposed extras: dev, robosuite, openai, anthropic, gemini. Document Isaac Lab installation separately in its supported environment. Import optional SDKs lazily and return actionable missing-dependency messages. API keys must never enter committed files or trace artifacts.

Proposed source layout:

Path What to implement
pyproject.toml Package metadata, CLI entry point, tested Python range, optional dependencies
src/nomad/__init__.py Small public API: Agent, RunConfig, contracts
src/nomad/core/ Schemas, protocols, manifest validation, status enums
src/nomad/runtime.py Budgeted observe/propose/validate/execute loop
src/nomad/context.py Image/state/history context construction
src/nomad/validation.py Protocol and capability checks
src/nomad/models/ Fake, OpenAI, Anthropic, Gemini adapters
src/nomad/embodiments/ Fake, robosuite, Isaac Lab adapters
src/nomad/tracing.py Event serialization, image references, artifact finalization
src/nomad/evaluation.py Episode runner, outcome aggregation, metrics
src/nomad/cli.py run, eval, backend/provider diagnostics
examples/ Offline demo, scripted controller baseline, model rollouts, custom adapter
benchmarks/ Versioned task, seed, budget, observation-mode configurations
tests/ Contract, validation, loop, adapter-conversion, trace tests
docs/ Protocol specification, installation matrix, integration guide

Implement context management for adapters; always close cameras, simulator applications, and provider resources. Avoid simulator-specific imports in src/nomad/__init__.py.

7. Workstreams and dependencies

These roles describe ownership boundaries. One partner can work through them sequentially; a larger team can split them once the contracts are agreed.

Workstream Deliverables Depends on Acceptance check
Core/runtime Schemas, fake world/model, validation, loop Agreed action/observation semantics Offline run handles success, rejection, timeout, cancellation
First backend robosuite adapter and scripted baseline Core schemas Canonical actions move and grasp as documented
Providers/context Three provider adapters, prompt, image/history handling Proposal schema and sample observation Same recorded context yields valid proposals from each configured provider
Second backend Isaac Lab adapter and launch/config guide Stable canonical actions Same runtime executes canonical commands without backend branches
Evaluation/release Seeded runs, metrics, videos, reproducible instructions Working rollout paths Another developer can reproduce a documented run
flowchart TD
    A["Core contracts and fake rollout"] --> B["robosuite controller baseline"]
    A --> C["First provider and context"]
    B --> D["Complete model rollout"]
    C --> D
    D --> E["Isaac Lab adapter"]
    D --> F["Remaining provider adapters"]
    D --> G["Evaluation and tracing"]
    E --> H["Release acceptance"]
    F --> H
    G --> H

Do not delay the first complete rollout to build a broad registry, UI, cloud service, or general plugin system.

8. Seven-day delivery plan

Day Build End-of-day acceptance gate
1 Package skeleton, six contracts, manifest, fake backend/model, runtime skeleton Offline loop produces a trace, respects budgets, rejects invalid commands
2 robosuite Lift, camera/state extraction, target conversion, scripted baseline Seeded scripted cube-lift succeeds; video and final predicate saved
3 OpenAI adapter, context builder, single-action feedback loop Real model receives RGB/state and completes a fully traced episode; report outcome even if unsuccessful
4 Isaac Lab adapter; start remaining providers Scripted canonical actions run in Isaac Lab; runtime remains unchanged
5 Anthropic/Gemini smoke tests, evaluation runner, failure feedback Valid provider proposals and a measured integration matrix; partial integrations labeled
6 Seeded evaluations, harness ablations, installation docs, demo recording Raw episode results, failures, configuration, and selected videos available
7 Fresh-environment reproduction, README status update, release candidate Acceptance checklist passes for every integration advertised as supported

Schedule rule: if the first backend is unstable, finish it before adding the second. If API access or Isaac installation blocks the full target, ship a clearly scoped preview with the working path and mark the rest pending. A functioning adapter is more useful than a support badge over a stub.

The first engineering task

Create a small first PR named “Implement core contracts and offline runtime demo.” It should contain:

  • pyproject.toml, package skeleton, and optional-dependency structure.
  • Typed observation, action, result, proposal, manifest, context, and run records.
  • FakeEmbodiment with deterministic state changes and an independent task predicate.
  • FakeModel with a scripted valid proposal and a scripted invalid proposal.
  • Runtime loop with single-action execution, termination checks, and budgets.
  • A JSONL trace writer and final metrics.json.
  • Tests for unknown actions, nonfinite/out-of-range values, stale observation references, provider errors, partial execution failures, and stopping at a budget.
  • examples/offline_demo.py runnable without API keys, MuJoCo, Isaac Sim, or a GPU.

Then implement the robosuite scripted baseline in the next PR. Review each change against the contracts in this README; avoid adding abstractions without a working use case.

9. Tracing and evaluation

Every episode must save the following artifacts under a unique run directory:

Artifact Required contents
config.json Repository commit, dependency versions, backend/task/controller/camera config, model ID, seed, prompts, action limits, budgets, timing and observation mode
trace.jsonl Timestamped observations, proposals, validation decisions, effective actions, execution results, and outcome checks
frames/ Selected input images referenced by observation ID; do not embed large image blobs in JSONL
video.mp4 Simulator rollout with timing mode documented; optional in offline unit tests
metrics.json Task outcome and reason, decisions, calls, latency, tokens, rejections, execution failures, wall and simulation time

Never log API keys or authorization headers. Trace the actual prompt/context sent to the provider after credential redaction. Record externally hosted image references carefully enough to reproduce inputs before they expire.

Initial experiment matrix

  1. Integration smoke tests: one task per backend with each provider, fixed budgets. Demonstrate interoperability without treating a few episodes as a leaderboard.
  2. Within-backend model comparison: fixed task configuration, action abstraction, observation mode, prompt policy, and episode seeds. Record actual model IDs and run dates.
  3. Harness ablation: freeze the model and backend; compare feedback/history off versus on, then compare single-action feedback with a bounded short sequence. Keep primitive controllers and grounding identical.
  4. Action-level comparison later: Cartesian versus semantic skills. Disclose what each skill contains, including policies, perception, privileged state, or motion planning.

Start with at least 10 seeds for a pilot and report sample count and uncertainty. Expand the sample count for publication-grade claims. Include unsuccessful episodes and provider failures with predefined treatment; do not selectively drop them.

Separate protocol interoperability from policy competence. Zero-shot model use means no task-specific model training; it does not mean no task-specific controller, prompt, perception module, or skill engineering. Document all of those choices.

Use the environment/task evaluator as the ground-truth success check in the initial benchmarks. A model's statement that it finished is not an evaluation label. Never publish illustrative success percentages as measured results.

For the first report, show per-task results for each backend rather than pooling different tasks into one cross-simulator success score. Add LIBERO or a larger suite after the basic adapter is stable, preserving the benchmark's original task definitions.

Research direction after v0.1

The system can support a paper if experiments establish a useful result. Test these questions:

  • How much does execution feedback improve task completion with the model held fixed?
  • What is the success/cost/latency tradeoff between short actions and longer action sequences?
  • Can the same context and action contract transfer across simulators, and later across robot embodiments?
  • Which failures arise from visual grounding, planning, controller execution, or invalid action proposals?

A second simulator alone is not evidence of universal embodiment transfer. Add a genuinely different robot and a documented capability mapping before making that claim. Do related-work review before any novelty or “first framework” claim.

10. Release acceptance checklist

  • A fresh checkout can run the offline example using documented commands.
  • Importing nomad requires no simulator installation and starts no application.
  • robosuite scripted and model-driven episodes save valid traces and outcomes.
  • Isaac Lab scripted and model-driven episodes use the same runtime contract, if listed as supported.
  • Each advertised provider passes a live multimodal structured-proposal test using an accessible model. (OpenAI passes; Anthropic and Gemini pending.)
  • Adapter tests confirm units, frames, rotation composition, gripper mapping, and total-displacement semantics.
  • Invalid and stale commands cannot reach execution; timeouts and partial failures have explicit outcomes.
  • Task success comes from a defined evaluator, and model-visible privileged state is labeled.
  • Budgets bound every run, resources close reliably, and traces finalize on failure.
  • Evaluation configurations, real sample counts, raw metrics, and videos match the advertised results.
  • Installation requirements and exact tested versions are documented separately for each simulator.
  • The project owner selects and adds a license before calling the repository open source or publishing a package. (MIT)
  • Package-name availability is checked before publishing; release credentials remain outside the repository. (Trusted Publishing)
  • README status and support table are updated to reflect implemented behavior.

11. Scope after the first week

Prioritize a second robot adapter, an integration conformance suite, richer failure feedback, and a standard benchmark runner. Then consider semantic skills, ROS 2/LeRobot execution, trace replay, and longer-horizon context.

Defer RL training, fine-tuning, fleet management, multi-agent orchestration, cloud scheduling, and a dashboard until the local package has users and reliable rollouts. The first product is the importable interface and working execution loop.

The launch message should state verified support precisely: “One runtime and action contract for the tested model providers and simulation backends.” Expand that statement as integrations pass the same checks.

12. Implementation references

Consult the official sources for the versions selected during development:

The Nomad contracts and roadmap in this README are design decisions for this project. The linked frameworks supply backend functionality; they do not implement the Nomad runtime.

License

MIT. See LICENSE.

Metadata

Release files for nomad-harness 0.1.0.dev1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nomad-harness 0.1.0.dev1
File Size Uploaded
nomad_harness-0.1.0.dev1.tar.gz 456.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nomad-harness 0.1.0.dev1
File Interpreter ABI Platform
nomad_harness-0.1.0.dev1-py3-none-any.whl Python 3 none any Details

Total release size: 547.4 kB

Release files / nomad_harness-0.1.0.dev1.tar.gz

Download URL nomad_harness-0.1.0.dev1.tar.gz
Size 456.2 kB
Tags Source
SHA-256 checksum
How to use checksums
951b4a9b40b77245bcacb0ac2a2f50cb3f325cdfb7c996549fa0730d5e72eeb8
BLAKE2b-256 checksum
How to use checksums
83909b615654aaec45515622663ce04d7167350dc990ba2b49a23d2c12c56f2f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / nomad_harness-0.1.0.dev1-py3-none-any.whl

Download URL nomad_harness-0.1.0.dev1-py3-none-any.whl
Size 91.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
86b018cb5f24baaf32e565962b7051002c840442b4a2bf655707dba22bccb572
BLAKE2b-256 checksum
How to use checksums
f440c24e5657e1c20f9d6413e088d82e4dff50a6afe3449f7e1893b5e2898a6f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page