Regression testing for robot world-model checkpoints.
Project description
WorldBench Robotics
Testing and regression infrastructure for robotics world models.
Bring your own robot rollout and predicted futures. WorldBench evaluates the checks supported by the available data, marks unsupported metrics N/A, and saves reproducible results.
Real-Model Integration Proof
A single-rollout integration proof using one NanoWM RT-1 rollout and eight generated future frames.
| Field | Verified value |
|---|---|
| Model | NanoWM B2 RT-1 300K (knightnemo/nanowm-b2-rt1-300k) |
| Dataset | RT-1 / Fractal |
| Scope | one rollout |
| Generated future frames evaluated | eight |
| Resolution | 256x256 RGB |
| FPS | 3 |
| Metric | Result |
|---|---|
| Overall | 92.4 |
| Visual Similarity | 89.2 |
| Temporal Stability | 96.3 |
| Action Consistency | N/A |
| Object Permanence | N/A |
| Contact Realism | N/A |
This is a single-rollout integration proof, not a standardized leaderboard result and not a claim that NanoWM is 92.4% accurate.
Artifact: artifacts/real_model_eval/nanowm_rt1_episode0.json
Real Checkpoint Regression Proof
WorldBench also compares checkpoints from the same model family on the same fixed episode suite.
| Verified setup | Value |
|---|---|
| Model family | NanoWM-B/2 |
| Baseline checkpoint | knightnemo/nanowm-b2-rt1-abl-pred-v-50k |
| Candidate checkpoint | knightnemo/nanowm-b2-rt1-300k |
| Dataset | RT-1 / Fractal via LeRobot |
| Fixed episode count | 10 |
| Episode range | 0 through 9 |
| Verified result | Value |
|---|---|
| Baseline Composite Score mean | 85.67 |
| Candidate Composite Score mean | 87.28 |
| Composite Score delta | +1.61 |
| Visual Similarity delta | +2.19 |
| Temporal Stability delta | +0.89 |
| Improved / regressed / unchanged episodes | 9 / 1 / 0 |
| Gate result | strict PASS; engineering-threshold PASS |
WorldBench detected that the candidate improved in aggregate while still surfacing the episode that regressed: episode_002.mp4 changed by -0.33.
This is a fixed 10-episode validation proof, not a standardized leaderboard result or universal model ranking.
Artifacts and documentation: artifacts/checkpoint_validation/, docs/checkpoint_validation.md, and docs/checkpoint_regression.md.
Quickstart
From a repository checkout, install the package and run the committed sample dataset:
python -m pip install -e ".[video]"
worldbench validate examples/demo_dataset
worldbench eval examples/demo_dataset --predictions examples/demo_dataset/good_model --output-root .worldbench/readme-quickstart
worldbench report artifacts/real_model_eval/nanowm_rt1_episode0.json --output .worldbench/readme-quickstart/nanowm_report.md
The sample dataset is synthetic and exists to verify the local setup. The real-model proof above is the committed NanoWM artifact.
How It Works
WorldBench has three verified evaluation paths:
evalscores a WorldBench frame dataset and an optional prediction folder.eval-videoscores one aligned ground-truth/prediction video pair after removing--skip-contextframes.eval-batchscores a checkpoint folder across matching episode videos, andgatecompares baseline and candidate batch artifacts.
Results include JSON output, metric coverage, effective normalized weights, unavailable metric reasons, and per-horizon metric summaries when enough frames are available. compare, report, and dashboard provide local comparison, Markdown, and HTML inspection surfaces.
Supported Metrics And N/A Behavior
WorldBench does not invent semantic scores. A metric returns N/A when the rollout does not provide the semantics or signals needed for reliable evaluation. Unsupported metrics are excluded from the weighted overall score, and the remaining available weights are renormalized.
| Metric | Required input | Current unavailable behavior |
|---|---|---|
| Visual Similarity | aligned ground-truth and predicted RGB frames | if no frame pairs exist, returns 0.0 with issue No aligned frame pairs available. |
| Temporal Stability | at least two predicted future frames | if fewer than two frames exist, returns 0.0 with issue Need at least two predicted frames.; per-horizon t+1 marks it unsupported |
| Action Consistency | at least two predicted frames plus string actions or explicit dx/dy |
N/A for raw numeric action vectors without an adapter, too few frames, or no aligned actions |
| Object Permanence | synthetic-labeled rollout and detectable object pixels | N/A when reliable object tracking is unavailable |
| Contact Realism | synthetic-labeled rollout and detectable robot/object centroids | N/A when reliable robot and object tracking are unavailable |
Default weights are Visual Similarity 0.25, Action Consistency 0.30, Temporal Stability 0.20, Object Permanence 0.15, and Contact Realism 0.10.
The NanoWM artifact had Visual Similarity and Temporal Stability available. Its rounded score comes from:
(89.24097498284189 * 0.25 + 96.32682361785054 * 0.20) / 0.45 = 92.39024104284573
Details: docs/metric_support.md
Real-Data Validation
The repository records native LeRobot validation against chocolat-nya/yaskawa-untangle-dataset episode 0 in network-marked integration tests. Those tests are not part of the default offline pytest run.
| Recorded check | Video timeline | Control timeline |
|---|---|---|
| Video timeline frames | 900 | 4,952 exported rows |
| Control timeline rows | 4,952 source rows | 4,952 exported rows |
| Actions | 7D | 7D |
| States | 7D | 7D |
| Source video | 640x480 RGB | 640x480 RGB |
Source evidence: tests/test_lerobot_integration.py and docs/real_data_validation.md.
LeRobot Support
import-lerobot supports native Hugging Face LeRobot datasets through the optional lerobot extra and a legacy local LeRobot-style folder converter.
worldbench import-lerobot --repo-id chocolat-nya/yaskawa-untangle-dataset --episodes 0:1 --camera observation.images.fixed_cam1 --timeline video --out examples/yaskawa_video
worldbench import-lerobot --repo-id chocolat-nya/yaskawa-untangle-dataset --episodes 0:1 --camera observation.images.fixed_cam1 --timeline control --out examples/yaskawa_control
--timeline video exports one timestep per unique source camera frame. It aligns actions with latest_at_or_before_timestamp and states with nearest_timestamp.
--timeline control exports one timestep per source control row. It aligns actions and states with source_control_row.
Imported actions, states, and metadata preserve source control indices, source control timestamps, source video frame indices, source video timestamps, repo id, camera key, timeline, and alignment strategy when available.
Details: docs/real_data_validation.md
Corruption Validation
Committed corruption artifacts were generated from the recorded Yaskawa video-timeline dataset. Scores decreased monotonically across the tested corruption severities. Temporal scrambling produced a smaller effect than frame freezing in these artifacts.
Frame-freeze artifact: artifacts/frame_freeze_benchmark.json
| Severity | Overall | Temporal |
|---|---|---|
| 0% | 99.68 | 99.28 |
| 5% | 99.40 | 98.66 |
| 15% | 99.09 | 97.97 |
| 30% | 98.81 | 97.36 |
Temporal-scramble artifact: artifacts/temporal_scramble_benchmark.json
| Severity | Overall | Temporal |
|---|---|---|
| 0% | 99.68 | 99.28 |
| 5% | 99.64 | 99.19 |
| 15% | 99.58 | 99.07 |
| 30% | 99.51 | 98.96 |
CLI Reference Summary
| Command | Verified purpose |
|---|---|
worldbench validate DATASET_PATH |
validate a WorldBench frame dataset |
worldbench eval DATASET_PATH --predictions PATH |
score a frame dataset against prediction frames |
worldbench eval-video --ground-truth PATH --prediction PATH --skip-context INTEGER |
score one matching video pair |
worldbench eval-batch --ground-truth PATH --predictions PATH --name TEXT |
score a checkpoint folder across matching episode videos |
worldbench gate --baseline PATH --candidate PATH |
compare batch artifacts and return PASS or FAIL |
worldbench import-lerobot --out PATH |
import LeRobot data into WorldBench format |
worldbench compare TARGET RUN_B |
compare result files or two model folders |
worldbench report RESULT_JSON --output PATH |
generate a Markdown report |
worldbench dashboard RESULT_JSON_OR_DATASET_PATH --no-open |
launch a local dashboard |
Run worldbench COMMAND --help for the installed version's exact options.
Python API Summary
from worldbench import WorldBench
result = WorldBench("examples/demo_dataset").evaluate(
predictions="examples/demo_dataset/good_model"
)
print(result.score)
result.save_json(".worldbench/readme-quickstart/result.json")
Public SDK exports include WorldBench, WorldModelRun, Metrics, evaluate, load_dataset, EvaluationResult, and MetricResult.
Dataset Format
The frame-dataset layout is:
dataset/
episode_001/
frames/
predictions/
actions.json
states.json
metadata.json
actions.json records timestep t, optional timestamps/provenance, an action value, optional dx/dy, and optional gripper state. states.json records timestep t, optional timestamps/provenance, optional observation_state, and optional synthetic tracker coordinates. metadata.json stores episode identity, robot/task labels, FPS, and optional LeRobot provenance fields.
The video workflow accepts one video per episode. Ground truth and prediction folders are paired by identical relative POSIX paths, with matching future-frame count after --skip-context, resolution, and FPS.
Details: docs/DATA_FORMAT.md
Current Limitations
- The public single-rollout NanoWM artifact covers one rollout and eight generated future frames.
- The NanoWM artifact is not a standardized leaderboard result or model accuracy claim.
- Arbitrary numeric robot actions require explicit semantics or an action adapter before Action Consistency is meaningful.
- Real-world Object Permanence and Contact Realism require reliable tracking support.
- Temporal scrambling currently causes a smaller score decrease than frame freezing in the committed corruption artifacts.
- WorldBench evaluates saved predictions; it does not run arbitrary model inference.
- Normal CI does not download LeRobot datasets or run network integration tests.
Roadmap
Working now
- direct video-pair evaluation
- multi-episode batch evaluation
- per-horizon metric summaries
- regression gate
- native LeRobot import
- single-rollout NanoWM integration artifact
- frame-freeze and temporal-scramble corruption artifacts
Next
- explicit action-adapter registry
- evaluate more real-model rollouts
- evaluate a second real model
- get an external user
Later
- simulator adapters
- ROS bags
- shared reports
- standardized leaderboard
Details: docs/ROADMAP.md
Contributing
Use a supported Python version, install the development dependencies, and run the local checks before sending changes:
python -m pip install -e ".[dev]"
ruff check .
pytest
python -m build
Network LeRobot tests are marked integration and are excluded from the default test command.
License
WorldBench is licensed under Apache-2.0. See LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file worldbench-0.4.0.tar.gz.
File metadata
- Download URL: worldbench-0.4.0.tar.gz
- Upload date:
- Size: 88.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
348d6febb295e1b8eac0fef5f4177f1c3b8bace7ebdd4c408ddd5a39aa1532c1
|
|
| MD5 |
39221aee70b87300c27d71f04aecf063
|
|
| BLAKE2b-256 |
e8aec348b82c92e62851b207995ce1183058045024b9f13bdd5fc16ea67aa93c
|
Provenance
The following attestation bundles were made for worldbench-0.4.0.tar.gz:
Publisher:
publish.yml on tigee1311/worldbench
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
worldbench-0.4.0.tar.gz -
Subject digest:
348d6febb295e1b8eac0fef5f4177f1c3b8bace7ebdd4c408ddd5a39aa1532c1 - Sigstore transparency entry: 2207325651
- Sigstore integration time:
-
Permalink:
tigee1311/worldbench@0edacbbda9f3640280252739976f40d5db7fec6e -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/tigee1311
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@0edacbbda9f3640280252739976f40d5db7fec6e -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file worldbench-0.4.0-py3-none-any.whl.
File metadata
- Download URL: worldbench-0.4.0-py3-none-any.whl
- Upload date:
- Size: 83.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1c069ecc7d26b2b81d79ec8a60f8552993dba2ed3d5960223ea443b736f3344d
|
|
| MD5 |
e146bba8e6419f59d66e96d90a511141
|
|
| BLAKE2b-256 |
d10c139edcc30af10a0b9f6c4bafcddf21e8a474877439152a638c27bd4b6e68
|
Provenance
The following attestation bundles were made for worldbench-0.4.0-py3-none-any.whl:
Publisher:
publish.yml on tigee1311/worldbench
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
worldbench-0.4.0-py3-none-any.whl -
Subject digest:
1c069ecc7d26b2b81d79ec8a60f8552993dba2ed3d5960223ea443b736f3344d - Sigstore transparency entry: 2207325668
- Sigstore integration time:
-
Permalink:
tigee1311/worldbench@0edacbbda9f3640280252739976f40d5db7fec6e -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/tigee1311
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@0edacbbda9f3640280252739976f40d5db7fec6e -
Trigger Event:
workflow_dispatch
-
Statement type: