Skip to main content

Control-aware evaluation toolkit for robotics world models.

Project description

WorldBench Robotics

Testing and regression infrastructure for robotics world models.

Bring your own robot rollout and predicted futures. WorldBench scores the behaviors it can reliably measure, marks unsupported metrics N/A, compares runs, and saves reproducible evaluation artifacts.

WorldBench is not another world model. It is a local evaluation toolkit for checking whether generated futures are useful for robotics workflows.

WorldBench demo showing robot world-model evaluation

Badges

Tests Python Version License

Real Model Evaluation Proof

WorldBench has been run on a real pretrained world model and a real robot rollout:

Field Value
Model NanoWM B2 RT-1 300K
Checkpoint knightnemo/nanowm-b2-rt1-300k
Data RT-1 real robot rollout
Scope 1 rollout
Generated future 8 genuinely generated future frames
Metric Score
Overall 92.4
Visual Similarity 89.2
Temporal Stability 96.3
Action Consistency N/A
Object Permanence N/A
Contact Realism N/A

This is a single-rollout integration proof, not a standardized leaderboard result and not a claim that NanoWM is 92.4% accurate.

Compact artifact: artifacts/real_model_eval/nanowm_rt1_episode0.json

More detail: docs/real_model_evaluation.md

One Quickstart

WorldBench is installed from a source checkout today. This README does not claim a PyPI install path.

git clone https://github.com/tigee1311/worldbench.git
cd worldbench
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev,video]"

worldbench --help
worldbench demo
worldbench validate examples/demo_dataset
worldbench eval examples/demo_dataset --predictions examples/demo_dataset/good_model
worldbench report .worldbench/runs/latest/result.json
worldbench dashboard .worldbench/runs/latest/result.json --no-open

The demo creates a synthetic rollout plus good and bad prediction folders. It is useful for smoke testing the CLI and reports, but the project is no longer synthetic-only.

How WorldBench Works

WorldBench evaluates a rollout dataset against predicted future frames.

Input:

  • ground-truth robot rollout frames
  • predicted future frames
  • action logs
  • state logs when available
  • episode metadata and source provenance

Output:

  • per-metric scores
  • N/A status and reasons for unsupported metrics
  • an overall score computed across available metrics
  • per-episode evidence and issues
  • JSON artifacts, Markdown reports, comparison artifacts, and a local dashboard

Default metric weights:

Metric Weight
Visual Similarity 25
Action Consistency 30
Temporal Stability 20
Object Permanence 15
Contact Realism 10

Overall scores are weighted over available metrics only. In the NanoWM proof, Visual Similarity scored 89.2 with weight 25 and Temporal Stability scored 96.3 with weight 20. Action Consistency, Object Permanence, and Contact Realism were N/A, so WorldBench renormalized over the 45 available weight points:

(89.2 * 25 + 96.3 * 20) / 45 = 92.4 overall

Real Data Validation

WorldBench has an opt-in integration test against a public LeRobot Yaskawa cable-untangling dataset:

Check Value
Video timeline 900 frames
Control timeline 4,952 rows
Actions 7D
States 7D
Source video 640x480 RGB

The normal test suite does not download this dataset. The integration test is marked separately because it depends on Hugging Face data access.

Detailed methodology: docs/real_data_validation.md

Metric Availability And N/A Behavior

WorldBench does not invent scores.

A metric returns N/A when the rollout does not provide the semantics required for reliable evaluation. Unsupported metrics are excluded from the overall score, and the remaining metric weights are renormalized.

Metric Available when Returns N/A when Notes
Visual Similarity Ground-truth and predicted image pairs can be aligned. It currently returns 0 with an issue if no image pairs are available. Uses MSE, PSNR, and SSIM-style structure. Falls back to a NumPy SSIM approximation if scikit-image is absent.
Temporal Stability At least two predicted frames are available. It currently returns 0 with an issue if fewer than two predicted frames are available. Measures frame-to-frame deltas, jumps, flicker, and variance.
Action Consistency Actions are string commands such as move_right, or records provide explicit dx and dy. Raw arbitrary numeric action vectors are present without an action adapter. This prevents 7D robot actions from being misread as zero-motion commands.
Object Permanence The rollout is explicitly synthetic and supports the current color/blob tracking heuristic. Real rollouts or other data without reliable object tracking are used. Real-world object permanence needs reliable tracking support before scoring.
Contact Realism The rollout is explicitly synthetic and supports robot/object tracking. Real rollouts or other data without reliable robot and object tracking are used. Real-world contact realism needs reliable tracking support before scoring.

Full details: docs/metric_support.md

Native LeRobot Import

WorldBench supports native Hugging Face LeRobot datasets through the optional lerobot extra. LeRobot is not a mandatory base dependency.

python -m pip install -e ".[lerobot,video]"
worldbench import-lerobot \
  --repo-id chocolat-nya/yaskawa-untangle-dataset \
  --episodes 0:1 \
  --camera observation.images.fixed_cam1 \
  --timeline video \
  --out examples/yaskawa_video

Implemented behavior:

  • --timeline video is the default and exports one WorldBench timestep per unique source camera frame.
  • --timeline control exports one WorldBench timestep per source control row and may repeat camera frames when control runs faster than video.
  • Video timeline action alignment uses the latest control action at or before the video timestamp.
  • Video timeline state alignment uses the nearest source state timestamp.
  • Control timeline action and state alignment use the source control row.
  • Exported actions and states include source provenance fields when available: source control index/timestamp and source video frame index/timestamp.
  • Episode metadata records source repo id, episode index, camera key, timeline, FPS, alignment strategy, source control steps, source video frame counts, and exported timestep counts.
  • The legacy local LeRobot-style folder converter remains available through worldbench import-lerobot <input_path> --out <output_path> and worldbench import-lerobot --demo --out <output_path>.

CLI help confirms the current flags:

worldbench import-lerobot --help

Corruption Validation

WorldBench includes compact corruption benchmark artifacts for real Yaskawa video-timeline data. These are reproducible JSON summaries, not generated video or frame directories.

Frame-freeze benchmark: artifacts/frame_freeze_benchmark.json

Severity Overall Temporal
0% 99.68 99.28
5% 99.40 98.66
15% 99.09 97.97
30% 98.81 97.36

Temporal-scramble benchmark: artifacts/temporal_scramble_benchmark.json

Severity Overall Temporal
0% 99.68 99.28
5% 99.64 99.19
15% 99.58 99.07
30% 99.51 98.96

The temporal-scramble response is currently weaker than the frame-freeze response.

CLI

Current commands:

worldbench benchmark         Run WorldBench benchmark scenarios.
worldbench compare           Compare result files or two model folders inside a dataset.
worldbench dashboard         Launch a local WorldBench dashboard.
worldbench demo              Generate a complete synthetic demo dataset and good/bad model outputs.
worldbench eval              Run all WorldBench metrics and save result.json.
worldbench import-lerobot    Import LeRobot data into WorldBench format.
worldbench init              Create a sample WorldBench dataset folder structure.
worldbench make-demo-video   Generate README demo MP4, GIF, and thumbnail assets.
worldbench make-screenshots  Generate README dashboard and report screenshot assets.
worldbench report            Generate a Markdown report from a result JSON file.
worldbench validate          Validate a WorldBench dataset.

Common command shapes:

worldbench init <path>
worldbench demo [output]
worldbench validate <dataset_path>
worldbench eval <dataset_path> --predictions <predictions_path>
worldbench compare <dataset_path> --models good_model bad_model
worldbench compare <run_a/result.json> <run_b/result.json>
worldbench benchmark --demo
worldbench benchmark <benchmark_path>
worldbench import-lerobot --repo-id <user/dataset> --episodes 0:1 --camera <camera_key> --timeline video --out <output_path>
worldbench report <result_json> --output <report.md>
worldbench dashboard <result_json_or_dataset_path> --host 127.0.0.1 --port 8765 --no-open

worldbench eval writes timestamped results under .worldbench/runs/ and updates .worldbench/runs/latest/result.json.

Python SDK

from worldbench import WorldBench

bench = WorldBench("examples/demo_dataset")
result = bench.evaluate(predictions="examples/demo_dataset/good_model")
print(result.score)
result.save_json("result.json")
result.save_report("report.md")

Composable metrics:

from worldbench import Metrics, WorldBench

bench = WorldBench("examples/demo_dataset")
result = bench.run(
    metrics=[
        Metrics.visual_similarity(),
        Metrics.temporal_stability(),
        Metrics.action_consistency(),
    ],
    predictions="examples/demo_dataset/good_model",
)

Dataset Format

WorldBench datasets are episode folders with frames, optional in-episode predictions, action records, state records, and metadata.

dataset/
  episode_001/
    frames/
      000001.png
      000002.png
    predictions/
      000001.png
      000002.png
    actions.json
    states.json
    metadata.json

actions.json is a list of action records. Current fields include:

{
  "t": 0,
  "timestamp": 0.0,
  "source_control_index": 12,
  "source_control_timestamp": 0.4,
  "source_video_frame_index": 12,
  "source_video_timestamp": 0.4,
  "action": "move_right",
  "dx": 1.0,
  "dy": 0.0,
  "gripper": "open"
}

states.json is a list of state records. Current fields include:

{
  "t": 0,
  "timestamp": 0.0,
  "source_control_index": 12,
  "source_control_timestamp": 0.4,
  "source_video_frame_index": 12,
  "source_video_timestamp": 0.4,
  "observation_state": [0.0, 0.1],
  "robot_x": 20,
  "robot_y": 50,
  "object_x": 80,
  "object_y": 50
}

metadata.json records episode-level information such as name, robot, task, fps, description, and LeRobot provenance fields when imported from LeRobot.

The result schema is represented by EvaluationResult:

  • dataset_path
  • predictions_path
  • created_at
  • score
  • metrics
  • episodes
  • weights
  • issues
  • main_failure

Each metric result contains name, score, status, reason, details, and issues.

Current Limitations

  • The NanoWM proof is one real-model rollout so far.
  • The NanoWM proof evaluates eight generated future frames.
  • The NanoWM score is not a standardized leaderboard result.
  • Arbitrary numeric action vectors require explicit action adapters before action consistency can be scored.
  • Real-world object permanence requires reliable object tracking support.
  • Real-world contact realism requires reliable robot and object tracking support.
  • Temporal scrambling currently produces a weaker score response than frame freezing.
  • Normal CI does not download large LeRobot datasets or rerun expensive model inference.

Roadmap

Working now:

  • synthetic demo
  • external prediction evaluation
  • native LeRobot import
  • video/control timelines
  • real robot rollout evaluation
  • model comparison
  • unavailable-metric handling
  • corruption validation
  • real NanoWM evaluation
  • reports
  • dashboard

Next:

  • direct video-pair evaluation
  • multi-episode aggregation
  • per-horizon curves
  • regression gate
  • explicit action-adapter registry
  • second real world model
  • external users

Later:

  • ManiSkill/RLBench
  • ROS bags
  • shared run reports
  • standardized leaderboard

Full roadmap: docs/ROADMAP.md

Contributing

WorldBench is intentionally small and inspectable. Useful contributions include:

  • action adapters for real robot action spaces
  • tracking adapters for real-world object/contact metrics
  • additional compact real-model artifacts
  • offline tests for importer and metric edge cases
  • report and dashboard polish

Before opening a PR:

python -m pip install -e ".[dev]"
ruff check .
pytest

Keep large generated media, extracted frame folders, temporary datasets, and model prediction directories out of commits.

License

Apache-2.0. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

worldbench-0.2.0.tar.gz (68.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

worldbench-0.2.0-py3-none-any.whl (65.2 kB view details)

Uploaded Python 3

File details

Details for the file worldbench-0.2.0.tar.gz.

File metadata

  • Download URL: worldbench-0.2.0.tar.gz
  • Upload date:
  • Size: 68.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for worldbench-0.2.0.tar.gz
Algorithm Hash digest
SHA256 c0ccb07d43ab957485b73968d2e107dca7bd5c98d58853645849e8c360a43af0
MD5 4ca75cf552a3c59213950a44fb6f2d76
BLAKE2b-256 d85a898996e17a61d9498b12baa79b2d660517464cc69fd1b731084874f76612

See more details on using hashes here.

Provenance

The following attestation bundles were made for worldbench-0.2.0.tar.gz:

Publisher: publish.yml on tigee1311/worldbench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file worldbench-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: worldbench-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 65.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for worldbench-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3e3870c8641e1bea8351e3533fc0c343180c4f6a1aaf26525d46fad7deb5cd9b
MD5 4e6ecfe957452d57e033023f740dffa4
BLAKE2b-256 0572292a0eaac5b47c243120412ce3cd615b0fd2c2f72c32cdaaa54ce4d546ea

See more details on using hashes here.

Provenance

The following attestation bundles were made for worldbench-0.2.0-py3-none-any.whl:

Publisher: publish.yml on tigee1311/worldbench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page