Give Physical AI a replay button.
Recorded policy runs in simulation. Films you can inspect. 3D scenes you can edit.
Watch the behavior, step through the evidence, and take the scene with you.
◉ Explore the Stress Lab · ✦ Enter the Butterfly Lab · ◐ Try the before / after · Download the demos ↓ · 中文
No account or install to watch. Preview panels show independent recorded runs.
Review an AI explanation against the run
Open the Qwen3.8 robot review: compare six actual model outputs with three original SmolVLA episodes, sampled images and recorded facts. Switch between images only and images plus the outcome record.
robot-reel review-claims trace.json claims.json --output review.json
Check outcomes, action counts and cited frame labels offline. Explanations remain unverified interpretations. Method, complete records and tutorial.
Choose your workflow
Robot Reel packages policy simulations and 3D scene edits into replayable recordings, source data and editable files. Explore an existing experiment, then use its guide to record or edit your own scene.
| What you need to do | Start here | What you can deliver |
|---|---|---|
| Replay your own LeRobot dataset episode | robot-reel lerobot · SO-101 example |
An offline page with every camera, commanded vs. measured joints and fingerprinted source files |
| Review a policy under changed conditions | SmolVLA Stress Lab | Paired outcomes, camera views, action traces and an offline experiment |
| Inspect simulation parameters and numerical error | Genesis × Newton Solver Lab | Error curves, original samples and editable OpenUSD |
| Edit captured assets and review the change | Blender Scene Lab | Baseline and edited projects, renders and edit parameters |
Each experiment documents its method, checked quantities and runtime requirements. For your own policy or simulator, start with the recording guide and validate exports against the original samples.
Quick start
Try Robot Reel on Hugging Face: compare 30 SmolVLA simulation trials on one task (10 initial states × 3 conditions), orbit recorded GPU cloth, and explore twelve Newton worlds and both Microduck walks. No installation or model account needed. The Space hosts the original recordings; build and publication details include their source commit and checksums. Model & data collection · Feedback & discussion.
New here? Take the three-step tour: compare a real paired outcome, inspect its native Rerun workspace, then verify the full experiment locally. The demo gallery filters policy runs, comparison experiments and 3D creation; previews play on request.
Install the released CLI from PyPI with Python 3.12+:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install robot-reel==0.16.0
robot-reel --help
To check the recordings included in a source checkout, Python 3.12+ and the standard library are enough:
git clone https://github.com/noteflowai/robot-reel.git
cd robot-reel
# Check every trial in the paired policy experiment.
python3 -m robot_reel.cli stress docs/stress
# Check the recorded policy episode and its evidence.
python3 -m robot_reel.cli vla docs/vla
# Turn the included braking comparison into a checked storyboard.
python3 -m robot_reel.cli direct docs/compare/braking \
--plan examples/contact-storyboard.json --output artifacts/director
Or use the verified installation packages or non-root Docker image.
Recording new runs needs the full runtime.
The hosted replays open in a browser without a local simulation environment.
Your LeRobot dataset. One command.
Point Robot Reel at any LeRobotDataset
episode, on the Hugging Face Hub or on disk, and get a folder that opens
offline: every camera on one clock, the language task, and each joint's command
beside its measurement with a jump to the largest difference. The export pins
the Hub commit and the SHA-256 of every source file; --check-source re-reads
the dataset and compares every value. Supports v3.0 and v2.x; LeRobot itself
is not required.
python -m pip install 'robot-reel[lerobot]==0.16.0'
robot-reel lerobot lerobot/svla_so101_pickplace --episode 0 --output artifacts/so101
robot-reel lerobot artifacts/so101 --verify --check-media --check-source
Open the SO-101 example ↗ ·
Guide, checks and limits ·
Source: episode 0 of lerobot/svla_so101_pickplace
(Apache-2.0), the real SO-101 data SmolVLA was fine-tuned on.
One frame. Back in the scene.
Scene Lab — Microduck × captured terrain × Blender. Step through a recorded robot on the original and edited coastal scene. Follow its trajectory and contacts, keep its source camera, or orbit the full robot. Compare the matching MuJoCo and Blender video frames, then download the animated native project or export a source-checked frame JSON.
Both six-second runs are retained, including the falls and slides. All 362 source frames have native transform and virtual-camera checks. Full geometry, body bounds/proxy and recorded-video modes support different rendering limits. Reproduction and measured performance · Portable Blender projects.
One launch. Mind the timestep.
Solver Lab — Genesis × Newton. Six independent L40S / CUDA flights use the same initial state and gravity at 30, 120 and 480 integration steps per second. Compare each recorded arc with the analytic solution, inspect position and velocity errors, and follow specific-energy drift. Smaller steps reduce this pilot's maximum position error from 32.70 cm to 2.05 cm.
Open Solver Lab ↗ · Complete offline experiment · Scene, equations and reproduction
All 366 recorded position/velocity states remain downloadable as JSON/CSV. Genesis native trajectories were reopened and checked; the editable OpenUSD retains every sample. The installed CLI verifies and exports the lab without a GPU. Both engines produce matching values in this simple no-contact, no-drag flight; it is an integration diagnostic, not a ranking of simulators.
Inside a learned Microduck walk.
Share a verifiable Microduck frame. Frame links identify the trace and model
revision. Download the matching experiment and frame JSON, then check them with
robot-reel microduck-review. The installed verifier needs no source checkout
or GPU. Offline workflow.
Microduck Motion Lab. Tap a 3D joint to inspect it, drag to orbit, overlay policy targets, click a 14-joint residual heatmap, and follow the original video. Switch between 0.3 / 0.5 m/s speed commands and inspect all 8,400 measured joint samples. Share a frame, export JSON/CSV, or reopen a received frame JSON after checking every fact against the recording. Take both complete walks offline. Playback stays beside the schematic; four view buttons work from the keyboard. Retry a failed video without losing the selected frame or joint. Review a frame with your agent: load a focused skill through Skills Anywhere, run the read-only source check, then explain the verified facts.
Enter the Microduck Motion Lab ↗ · Both recordings and offline viewer · Methods and checks · Interaction inspiration: mishig's Microduck Anatomy
All 18,000 body transforms were checked against MuJoCo. The schematic fixes the floating root because the original recordings did not save root orientation; it does not infer foot contact. This is simulation with PD-actuator fallback, not hardware. Model-derived geometry and footage retain upstream noncommercial/ share-alike terms. Original implementation; no code or assets copied from the reference Space.
Same sheet. Three ways to fall.
Cloth Lab. Release three independent Newton cloth simulations on NVIDIA L40S, changing only the bending coefficient. Orbit the deforming meshes, overlay them on the same clock, and compare measured vertex motion. Every one of the 42,471 vertex samples is retained in the source data and checked through OpenUSD and Blender. The current lab also exports 1920 × 1080 figures with measured deformation and source fingerprints, plus full-precision sample JSON. Shared links preserve the camera angle so teammates can reopen the same view. Open a received sample JSON to check its facts and restore that view offline, or verify it independently against the source vertices with the current CLI.
Release the sheets ↗ · Offline experiment · Editable OpenUSD · Method, limits & reproduction
The coefficients are solver settings, not calibrated fabric properties. Colors identify cases; the page reports geometric diagnostics and preserves original float32 positions and velocities. No collisions or self-contact are modeled. The browser draws saved meshes with Canvas 2D; no CUDA or Newton installation is needed for playback. Robot Reel 0.7.0+ includes the cloth CLI and complete offline export in its installation package. Recording new runs uses the optional Newton runtime.
Same task. Change the view.
The Stress Lab. SmolVLA runs the same task under reference lighting, reduced light and a shifted camera. Explore 30 real closed-loop trials across ten paired initial states. Select any outcome in the matrix, compare both policy cameras, and jump to the largest measured trajectory difference. Recorded with NVIDIA L40S / CUDA inference, with hardware and timing in every trace.
Compare the policy runs ↗ · Complete offline lab ↓ · MCAP telemetry ↓ · Open it in Foxglove · Reproduce & inspect
Share a moment for review. Export a selected pair as JSON or readable Markdown with your own note. Reopen the JSON to restore the exact source samples, or verify its recorded facts against the full local collection. Held final observations and the complete experiment's counts stay explicit. Review workflow and CLI.
See what a net score hides. The live lab groups every paired seed by outcome. The camera condition's net gain of two successes includes three gains and one loss. Select either group, jump to its recordings, and export the full paired report for independent verification with the 0.8.0+ installed CLI. Compare paired outcomes · Report method and CLI.
Inspect every failure. Filter the recorded classifications and jump to each final motion window, with measured end-effector travel.
Failure analysis. All 14 unsuccessful trials reached the action limit. Every trial remained above the 1 mm stall threshold: end-effector travel over the final tenth of each episode ranged from 44.8 mm to 138.6 mm. These measurements establish motion at the cut-off; task progress and success with a larger action budget require separate evaluation.
Repeatability check. One repeat of the full plan on the same L40S matched 30 / 30 outcomes, action counts, recorded robot states and actions, and 360 / 360 rendered frames consumed by policy calls. One of 3,195 recording-only frames differed; that frame was not used for inference. The result documents repeatability under these recorded conditions. It does not establish general determinism or, by itself, causal attribution. Taxonomy, reproducibility and their limits.
Browse the results on Hugging Face Datasets: 30 trial rows and 20 paired rows, with source hashes, units and the full method. This is the recorded pilot's tabular evidence, not a training dataset or official benchmark.
The 0.8.0 offline lab includes these review tools. Download the ZIP and the sample review JSON, then follow the quick start guide. No installation is needed to replay; the matching release wheel enables independent CLI checks.
One task × three native scene conditions × ten paired seeds. Fixed budget: 160 actions / 8 simulation seconds per trial. Every trial is retained; execution errors stay in the attempt ledger. Inspect applied controls, measured state, separate inference/simulation timings and per-condition confidence intervals. This is a controlled diagnostic, not an official LIBERO benchmark score. Preview plays on simulation time; shorter runs explicitly hold their final sample.
Same seed. Different endings.
Native Rerun inspection. Open three paired Stress Lab trials with six embedded camera videos, measured 3D end-effector paths, applied controls and policy-timing curves on one clock. The portable recording keeps the original JSON and has been read back against every source sample.
Open the Rerun workspace ↗ · Portable recording ↓ · Rebuild & verify
Selected seed 09: reference succeeds; dim lighting and the shifted camera reach the step limit. 405 observations · 41 policy calls · 6 embedded videos. The full experiment still contains 30 trials. Desktop browser recommended; the downloaded file opens locally in Rerun 0.37.2.
0.05° apart. Worlds apart.
The Butterfly Lab. Twelve isolated Newton worlds begin at nearly identical angles. Their recorded paths become a luminous 3D time sculpture. Drag to orbit, switch to a motion overlay, and find the moment a tiny release difference becomes a 6.26 m gap.
Explore the Butterfly Lab ↗ · OpenUSD scene ↓ · Offline experiment ↓ · Reproduce & inspect
601 samples × 12 worlds. All 14,424 body poses checked in native Blender. Adjacent release offsets are 0.05°; the sweep spans 0.55°. The largest recorded gap is world 04 versus 01 (+0.15°), at 12.5 s. Sculpture depth represents time, not physical travel. Preview plays at 3.33×; the interactive replay defaults to 1×.
One recording. Two looks.
Drag between the original MuJoCo simulation and its Blender replay. Jump to the recorded contact, step both views together, then take the editable scene into your own project.
Drag to compare ↗ · How the samples match · Download the Blender scene
Choose your front-row seat
More scenes: Early vs. late braking · SO-100 arm studio · Editable braking scene
Keep the run behind the film
| Watch | Inspect | Reuse |
|---|---|---|
| Browser replays, synchronized views, shot navigation and frame links. | Applied actions, measured poses, recorded outcomes and source revisions. | Offline episodes, MP4s, JSON traces, editable Blender projects and OpenUSD scenes. |
The evidence travels with the demo. Validators check file hashes, timestamps, frame mappings and outcome consistency. Native Blender checks cover all 420 vehicle samples in the directed film and all 362 body transforms in the Newton import. Director check · Newton check.
The browser plays recorded simulations. The VLA example is one seeded rollout, with inference waiting time omitted; new director briefs use your connected agent and a new render. Microduck uses the XML PD-actuator fallback. Scope, provenance and asset terms.
Agent workflow experiments
Robot Reel recordings also provide source data for testing agent workflows. Skills Anywhere delivers instructions; EvalArc checks the resulting programs against the original coordinates and clocks.
| Experiment | Recorded result | Inspect the evidence |
|---|---|---|
| Skill delivery: 27 Qwen3-8B attempts across three engineering profiles | No direct-delivery or MCP attempt fully resolves the task. The no-skill condition resolves 2/3 attempts in the final profile. | Instructions, candidate files and task checks |
| Session handoff: six Qwen3-4B continuations, three per condition | All six retrieval calls succeed. All six programs remain unchanged; each condition resolves 0/3 tasks. | Retrieved history, commands and coordinate checks |
Review independent-source SWE tasks: a separate fixed cohort records 36 attempts on Astropy, pytest and SymPy across four workflow conditions. Direct and MCP preloads deliver identical guidance; an unrelated MCP control matches the prompt length. No attempt obtains native acceptance: 31 have assessable reports and five remain uncertain after upstream infrastructure flags. Eight attempts produce nonempty patches. Methods and offline records keep tool failures, native labels and incomplete usage visible.
Carry the reviewed skill into a new session: a separate six-attempt cohort reuses the predecessor's exact skill bytes through workflow MCP preloads. All six preloads succeed; the memory group retrieves six results. All six programs remain unchanged and no task passes full acceptance. The report connects original pins, delivery receipts, retrieved history and task checks.
These are small studies on public tasks. Each profile and cohort retains its own results, including failures; they do not establish general skill or memory benefits. The research guide documents the methods, source versions and related scene-editing experiments.
Build your own scene
Choose the workflow you want to build:
| I want to… | Start here |
|---|---|
| Run paired VLA stress trials on GPU | CUDA setup, fixed experiment and evidence checks |
| Inspect paired trials in Rerun | Portable video, 3D paths and native readback |
| Run SmolVLA locally | Isolated CPU environment + pinned models |
| Let an agent direct a film | MCP setup + Blender build/render |
| Export recorded simulations to a DCC | Newton → OpenUSD → Blender |
| Explore a physics parameter sweep | Butterfly Lab → twelve isolated worlds |
| Record Microduck, braking or the arm | Recording packs + runtime setup |
| Compare two captured runs | Comparison contract + CLI |
| Replay a LeRobot dataset episode | Hub or local dataset → offline replay |
Make the next scene
Contributions with a working replay and inspectable source data are welcome: new simulation adapters, measured policy comparisons, accessible viewers and editable 3D exports. Contributing · Development and checks · Issues.
Built with LeRobot, MuJoCo, Newton, Blender, OpenUSD, Strands Robots, and Pollen Robotics.
Recorder code: Apache-2.0. VLA images retain their upstream attribution and asset terms; Microduck media retains its noncommercial/share-alike terms. The cover's source mapping and media notice accompany the preview. Independent project; no upstream endorsement is implied.
Release files for robot-reel 0.16.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| robot_reel-0.16.0.tar.gz | 319.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| robot_reel-0.16.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 603.3 kB
Release files / robot_reel-0.16.0.tar.gz
| Download URL | robot_reel-0.16.0.tar.gz |
|---|---|
| Size | 319.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
60de58ca50a67be64f79d7bc1959d5853fefc8b086751b3e1e1522a3c42149b5
|
|
BLAKE2b-256 checksum How to use checksums |
4d5fbe8a253b2d6a574730d07c3e0678549f552a17803e321719e10ec5a2ba82
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / robot_reel-0.16.0-py3-none-any.whl
| Download URL | robot_reel-0.16.0-py3-none-any.whl |
|---|---|
| Size | 283.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d6814b7ce70a7348248d5831ead7f7db9f5a4a91cc4c2e188ec6ec89e0c5c115
|
|
BLAKE2b-256 checksum How to use checksums |
a1e087c97b3df29a7b74c37eba05d5f9564534af076c646e204007bc75da1fe2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log