OpenRUA
Let Your Claude Code or Codex Control Any Robot, Real or Simulated
Through the standard ROS 2 CLI and client library, without relying on any VLA model.
Task sentence here, model here, outcome here. More in docs/demos.md.
A robot-use agent
uses a robot just as a computer-use agent uses a computer. OpenRUA is
the open harness for one: type openrua run panda "pick up the red cube" and
Claude Code or Codex opens in a terminal on the robot's ROS 2 graph,
lists the topics, reads the docs in its workspace, writes a script with
rclpy, runs it, and checks the camera.
OpenRUA gives you one command for three things:
- Play in simulation.
openrua runbrings up a Franka Panda in MuJoCo; the agent drives it the same way it would a real one. - Put an agent on your robot. Draft its file from the robot's live
graph, finish the
TODOlines,openrua run <name>. See docs/your-own-robot.md. - Run experiments.
openrua benchplays a benchmark across tasks and seeds with a fresh sandbox per trial and archives every command the agent ran. See docs/running-experiments.md.
OpenRUA turns your robot into a coding project: your agent explores it like a live codebase, pulls sensor streams into files for reading, and runs commands and programs to move it.
Start playing with your robot like you code a project :)
Quick start
Install, choose a simulator and an agent once, name a robot, run:
pip install git+https://github.com/terminalworld/OpenRUA
openrua config set --sim robosuite --agent claude-code # your defaults
openrua build # the three images, once
openrua install --sim robosuite # the simulator's checkout and venv, once
openrua run panda "pick up the red cube"
The robot comes up on its ROS 2 graph, the agent opens on its
terminal with that sentence, and the robot powers off when you leave;
run first prints which robot, scene and agent it picked.
openrua doctor tells you what is missing before the first run
(Docker, the three images, the simulator install, an agent login); the
details are in docs/install.md.
Choosing what to run
- The world. Without
--benchthe scene is the simulator's own (robosuite: a table and a cube).--bench libero_proloads a benchmark's world instead, a LIBERO kitchen with a bowl, a plate, a wine bottle, a drawer and a stove;--task-suiteand--task-idpick a scene inside it. - The agent.
--agent codexopens Codex;--modelpicks the model. - Once or every time. Whatever
config setstored can also be given on the command line:openrua run panda --sim robosuite --bench libero_pro "pick up the bowl".openrua robots,openrua simulators,openrua benchmarksandopenrua agentslist the choices;openrua benchmarks libero_prolists one benchmark's suites and tasks. - A robot that stays up. The same three steps as separate commands:
openrua up, thenopenrua agent "..."in a second terminal, thenopenrua down.
How it works
your terminal the robot (real or simulated)
┌─────────────────────────┐ ┌──────────────────────────────┐
│ openrua run │ │ ROS 2 graph │
│ └─ Claude Code / Codex │ DDS │ /joint_states /tf /camera │
│ in a sandbox with │◄──────►│ FollowJointTrajectory │
│ ros2 · rclpy · docs│ │ GripperCommand MoveIt │
└─────────────────────────┘ └──────────────────────────────┘
- The interface is the robot's own. The agent sees the topics,
actions and services the robot exposes, plus
machine.yaml(joints, limits, frames, ports) and four short docs. It never sees OpenRUA. - The sandbox is a plain Ubuntu + ROS 2 container with the agent installed, a workspace mounted, and a whitelist proxy as its only way out (the model API; nothing else).
- A robot is a file (
openrua/configs/robots/): the facts true of it wherever it runs (joints, limits, frames, gripper, ports). Your real robot is the same kind of file with amachine:section that says how to reach its graph, passed by path. - A simulator is a file (
openrua/configs/simulators/): the engine, its install, its native scene, and how it drives each robot it embodies. - A benchmark is a file (
openrua/configs/benchmarks/): which robot and simulator, which suites and init states to load, how a trial runs and stops.openrua benchruns trials, checks every promise the workspace docs make before the agent starts, and records each trial with full provenance. - Everything is checked against one schema (
openrua config schema): a misspelled key in any file is an error, never a silent no-op. Your defaults live in~/.openrua/config.yaml.
Supported robots
A robot is what is true of it wherever it runs: joints, limits, frames, gripper, ports, planner. Which simulator embodies it, and what surrounds it, come from the other two kinds of file.
| Robot | Model | Embodied by |
|---|---|---|
panda |
Franka Emika Panda | robosuite, maniskill, calvin, vlabench |
panda-omron |
Panda on an Omron mobile base | robosuite through robocasa / robocasa365's assets |
widowx |
Trossen WidowX 250S | maniskill (the Bridge dataset's arm, through simpler) |
aloha-agilex |
AgileX Cobot Magic with two ARX X5 arms | robotwin |
| your robot | any ROS 2 arm or mobile manipulator, real | its own file; see docs/your-own-robot.md |
openrua robots prints this list from the files on disk, yours included.
Supported simulators
| Simulator | Engine | Robots | Native scene |
|---|---|---|---|
robosuite |
robosuite 1.5 on MuJoCo | panda |
Lift: a table and a cube |
maniskill |
ManiSkill 3 on SAPIEN 3 (PhysX, CPU) | panda, widowx |
PickCube-v1: a table, a cube and a goal marker |
robotwin |
RoboTwin 2.0's harness on SAPIEN 3 (PhysX, CPU) | aloha-agilex |
none: name a benchmark |
calvin |
calvin_env on PyBullet (TinyRenderer, CPU) | panda |
none: name the benchmark |
vlabench |
VLABench's dm_control environments on MuJoCo 3.2 | panda |
none: name the benchmark |
A simulator file knows the engine and how it drives each robot it
embodies; it knows no benchmark. Its install (a venv under
~/.openrua/simulators/) and the benchmarks' own are described in
docs/simulation.md.
Supported benchmarks
| Benchmark | Robot | Simulator | Brings |
|---|---|---|---|
LIBERO (libero) |
panda |
robosuite |
the four standard suites and LIBERO-90, on LIBERO's robosuite 1.4 fork, ROS 2 Jazzy |
LIBERO-PRO (libero_pro) |
panda |
robosuite |
LIBERO's scenes under five perturbation axes, same fork and venv as libero |
LIBERO-Plus (libero_plus) |
panda |
robosuite |
~10,000 perturbed variants of the four suites, its own fork and assets, ROS 2 Jazzy |
LIBERO-Mem (libero_mem) |
panda |
robosuite |
ten non-Markovian tasks with subgoal sequences, its own fork, ROS 2 Jazzy |
RoboCerebra (robocerebra) |
panda |
robosuite |
long-horizon tabletop cases on its LIBERO fork, the Ideal protocol, ROS 2 Jazzy |
CaP-Bench (capbench) |
panda |
robosuite |
CaP-X's tabletop scenes on robosuite 1.5, ROS 2 Humble |
RoboCasa (robocasa) |
panda-omron |
robosuite |
the original release's 24 atomic kitchen tasks (v0.2 on robosuite 1.5.0), ROS 2 Humble |
RoboCasa365 (robocasa365) |
panda-omron |
robosuite |
the 365-task release's kitchens and the Panda-Omron body, ROS 2 Humble |
ManiSkill (maniskill) |
panda |
maniskill |
the eleven table-top Panda tasks that ship with ManiSkill 3, seeded resets, ROS 2 Jazzy |
SimplerEnv (simpler) |
widowx |
maniskill |
the four WidowX Bridge tasks as their authors ported them to ManiSkill 3 (the SAPIEN 2 original needs a GPU; its Google Robot tasks are not ported), the visual-matching placement grid, ROS 2 Jazzy |
MIKASA-Robo (mikasa) |
panda |
maniskill |
the 90 language-conditioned memory tasks (remember, shell game, intercept, ...), its own venv on ManiSkill 3.0.1, ROS 2 Jazzy |
RoboTwin 2.0 (robotwin) |
aloha-agilex |
robotwin |
the fifty dual-arm tasks under the Easy protocol (demo_clean); the Hard protocol needs its 11 GB textures and is not declared; ROS 2 Jazzy |
CALVIN (calvin) |
panda |
calvin |
the 1000 five-subtask chains of the long-horizon evaluation on play table D, each with its fixed initial condition and the benchmark's task oracle, ROS 2 Humble |
VLABench (vlabench) |
panda |
vlabench |
every task registered in the pinned checkout (5 GB of objects and scenes), seeded resets, the task's own termination as success, ROS 2 Humble |
A benchmark names its robot and simulator and brings its own world:
install: (its venv and ROS distro) and scenes: (scene cameras, and
robot embodiments its assets add). openrua benchmarks prints this
list; openrua bench --config <name> runs one.
Supported agents
| Agent | Status |
|---|---|
| Claude Code | supported (--agent claude-code) |
| Codex | supported (--agent codex) |
Bring your own agent. An agent is a manifest (how to install its CLI in the sandbox, which
hosts it talks to, how it logs in) and a small hooks class (how to
launch it); everything else is optional. Pass yours as a path
(--agent ./my-agent.yaml) or send a pull request;
openrua agents lists what is available and what each can do. See
docs/agents.md.
Use your own robot
Draft a profile from the robot's live graph, finish the TODO lines,
and point up at it:
openrua probe --host > my-ur5.yaml # joints, limits, frames, ports, cameras from the graph
openrua run ./my-ur5.yaml "..." # or: openrua config set --robot ./my-ur5.yaml
The profile's machine: section is what the agent's machine.yaml is
generated from: model, joint names and limits, frames,
gripper, and the ports (trajectory, gripper, twist, wrench) the
robot serves. Details in docs/your-own-robot.md.
Simulation and benchmarks
The simulated robots run the community benchmark scenes unchanged; their original success predicates score the trial in place.
openrua bench --config libero_pro --run-id demo \
--task-suite libero_goal_task --task-ids 0,1 --seeds 0 --operator agent
Every trial writes result.json (verdict, preflight, termination,
token accounting), provenance.json (code and simulator commits, image
digests, config and prompt hashes), the agent's full transcript, and
the workspace it left behind; the run directory keeps a SUMMARY.md
regenerated from those files after every trial. Building the simulator checkouts:
docs/simulation.md.
A trial replays from its own commands.sh, and a replay with
--record renders as a video, terminal on the left, cameras on the
right (openrua demo <trial>; see
running-experiments.md).
Architecture
Eight units, one direction of dependency: robot/ (the machine,
simulated or real), sandbox/ (the agent's terminal and workspace),
proxy/ (the only route out), agents/ (the agent contract, registry
and launcher), runner/ (bring-up, preflight, operator, verdict,
record), demo/ (a recorded trial's files rendered into a video),
cli/ and doctor/; under all of them the shared leaves config/
(schema, loader, where things live), errors.py and testing.py. Who
may import whom is enforced by CI (import-linter and
tests/architecture/). The prose is docs/architecture.md.
Documentation
| page | read when |
|---|---|
| docs/install.md | setting a machine up: images, logins, doctor |
| examples/first-task.md | your first task on the simulated Panda |
| docs/your-own-robot.md | describing your robot in one profile |
| examples/real-robot.md | the same flow on a real ROS 2 arm |
| docs/simulation.md | the simulator checkouts and GPU rendering |
| docs/podman.md | machines without Docker |
| docs/demos.md | recorded trials rendered as videos, one per task type |
| docs/running-experiments.md | openrua bench, the runs/ layout, every record field, replays and demo videos |
| docs/cli.md | every verb and flag, exit codes (generated) |
| docs/config.md | every config key (generated) |
| docs/agents.md | adding a coding agent |
| docs/architecture.md | the units and the layering contract |
| CONTRIBUTING.md | conventions for code, names and docs |
Results
Claude Code with Claude Opus 5 (effort high), one trial per task and seed, the benchmark's own success predicate, 240 minutes of active wall clock per trial. Details, protocol and per-model runs in docs/demos.md and the paper.
| Benchmark | Trials | Success |
|---|---|---|
| CaP-Bench (7 tasks × 100 seeded resets) | 700 | 693 (99.0%) |
| LIBERO-PRO (8 cells × 10 tasks × 10 initial states) | 800 | 696 (87.0%) |
| RoboCasa365 (50 tasks × 10 seeds) | 500 | in progress |
Runs dated 2026-08 to 2026-09.
Citation
The paper will be released soon; until then, cite the software:
@software{openrua,
author = {{Terminal World Labs}},
title = {OpenRUA},
year = {2026},
url = {https://github.com/terminalworld/OpenRUA},
license = {Apache-2.0}
}
License
Apache-2.0
Metadata
Release files for openrua 0.0.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| openrua-0.0.7.tar.gz | 661.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| openrua-0.0.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.3 MB
Release files / openrua-0.0.7.tar.gz
| Download URL | openrua-0.0.7.tar.gz |
|---|---|
| Size | 661.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
811d358efd8f2b193abb7f61ff2609a079de7794004793f93c2486a3b7c8065c
|
|
BLAKE2b-256 checksum How to use checksums |
d692da4d552996b3f9c913b67fb42375e3b2d52c7f4cad5d208f398881feae95
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|
Release files / openrua-0.0.7-py3-none-any.whl
| Download URL | openrua-0.0.7-py3-none-any.whl |
|---|---|
| Size | 657.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f340fb08e4dac658c13f161c02548a7f2357fb0aa98529f203b4d1c951e4be96
|
|
BLAKE2b-256 checksum How to use checksums |
2bd9910ffaf9c04e98076bc3b248abb26b5f6dc19f824f6538bedad573ce1b22
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|