robouse
robo-use: agent as a policy for embodied agents. A BenchFlow extension that turns robot manipulation and navigation benchmarks into tasks an LLM agent harness solves by driving the robot itself, one shell command at a time. Think Terminal-Bench plus ARC, for robots.
- Tasks are BenchFlow-native task folders adapted from existing benchmarks (Meta-World, Gymnasium-Robotics, LIBERO-style scenes, ARC-style rule inference, robo-use families, RoboHarm-style safety pairs). Every task ships a reference solution that scores 1.0.
- Harnesses are the agents people already use (Claude Code, Codex, mini-swe-agent, ...) with any model. No harness integration is needed beyond a shell: the robot is driven through the
robocommand. - The episode server is trusted. It owns the simulator, enforces the step budget, records video and judges success from the physical state. The agent never touches simulator objects.
- Every trial is recorded as a BenchFlow trial directory with video, so it opens in the BenchFlow viewer.
Site: robouse.ai. Source: github.com/benchflow-ai/robouse. Status log: STATUS.md.
Install
Python 3.11 or newer. The base package is light (numpy and PyYAML): it gives you the task loader, the robouse command, and the agent-facing robo command, which uses only the standard library. Simulators are extras:
pip install robouse # task loader, `robouse` and `robo` commands
pip install 'robouse[sim]' # + MuJoCo and video recording: tabletop, ARC-style, hard, safety, vision,
# robo-use families, RoboHarm, menagerie and drone suites
pip install 'robouse[metaworld]' # + Meta-World
pip install 'robouse[gymrobotics]' # + Gymnasium-Robotics (Fetch, PointMaze)
pip install 'robouse[robosuite]' # + robosuite 1.5
pip install 'robouse[all]' # all of the above
LIBERO, RoboCasa, BEHAVIOR and DexJoco run in their own environments; see the suite docs linked below.
Tasks ship with the package. All task folders (about 5 MB of text: task.md, reference solution, verifier) are inside the wheel, so a task can be named by its id. robouse tasks lists them and robouse tasks --path prints where they are; copy that folder if you want to edit tasks. Robot meshes are fetched, not shipped: the menagerie, RoboHarm and drone suites use MuJoCo Menagerie models (about 43 MB); run robouse fetch-assets once to download them from the pinned upstream commit into ~/.cache/robouse/menagerie, with a SHA-256 check per file.
Quickstart
# list the bundled tasks (id, backend, env)
robouse tasks
# check a task with its reference solution (reward should be 1)
robouse run --task arc-gravity --harness oracle --out runs/try
robouse run --task metaworld-reach --harness oracle --out runs/try # needs robouse[metaworld]
# let an agent be the policy (needs the `claude` or `codex` CLI and credentials; see docs/harnesses.md)
robouse run --task metaworld-push --harness claude-code --model claude-opus-5-5 --out runs/try
robouse run --task metaworld-push --harness codex --model gpt-6-astra --out runs/try
# a whole suite, 4 at a time (a task folder path works anywhere a task id does)
robouse run-many --tasks "$(robouse tasks --path)/metaworld" --harness codex --model gpt-6-astra --out runs/mw-codex --concurrency 4
To work on robouse itself, install a checkout in editable mode; it then uses the checkout's tasks/ and assets/:
git clone https://github.com/benchflow-ai/robouse.git && cd robouse
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e ".[metaworld,gymrobotics]"
Each trial writes runs/<job>/<task>__<harness>__<id>/ with result.json, the agent's raw output and trajectory, the episode trace, verifier/reward.txt and artifacts/recording.mp4.
How it works
harness (Claude Code, Codex, ...) trusted episode server (robouse serve)
┌─────────────────────────────────┐ JSON ┌────────────────────────────────────┐
│ reads instruction.md │ over a │ MuJoCo simulator (Meta-World, │
│ runs `robo observe`, `robo act` │◄─────────►│ tabletop, Gymnasium-Robotics) │
│ ... `robo done` │ Unix │ step + wall-clock budgets │
└─────────────────────────────────┘ socket │ video + trace recording │
│ success judged from physical state │
└────────────────────────────────────┘
The agent sees only what robo returns: robot and object state as numbers, and camera images on request. See docs/robo-cli.md.
Docs
- docs/task-format.md: task folders, the
robouse:settings block intask.md, observation and scoring modes - docs/robo-cli.md: the agent-facing
robocommand - docs/harnesses.md: supported harnesses and models, credentials, isolation
- docs/suites/: each task suite and where it comes from
- docs/results.md: solve rates per harness, model and suite (generated by
scripts/results.pyfrom the audit) - docs/viewer.md, docs/audit.md: looking at and checking runs
- docs/benchflow.md: running robo-use tasks through BenchFlow (
bench eval run) - docs/molmoact2.md: MolmoAct2 (a VLA) runs on a remote GPU
Metadata
Release files for robouse 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| robouse-0.1.0.tar.gz | 368.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| robouse-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.8 MB
Release files / robouse-0.1.0.tar.gz
| Download URL | robouse-0.1.0.tar.gz |
|---|---|
| Size | 368.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8f3b3458204549b53edd105db09c6176e1e73f5e49194d31cb100f6552d09847
|
|
BLAKE2b-256 checksum How to use checksums |
bef9a7a1f51dd087e9d7523665e73567242d3a13610a245676662b6fa5170b29
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / robouse-0.1.0-py3-none-any.whl
| Download URL | robouse-0.1.0-py3-none-any.whl |
|---|---|
| Size | 1.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
96cff510c800e71e5d7bfd9a04d66aa4209f31bcf8c63b9ac6116c32bc13b2eb
|
|
BLAKE2b-256 checksum How to use checksums |
7b92acd60180451a1b031a7c09163cbe5e2cdd1df4062cb51113566e2f464971
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|