Skip to main content

KitchenBench

A bimanual kitchen-manipulation benchmark for VLA models.

Built on Inspect Robots · part of WorldEvals, the "Inspect Evals for robotics".

Status: alpha CI Docs License: MIT Coverage Docs coverage Built on Inspect Robots

Note: This project is in early development. The API may change between releases, so pin a version before depending on it.

KitchenBench is 10 kitchen-manipulation tasks expressed as Inspect Robots Tasks. They are embodiment-agnostic, so you run them against any compatible policy/embodiment. The set emphasizes bimanual coordination: pouring, lid removal, folding, part-mating, a pure two-arm handover, and tool-mediated scooping, alongside classic pick-place / stacking / slotted insertion and a multi-instance sort.

It ships a dependency-free mock kitchen so the whole suite runs in CI, and is designed to point straight at real hardware (e.g. YAM bimanual arms driven by MolmoAct2).

The tasks

Task (--task) Goal Bimanual Category
kitchenbench/place_cutlery place the {cutlery} on the {dishware} pick-place
kitchenbench/stack stack the cups / bowls / plates stacking
kitchenbench/place_in_rack place the {dishware} into the dish rack insertion
kitchenbench/pour_pasta pour the dry pasta into the {vessel} granular
kitchenbench/open_container open the {container} articulated
kitchenbench/fold_cloth fold the {cloth} deformable
kitchenbench/seal_container seal the {container} with its lid mating
kitchenbench/handoff hand off the {item} from one arm to the other coordination
kitchenbench/sort_cutlery sort the cutlery into the correct tray compartments classification
kitchenbench/scoop_pasta scoop about {fill_target_g} g of the {pasta} with the {tool} and transfer it to the container granular+tool

Note: KitchenBench-Lite: place_cutlery, stack, place_in_rack, open_container, and scoop_pasta together form KitchenBench-Lite, a five-task subset of the full benchmark.

Note: Every value each task varies over (the objects a {placeholder} can name, counts, forces, placement jitter) is documented in the task reference on the KitchenBench docs site. The per-instance distributions live in specs.py.

Quick start

Install from PyPI (uv shown, plain pip works too; inspect-robots comes along as a dependency):

uv pip install kitchenbench

Installing registers the 10 tasks plus a dependency-free mock kitchen (the kitchen embodiment and the kitchen_scripted / kitchen_random / kitchen_noop policies) with Inspect Robots via entry points, so everything below runs immediately: no hardware, no simulator, no configuration.

Run one task (5 instances × 5 realizations = 25 rollouts, a few seconds on the mock):

inspect-robots list tasks    # the 10 kitchenbench/* tasks from the table above
inspect-robots run --task kitchenbench/pour_pasta --policy kitchen_scripted --embodiment kitchen

Run the whole benchmark from Python with eval_set:

from inspect_robots import eval_set

from kitchenbench import TASK_FACTORIES

ok, logs = eval_set(
    [f"kitchenbench/{key}" for key in sorted(TASK_FACTORIES)],
    "kitchen_scripted",
    "kitchen",
)
for log in logs:
    print(log.eval.task, log.results.metrics["task_success"])

The mock is abstract (it models progress toward the scene goal, like Inspect Robots's CubePick). Its job is to exercise the pipeline and give you a template. In the mock, success depends only on the seeded goal direction, so the sampled setup distributions have no causal effect (P̂ is degenerately 1.0 for the scripted oracle); the distribution content only bites on a real embodiment. The value is the task definitions, which run unchanged on a real robot: to evaluate a real VLA, swap in a real policy/embodiment pair (see "Run it on real hardware" and "Run it in simulation" below).

Task instances & realizations

KitchenBench follows the physical-automation methodology. The key ideas, top-down:

  • A task (e.g. pour_pasta) is a set of task instances.
  • A task instance is one concrete scenario written as a distribution: a stochastic setup (named random variables, each with a distribution) plus a goal (a natural-language success criterion that may reference the sampled variables). It is not a single fixed scene. It is a recipe for generating many.
  • A realization is one sample of that recipe: draw every random variable from its distribution to get one concrete environment (and a concrete goal sentence).
  • Running a (policy, embodiment) pair on K_realizations realizations and averaging the binary successes estimates the instance success probability P̂[Yᵢ = 1].
task  pour_pasta
 ├─ instance 1  (a distribution)  ──realize──▶  5 concrete environments  ──▶  P̂₁
 ├─ instance 2  (a distribution)  ──realize──▶  5 concrete environments  ──▶  P̂₂
 │  … 5 instances total …
 └─ instance 5  (a distribution)  ──realize──▶  5 concrete environments  ──▶  P̂₅

KitchenBench uses the methodology's recommended defaults: 5 instances per task (K_INSTANCES) and 5 realizations per instance (K_REALIZATIONS).

A worked example

This is one of pour_pasta's five instances (from specs.py, lightly reformatted):

TaskInstance(
    instance_id="pour_pasta/measuring-cup-to-bowl",
    goal="pour the dry pasta into the {vessel}",       # {vessel} is sampled
    setup={
        "vessel":         Categorical(("bowl", "cup", "pot")),
        "fill_g":         Uniform(80, 200),            # grams of pasta
        "pour_height_cm": Uniform(8, 15),
        "vessel_x_cm":    Normal(0.0, 3.0),            # placement jitter (cm)
        "vessel_y_cm":    Normal(0.0, 3.0),
    },
    language_vars=("vessel",),
    target_kind="pour_into",
    static={"substance": "dry_pasta"},
)

Realizing it with different seeds samples those distributions into concrete environments an operator can physically arrange (and a goal to give the VLA) (numbers rounded here for readability):

realize(seed=0)                          realize(seed=2)
  Goal: pour the dry pasta into the bowl   Goal: pour the dry pasta into the cup
  Setup:                                    Setup:
    vessel        = bowl                      vessel        = cup
    fill_g        = 156                        fill_g        = 111
    pour_height_cm = 9.9                       pour_height_cm = 10.1
    vessel_x_cm   = +0.3                        vessel_x_cm   = -7.3
    vessel_y_cm   = -1.6                        vessel_y_cm   = +5.4

Inspect and realize instances from Python:

from kitchenbench import SPEC_BY_KEY

inst = SPEC_BY_KEY["pour_pasta"].instances[0]
inst.setup_spec()
# {'fill_g': 'Uniform[80, 200]', 'pour_height_cm': 'Uniform[8, 15]',
#  'vessel': 'Categorical({bowl, cup, pot})', 'vessel_x_cm': 'N(0, 3²)', ...}

r = inst.realize(seed=0)
r.instruction     # 'pour the dry pasta into the bowl'
r.values          # {'vessel': 'bowl', 'fill_g': 156.43…, 'pour_height_cm': 9.88…, …}  (JSON-native)
r.setup_lines     # ('fill_g = 156.43…', 'pour_height_cm = 9.88…', 'vessel = bowl', …)

How it maps to a run

Each instance becomes one Inspect Robots Scene; the 5 realizations are the 5 epochs (Epochs(count=5, reducer="mean")), each seeded independently. Because the reducer is the mean, each scene's reduced task_success is the instance success probability P̂[Yᵢ = 1], exactly the methodology's estimator:

from inspect_robots import eval
(log,) = eval("kitchenbench/pour_pasta", "kitchen_scripted", "kitchen")
for s in log.samples:
    print(s.scene_id, s.reduced["task_success"])   # one P̂ per instance

On real hardware an embodiment (or operator tool) calls realize_scene(scene, seed) to get the concrete setup to arrange. Realization.setup_lines is the "arrange this" checklist, and Realization.instruction is the goal fed to the VLA.

Distribution types (in distributions.py): Uniform(a, b) continuous · Categorical((…), weights=None) over a finite set · Normal(μ, σ) Gaussian · Constant(v) fixed. Every sample is a builtin float/int/str (JSON-native), and Categorical preserves value types (an int category samples back as an int).

Validation status: read before trusting the numbers. The shipped instances are AI-authored drafts (Validation(source="opus-draft"), validated=False). The methodology's K_i = 5 is the count after human validation: 3 experts rating each instance on representativeness and quality, accepted only if both are ≥ 4. Run that commissioning pipeline before relying on the instances; we do not fabricate ratings. Also note eval()'s task-level metrics["task_success"] is the mean of P̂ over instances: a convenience aggregate, not a methodology output (the methodology sorts the per-instance P̂ into quantiles and fits the pTQ / automation-halvings curves).

Run it on real hardware (YAM arms + MolmoAct2)

KitchenBench tasks are embodiment-agnostic. To evaluate on real YAM bimanual arms with MolmoAct2, provide two Inspect Robots components (e.g. in your own adapter package such as robocurve/embodiments):

  • a Policy wrapping MolmoAct2: act(observation) -> ActionChunk (the scene's instruction is fed to the VLA verbatim);
  • an Embodiment for the YAM arms: reset/step/close, declaring its action space (e.g. two 7-DoF arms + grippers) and cameras. Because there is no privileged success oracle, the embodiment should turn the operator's confirmation at episode end into StepResult(terminated=True, termination_reason="success") (or set record.operator_judgement). KitchenBench's task_success scorer reads either. Declare the "self_paced" capability and pace the control loop inside step().
inspect-robots run --task kitchenbench/pour_pasta --policy molmoact2 --embodiment yam_arms

Inspect Robots checks (policy, embodiment) compatibility (action dims, semantics, camera/state keys) before any motion and writes an immutable EvalLog.

Run it in simulation

Every task instance carries a machine-readable sim annotation (see plans/0003-sim-support.md): kitchenbench.sim resolves it into a SceneBlueprint (what to spawn, where, with per-instance success parameters) and defines the benchmark's official quantitative success criteria per target kind, so a physics simulator needs no human operator. A sim embodiment (Isaac Lab, MuJoCo, …):

from kitchenbench import build_blueprint, make_success_checker

# in Embodiment.reset(scene, seed=...) — the same per-epoch seed eval() passes:
bp = build_blueprint(scene, seed)
# ...spawn bp.objects from your asset catalog (nominal dims in kitchenbench.ASSETS),
# apply bp.conditions, register the reserved names gripper_left/gripper_right...
# build the checker AFTER spawning, with the scene at rest: some criteria
# capture initial state at construction (fold's baseline footprint).
checker = make_success_checker(bp, world)   # world implements WorldState (4 queries)

# per step(): once checker(world).success, terminate with
# termination_reason="success" — the existing task_success scorer fires unchanged.

WorldState is four queries any simulator can answer (aabb, contained_fraction, contained_mass_g, opening_fraction); the criteria (nesting-aware stacking, mass-targeted pouring/scooping, transfer-detecting handoff, …) and their thresholds are versioned as SIM_CONTRACT_VERSION. Log it with your runs. Declare supported_target_kinds on your embodiment so Inspect Robots gates scene realizability before any rollout. The checkers work off poses, not sim internals, so a pose-tracked real rig can reuse them too.

Development

Dependency changes: after editing dependencies in pyproject.toml, run uv lock and commit the updated lockfile. CI installs with uv sync --locked and fails with "the lockfile needs to be updated" if you forget. Day-to-day conventions (PR-only main, the required ci-ok check, one-click releases) are documented in CLAUDE.md.

uv venv && uv pip install -e ".[dev]"     # inspect-robots resolved from PyPI (>= 0.6)
uv run pre-commit install
uv run pytest --cov                        # 100% coverage required
uv run ruff check . && uv run mypy

Every public module, class, and function needs a docstring, enforced by Ruff D1; state the contract instead of restating the symbol name.

Citation

If you use KitchenBench in your research, please cite it:

@software{kitchenbench,
  author  = {Robocurve},
  title   = {KitchenBench: A bimanual kitchen-manipulation benchmark for VLA models},
  year    = {2026},
  url     = {https://github.com/robocurve/kitchenbench},
  version = {0.3.0},
  license = {MIT}
}

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kitchenbench-0.6.0.tar.gz (205.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kitchenbench-0.6.0-py3-none-any.whl (47.4 kB view details)

Uploaded Python 3

File details

Details for the file kitchenbench-0.6.0.tar.gz.

File metadata

  • Download URL: kitchenbench-0.6.0.tar.gz
  • Upload date:
  • Size: 205.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for kitchenbench-0.6.0.tar.gz
Algorithm Hash digest
SHA256 c0b63aaa6b4fa671f21d1bf414b9a5f3129e4cfa208d0dee68a70af8e143c1f2
MD5 784c1a35c0451d01da4e9b0cea1de94b
BLAKE2b-256 4952be1c91853ae2859f7ccfdcbd9b5af637b2a8b8ee973d635671433a1fd108

See more details on using hashes here.

Provenance

The following attestation bundles were made for kitchenbench-0.6.0.tar.gz:

Publisher: release.yml on robocurve/kitchenbench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file kitchenbench-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: kitchenbench-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 47.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for kitchenbench-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c088f9566c6437dba6c60f57aa624120bc862c5980f11411aa344b5cc30eb615
MD5 1072415313b2ddec3e2931b4ce267d12
BLAKE2b-256 313dfc26e459b4b8d6293077356fe864781434e2f8665d5cddf61e1ea1b56224

See more details on using hashes here.

Provenance

The following attestation bundles were made for kitchenbench-0.6.0-py3-none-any.whl:

Publisher: release.yml on robocurve/kitchenbench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page