Skip to main content

cua-speedrun

Compare computer-use agents by performance, time, and cost on real desktop tasks.

Quickstart

Install and configure your credentials:

pip install cua-speedrun
cua-speedrun setup
cua-speedrun benchmark --dataset osworld-50 --agent qwen3vl

Models and desktop images are prepared on first use and cached. Data lives in ~/.local/share/cua-speedrun; set CUA_SPEEDRUN_HOME to use another location. Use cua-speedrun doctor to inspect your installation.

Run an evaluation

Evaluations run on Modal by default, using credentials from setup, your Modal CLI profile, or environment variables. For an API agent:

export ANTHROPIC_API_KEY="your-api-key"
cua-speedrun benchmark --dataset osworld-50 --agent claude

The command follows progress until the evaluation finishes. To launch and inspect evaluations in a browser, run cua-speedrun dashboard.

  • Add --parallel-evaluations 4 to run four agent replicas in parallel.
  • Bundled GPU agents select their GPU automatically; override it with --gpu L40S.
  • Add --no-preload to disable environment preloading.
  • To use local Linux hardware, add --compute local --environment local. Local desktops require KVM/QEMU.

Inspect results or download trajectories:

cua-speedrun evaluations
cua-speedrun status RUN_ID
cua-speedrun export RUN_ID

Use cua-speedrun catalog to list available agents and benchmarks, or cua-speedrun help benchmark for more options.

Benchmarks

Benchmark Tasks
cua-world-26 26
osworld-50 50
osworld2-52 52
my-pc-bench 38
cua-world-offline 143
osworld-offline 295
osworld2-offline 63

The offline variants contain the full offline task sets; the smaller variants are representative subsets. MyPCBench also requires a MYPCBENCH_JUDGE_API_KEY for its evaluator; see its setup instructions.

Bring your own agent

Start from an implementation in agents/. Each agent has two files:

  • init.py prepares dependencies or starts a model server before task timing begins.
  • agent.py receives the environment URL and task description, then interacts through Computer.

Submit the folder directly:

cua-speedrun validate --agent ./my-agent
cua-speedrun benchmark --dataset osworld-50 --agent ./my-agent

Additional packages can be installed by init.py; the submission uploads init.py and agent.py. An optional agent.json declares gpu and required_environment_variables. To contribute an agent, add its folder to agents/; it is discovered automatically.

Bring your own benchmark

A benchmark is a folder with a manifest.yaml, task folders containing task.yaml, and its environment setup and verifier. The manifest lists tasks:

name: my-benchmark
version: "1"
tasks: [tasks/my-task]

Each task.yaml specifies task_id, description, and an env mapping:

task_id: my-task
description: The task for the agent to complete.
env:
  kind: gym-anything
  env_dir: ${BENCHMARK_DIR}/environment
  task_id: my-task

Keep the Gym-Anything environment and its task setup/verifier inside the benchmark folder. Then:

cua-speedrun validate --dataset ./my-benchmark
cua-speedrun benchmark --dataset ./my-benchmark --agent ./my-agent

The benchmark is copied into the installation. Increase its version when changing a registered task set. To contribute it, add the folder under benchmarks/ and its name to catalog/benchmarks.yaml; packaging is automatic.

Repository structure

Release files for cua-speedrun 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cua-speedrun 0.3.0
File Size Uploaded
cua_speedrun-0.3.0.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for cua-speedrun 0.3.0
File Interpreter ABI Platform
cua_speedrun-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / cua_speedrun-0.3.0.tar.gz

Download URL cua_speedrun-0.3.0.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
a1d9c3d1b4337b0365a2401a30dd8e86e5b0ce6f3a2da219c39bdd660f13b939
BLAKE2b-256 checksum
How to use checksums
1164c79662805f1bcc69333770b0d93184b6fcaa67a0e2af01b070ecc0ad06b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / cua_speedrun-0.3.0-py3-none-any.whl

Download URL cua_speedrun-0.3.0-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
5b7ab1afcc0b476321d5e6a995f0874b773b5ee68eeed99c2a1d748a79994e84
BLAKE2b-256 checksum
How to use checksums
6e47d7aeaa1b24f52a864b8bf9d358c9f945555e276e62e269f0abeade0975d0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page