Skip to main content

cua-speedrun

Compare computer-use agents by performance, time, and cost on real desktop tasks.

Quickstart

Install and configure your credentials:

pip install cua-speedrun
cua-speedrun setup
cua-speedrun benchmark --dataset osworld-50 --agent qwen3vl

Models and desktop images are prepared on first use and cached. Data lives in ~/.local/share/cua-speedrun; set CUA_SPEEDRUN_HOME to use another location. Use cua-speedrun doctor to inspect your installation.

Run an evaluation

Evaluations run on Modal by default, using credentials from setup, your Modal CLI profile, or environment variables. For an API agent:

export ANTHROPIC_API_KEY="your-api-key"
cua-speedrun benchmark --dataset osworld-50 --agent claude

The command follows progress until the evaluation finishes. To launch and inspect evaluations in a browser, run cua-speedrun dashboard.

  • Add --parallel-evaluations 4 to run four agent replicas in parallel.
  • Bundled GPU agents select their GPU automatically; override it with --gpu L40S.
  • Add --no-preload to disable environment preloading.
  • To use local Linux hardware, add --compute local --environment local. Local desktops require KVM/QEMU.

Inspect results or download trajectories:

cua-speedrun evaluations
cua-speedrun status RUN_ID
cua-speedrun export RUN_ID

Use cua-speedrun catalog to list available agents and benchmarks, or cua-speedrun help benchmark for more options.

Benchmarks

Benchmark Tasks
cua-world-26 26
osworld-50 50
osworld2-52 52
my-pc-bench 38
cua-world-offline 143
osworld-offline 295
osworld2-offline 63

The offline variants contain the full offline task sets; the smaller variants are representative subsets. MyPCBench also requires a MYPCBENCH_JUDGE_API_KEY for its evaluator; see its setup instructions.

Bring your own agent

Start from an implementation in agents/. Each agent has two files:

  • init.py prepares dependencies or starts a model server before task timing begins.
  • agent.py receives the environment URL and task description, then interacts through Computer.

Submit the folder directly:

cua-speedrun validate --agent ./my-agent
cua-speedrun benchmark --dataset osworld-50 --agent ./my-agent

Additional packages can be installed by init.py; the submission uploads init.py and agent.py. An optional agent.json declares gpu and required_environment_variables. To contribute an agent, add its folder to agents/; it is discovered automatically.

Bring your own benchmark

A benchmark is a folder with a manifest.yaml, task folders containing task.yaml, and its environment setup and verifier. The manifest lists tasks:

name: my-benchmark
version: "1"
tasks: [tasks/my-task]

Each task.yaml specifies task_id, description, and an env mapping:

task_id: my-task
description: The task for the agent to complete.
env:
  kind: gym-anything
  env_dir: ${BENCHMARK_DIR}/environment
  task_id: my-task

Keep the Gym-Anything environment and its task setup/verifier inside the benchmark folder. Then:

cua-speedrun validate --dataset ./my-benchmark
cua-speedrun benchmark --dataset ./my-benchmark --agent ./my-agent

The benchmark is copied into the installation. Increase its version when changing a registered task set. To contribute it, add the folder under benchmarks/ and its name to catalog/benchmarks.yaml; packaging is automatic.

Repository structure

Release files for cua-speedrun 0.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cua-speedrun 0.3.2
File Size Uploaded
cua_speedrun-0.3.2.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for cua-speedrun 0.3.2
File Interpreter ABI Platform
cua_speedrun-0.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / cua_speedrun-0.3.2.tar.gz

Download URL cua_speedrun-0.3.2.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
d3401ea6ccf6e4ddffc273da43ebed16dfc0cd55da507d72371babc91f7649b3
BLAKE2b-256 checksum
How to use checksums
ff8adad532cac322460da9331806485f2615102ca4b62b3171f2e2eb1eeae131
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / cua_speedrun-0.3.2-py3-none-any.whl

Download URL cua_speedrun-0.3.2-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
f2cddf74b532cfb1220355e809e2538f5122e2996b23992c581ef796005ada7f
BLAKE2b-256 checksum
How to use checksums
e3bee70454286dbf5142e9035e69903f38a548b572de9526dacfc1302bffb951
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

0.3.4

2 release files

0.3.3

2 release files

This release

0.3.2 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page