Skip to main content

cua-speedrun

Compare computer-use agents by performance, time, and cost on real desktop tasks.

Quickstart

Install and configure your credentials:

pip install cua-speedrun
cua-speedrun setup
cua-speedrun benchmark --dataset osworld-50 --agent qwen3vl

Prebuilt desktop images are imported from Docker Hub and cached in Modal. Models and application setup are cached on first use. Data lives in ~/.local/share/cua-speedrun; set CUA_SPEEDRUN_HOME to use another location. Use cua-speedrun doctor to inspect your installation.

Run an evaluation

Evaluations run on Modal by default, using credentials from setup, your Modal CLI profile, or environment variables. For an API agent:

export ANTHROPIC_API_KEY="your-api-key"
cua-speedrun benchmark --dataset osworld-50 --agent claude

The command follows progress until the evaluation finishes. To launch and inspect evaluations in a browser, run cua-speedrun dashboard.

  • Add --parallel-evaluations 4 to run four agent replicas in parallel.
  • Bundled GPU agents select their GPU automatically; override it with --gpu L40S.
  • Add --no-preload to disable environment preloading.
  • To use local Linux hardware, add --compute local --environment local. Local desktops require KVM/QEMU.

Inspect results or download trajectories:

cua-speedrun evaluations
cua-speedrun status RUN_ID
cua-speedrun export RUN_ID

Use cua-speedrun catalog to list available agents and benchmarks, or cua-speedrun help benchmark for more options.

Benchmarks

Benchmark Tasks
cua-world-26 26
osworld-50 50
osworld2-52 52
my-pc-bench 38
cua-world-offline 143
osworld-offline 295
osworld2-offline 63

The offline variants contain the full offline task sets; the smaller variants are representative subsets. MyPCBench also requires a MYPCBENCH_JUDGE_API_KEY for its evaluator; see its setup instructions.

Bring your own agent

Start from an implementation in agents/. Each agent has two files:

  • init.py prepares dependencies or starts a model server before task timing begins.
  • agent.py receives the environment URL and task description, then interacts through Computer.

Submit the folder directly:

cua-speedrun validate --agent ./my-agent
cua-speedrun benchmark --dataset osworld-50 --agent ./my-agent

Additional packages can be installed by init.py; the submission uploads init.py and agent.py. An optional agent.json declares gpu and required_environment_variables. To contribute an agent, add its folder to agents/; it is discovered automatically.

Bring your own benchmark

A benchmark is a folder with a manifest.yaml, task folders containing task.yaml, and its environment setup and verifier. The manifest lists tasks:

name: my-benchmark
version: "1"
tasks: [tasks/my-task]

Each task.yaml specifies task_id, description, and an env mapping:

task_id: my-task
description: The task for the agent to complete.
env:
  kind: gym-anything
  env_dir: ${BENCHMARK_DIR}/environment
  task_id: my-task

Keep the Gym-Anything environment and its task setup/verifier inside the benchmark folder. Then:

cua-speedrun validate --dataset ./my-benchmark
cua-speedrun benchmark --dataset ./my-benchmark --agent ./my-agent

The benchmark is copied into the installation. Increase its version when changing a registered task set. To contribute it, add the folder under benchmarks/ and its name to catalog/benchmarks.yaml; packaging is automatic.

Repository structure

Release files for cua-speedrun 0.3.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cua-speedrun 0.3.4
File Size Uploaded
cua_speedrun-0.3.4.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for cua-speedrun 0.3.4
File Interpreter ABI Platform
cua_speedrun-0.3.4-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / cua_speedrun-0.3.4.tar.gz

Download URL cua_speedrun-0.3.4.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
d50a5942a933fe27f8b4d813f46448a37906fb304a673770750ffcadcecfc081
BLAKE2b-256 checksum
How to use checksums
4b64bee6b8f41be8811bbb136f0611d896ee664adf92d2fccf0173a0c053966f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / cua_speedrun-0.3.4-py3-none-any.whl

Download URL cua_speedrun-0.3.4-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
b0d3382f4a5648e4481fbdba2e277d29c9c5d3990a2adcd8b03efc02ed3b100d
BLAKE2b-256 checksum
How to use checksums
f9e13fc28f38c0f27023a9818809360cc3f448b361c75e7a961cb6eec9beee99
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

This release

0.3.4 This release

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page