Skip to main content

cua-speedrun

Compare computer-use agents by performance, time, and cost on real desktop tasks.

Quickstart

Install and configure your credentials:

pip install cua-speedrun
cua-speedrun setup
cua-speedrun benchmark --dataset osworld-50 --agent qwen3vl

Prebuilt desktop images are imported from Docker Hub and cached in Modal. Models and application setup are cached on first use. Data lives in ~/.local/share/cua-speedrun; set CUA_SPEEDRUN_HOME to use another location. Use cua-speedrun doctor to inspect your installation.

Run an evaluation

Evaluations run on Modal by default, using credentials from setup, your Modal CLI profile, or environment variables. For an API agent:

export ANTHROPIC_API_KEY="your-api-key"
cua-speedrun benchmark --dataset osworld-50 --agent claude

The command follows progress until the evaluation finishes. To launch and inspect evaluations in a browser, run cua-speedrun dashboard.

  • Add --parallel-evaluations 4 to run four agent replicas in parallel.
  • Bundled GPU agents select their GPU automatically; override it with --gpu L40S.
  • Add --no-preload to disable environment preloading.
  • To use local Linux hardware, add --compute local --environment local. Local desktops require KVM/QEMU.

Inspect results or download trajectories:

cua-speedrun evaluations
cua-speedrun status RUN_ID
cua-speedrun export RUN_ID

Use cua-speedrun catalog to list available agents and benchmarks, or cua-speedrun help benchmark for more options.

Benchmarks

Benchmark Tasks
cua-world-26 26
osworld-50 50
osworld2-52 52
my-pc-bench 38
cua-world-offline 143
osworld-offline 295
osworld2-offline 63

The offline variants contain the full offline task sets; the smaller variants are representative subsets. MyPCBench also requires a MYPCBENCH_JUDGE_API_KEY for its evaluator; see its setup instructions.

Bring your own agent

Start from an implementation in agents/. Each agent has two files:

  • init.py prepares dependencies or starts a model server before task timing begins.
  • agent.py receives the environment URL and task description, then interacts through Computer.

Submit the folder directly:

cua-speedrun validate --agent ./my-agent
cua-speedrun benchmark --dataset osworld-50 --agent ./my-agent

Additional packages can be installed by init.py; the submission uploads init.py and agent.py. An optional agent.json declares gpu and required_environment_variables. To contribute an agent, add its folder to agents/; it is discovered automatically.

Bring your own benchmark

A benchmark is a folder with a manifest.yaml, task folders containing task.yaml, and its environment setup and verifier. The manifest lists tasks:

name: my-benchmark
version: "1"
tasks: [tasks/my-task]

Each task.yaml specifies task_id, description, and an env mapping:

task_id: my-task
description: The task for the agent to complete.
env:
  kind: gym-anything
  env_dir: ${BENCHMARK_DIR}/environment
  task_id: my-task

Keep the Gym-Anything environment and its task setup/verifier inside the benchmark folder. Then:

cua-speedrun validate --dataset ./my-benchmark
cua-speedrun benchmark --dataset ./my-benchmark --agent ./my-agent

The benchmark is copied into the installation. Increase its version when changing a registered task set. To contribute it, add the folder under benchmarks/ and its name to catalog/benchmarks.yaml; packaging is automatic.

Repository structure

Release files for cua-speedrun 0.3.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cua-speedrun 0.3.3
File Size Uploaded
cua_speedrun-0.3.3.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for cua-speedrun 0.3.3
File Interpreter ABI Platform
cua_speedrun-0.3.3-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / cua_speedrun-0.3.3.tar.gz

Download URL cua_speedrun-0.3.3.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
e5cec39598ff55b8e774e0ce726345842e4e5befed47771a611b8ea012682047
BLAKE2b-256 checksum
How to use checksums
7b20e0f546a71f4f2f9b5c022b0656b13e0c6eff98497366c8467e507d64327e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / cua_speedrun-0.3.3-py3-none-any.whl

Download URL cua_speedrun-0.3.3-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
88db06a4f59ef389b23ffd13ff62ebed3afaa6d921073a846e9d490daf118351
BLAKE2b-256 checksum
How to use checksums
78438ab7b023dd7154bd93c09e6018be0ca5a9bb037126400b954be13774ebe1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

0.3.4

2 release files

This release

0.3.3 This release

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page