Skip to main content

openadapt-evals

Tests PyPI Python License: MIT

You changed something in the compiler and you need to know whether it got better or worse. This runs GUI agents and compiled OpenAdapt workflows against a benchmark, scores them with the benchmark's own verifier instead of asking the agent how it did, provisions the VMs that takes, and writes the numbers to a file you can commit.

It's for people working on OpenAdapt, and for anyone who wants to measure a GUI agent against Windows Agent Arena. You don't need it to record or replay a workflow, which is openadapt-flow.

Docs · CLI reference · Workflows and layout · Evidence boundary

Benchmark viewer

Try it without a VM

pip install openadapt-evals
openadapt-evals mock --tasks 5
11:24:42 [INFO] Task 5/5: mock_file_explorer_001
11:24:42 [INFO] Step 0: Agent chose action: click
11:24:42 [INFO] Step 1: Agent signaled task completion
11:24:42 [INFO] [SUCCESS] Task mock_file_explorer_001 completed successfully (score: 1.00)
11:24:42 [INFO] Saved summary: 5/5 outcomes succeeded (100.0%); 0 errors across 5 attempts

==================================================
Evaluation Results
==================================================
Attempts:     5
Outcomes:     5
Errors:       0
Success rate: 100.0% (outcomes only)
Avg score:    1.000 (outcomes only)
Avg steps:    1.0 (all attempts)

Abridged output from 0.94.0. The mock adapter and its deterministic agent both always succeed, so 100% here means the harness is wired up, and nothing else. Every other command in this repository can cost money.

Price a real run before you pay for it

openadapt-eval-flow is dry-run unless you pass --live. Dry runs provision nothing, start no VM, and make no network calls:

openadapt-eval-flow --mode replay --tasks 154 --dry-run
  154 tasks:
    Azure VM-hours:        2.82 vm-hours @ $0.19/hr  = $0.54
    Agent token cost:      $0.00  (paid tasks=0.0, 0 steps each)
    -> TOTAL:              $0.54   ($0.0035/task)
    Pure-agent baseline:   $39.06   (22 steps/task, all paid)
    Savings vs baseline:   $38.53

  HARD GUARDRAILS (enforced on any --live paid run):
    per-run cap:      $0.50
    total cap:        $5.00
    per-task tokens:  60,000
    billing-abort:    after 2 consecutive errors

Those caps aren't decoration. An early uncapped run cost real money, which is how they got there.

Drive it from Python

from openadapt_evals import (
    ApiAgent, WAALiveAdapter, WAALiveConfig,
    evaluate_agent_on_benchmark, compute_metrics,
)

adapter = WAALiveAdapter(WAALiveConfig(server_url="http://localhost:5001"))
agent = ApiAgent(provider="anthropic")

results = evaluate_agent_on_benchmark(agent, adapter, task_ids=["notepad_1"])
print(f"Success rate: {compute_metrics(results)['success_rate']:.1%}")

Run against a live WAA server

oa-vm pool-create --workers 1        # or --cloud aws
oa-vm pool-wait --qualification-dir ./proofs
openadapt-evals run --agent api-claude --task notepad_1
openadapt-evals view --run-name live_eval
oa-vm pool-cleanup -y                # stops billing

Forget the last line and the VM bills until you remember. pool-wait, pool-run, and pool-auto require --qualification-dir, a directory holding a fresh <worker>.identity.json and <worker>.egress.json for each worker; they refuse to run without it. Most pool commands take --cloud azure (the default) or --cloud aws, though pool-logs, pool-vnc, and pool-exec do not. The rest of the oa-vm subcommands, the oa and openadapt-evals command tables, and the configuration and AWS SSO setup are in docs/CLI.md.

What the current evidence says

Every published report pins one exact openadapt-flow wheel, so a Flow release can invalidate a report without a single commit landing here. docs/eval_results/PUBLISHED_EVIDENCE.json records which set is current and which release it was measured against, and scripts/check_published_evidence_freshness.py fails when that pin drifts. It runs offline on every pull request and against PyPI on a daily schedule.

The current set is current_flow_v1_34_0_local_20260829, measured 2026-08-29 against Flow 1.34.0 on macOS with headless Chromium. One synthetic MockMed workflow, three arms, three runs each:

Condition Compiled replay DOM positional DOM name-scoped
clean 3/3 3/3 3/3
theme 3/3 3/3 3/3
rename 3/3 0/3 0/3

The rename row is the one worth reading. It changes Open to View and Save Encounter to Submit Encounter, and both Playwright selector controls failed loudly at the first renamed locator before mutating anything. Those are unsupported-drift halts, not silent wrong writes. Compiled replay is also roughly thirty times slower per run than a working selector, 6.4s against 0.20s, which is the honest reminder that structural actuation should stay the preferred tier wherever an application gives you one.

Zero silent incorrect successes, zero wrong actions, zero over-halts, zero model calls, $0.00 across all 27 runs. Full report, caveats, and the dependency freeze: docs/eval_results/current_flow_v1_34_0_local_20260829/.

Publish a new set rather than editing an old one. Superseded reports stay reproducible against the wheel they were measured on, and none of the nine get deleted.

What this repository cannot tell you yet

  • There is no current Flow-versus-zero-shot number. The word current on an evidence set means release-fresh, not production-accepted. No Azure WAA VM was started for the 1.34.0 set and no model was called.
  • The live WAA path is not finished. scripts/eval_flow_on_waa.py leaves WAALiveAdapter.evaluate unwired on the replay path, so it cannot independently score success, and the hybrid live path returns before execution because its adapter isn't connected either. Wiring both is step one of any valid comparison run.
  • The evidence is local and synthetic. Bundled MockMed, one workflow, one macOS host. No hosted lifecycle, no Windows UIA, no RDP, no Citrix, and no real customer application is represented anywhere in it.
  • Only two benchmark families exist. BenchmarkAdapter is built to extend to OSWorld or WebArena. Today there is WAA, live and mock, and there is LocalAdapter for native desktop runs.
  • Some of it is deliberately absent. Deployment-derived thresholds, tuned adversary parameters, per-system-of-record oracle recipes, and real customer datasets stay out of this open repository.

The missing acceptance tracks and the exit condition for each are written down in docs/eval_results/PRODUCTION_READINESS.md.

What else is in here

A meta-benchmark harness runs record, compile, replay, heal, verify across any registered Environment and emits one metrics row per (env, task, mode), and exports to Inspect AI. Thirteen agents ship, including a dual-model PlannerGrounderAgent that separates what to do from where to click, and a ScrubMiddleware strips PII before any agent sees a screenshot. There's also a standalone GRPO trainer with no openadapt-ml dependency, an OpenEnv-compatible environment, and a four-pass pipeline that turns desktop recordings into structured workflows.

ExtraDup (python -m openadapt_evals.extradup) mutates a MockMed gold write and asks whether a checker notices. Gold is FAIL when the system of record has the wrong cardinality or an extra field. If the checker cannot kill ExtraDup, it cannot underwrite a write. See openadapt_evals/extradup/README.md.

openadapt_evals.reward wires TRL GRPO or verl to a reward endpoint that answers with signed openadapt-types receipts, drops unscored episodes instead of scoring them 0, and never calls a tier-0 read certified. python -m openadapt_evals.reward.proof scores scripted MockMed rollouts with a visual-only and a certified reward, no model needed. See docs/reward/README.md.

Runbooks for the demo-conditioned eval, the full evaluation runner, the UI-Venus grounder endpoint, GRPO training, and writing your own agent are in docs/WORKFLOWS.md, along with the package tree.

Contributing

git clone https://github.com/OpenAdaptAI/openadapt-evals.git
cd openadapt-evals
uv sync --extra dev
uv run pytest tests/ -v

This is research infrastructure and it moves fast. Branches and pull requests only, never a direct push to main. PR titles need conventional commit format, because python-semantic-release parses them to decide the version bump. CLAUDE.md has the development conventions and the WAA benchmark workflow.

License

MIT

Metadata

Release files for openadapt-evals 0.96.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for openadapt-evals 0.96.1
File Size Uploaded
openadapt_evals-0.96.1.tar.gz 62.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for openadapt-evals 0.96.1
File Interpreter ABI Platform
openadapt_evals-0.96.1-py3-none-any.whl Python 3 none any Details

Total release size: 63.2 MB

Release files / openadapt_evals-0.96.1.tar.gz

Download URL openadapt_evals-0.96.1.tar.gz
Size 62.1 MB
Tags Source
SHA-256 checksum
How to use checksums
8ae799afff840608d704bbd0f92675eb2e2af87aea1dd3fe789686cbfdadbc8c
BLAKE2b-256 checksum
How to use checksums
b1e46667bcd5c4e9434f1790a21b08dd1ae22b99266339942b353f033b5b03b9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.

Transparency log

Release files / openadapt_evals-0.96.1-py3-none-any.whl

Download URL openadapt_evals-0.96.1-py3-none-any.whl
Size 1.0 MB
Tags Python 3
SHA-256 checksum
How to use checksums
15fb9891829b6f3b21f4f921073132d85f989a748922d89a7d405463d18f6da8
BLAKE2b-256 checksum
How to use checksums
5ce079a9ffbdd28268ea3dc6a82f6399fb3cde8fa3abed2a6faeb861f95330d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.96.1 This release

2 release files

0.94.0

2 release files

0.93.0

2 release files

0.92.0

2 release files

0.91.2

2 release files

0.91.1

2 release files

0.90.3

2 release files

0.90.2

2 release files

0.90.1

2 release files

0.90.0

2 release files

0.89.1

2 release files

0.89.0

2 release files

0.88.0

2 release files

0.87.2

2 release files

0.87.1

2 release files

0.85.0

2 release files

0.84.0

2 release files

0.83.0

2 release files

0.82.4

2 release files

0.82.3

2 release files

0.82.2

2 release files

0.82.1

2 release files

0.82.0

2 release files

0.81.9

2 release files

0.81.8

2 release files

0.81.7

2 release files

0.81.6

2 release files

0.81.5

2 release files

0.81.4

2 release files

0.81.3

2 release files

0.81.2

2 release files

0.81.1

2 release files

0.81.0

2 release files

0.80.2

2 release files

0.80.1

2 release files

0.80.0

2 release files

0.79.2

2 release files

0.79.1

2 release files

0.79.0

2 release files

0.78.2

2 release files

0.78.1

2 release files

0.77.4

2 release files

0.77.3

2 release files

0.77.2

2 release files

0.77.1

2 release files

0.77.0

2 release files

0.76.3

2 release files

0.76.2

2 release files

0.76.1

2 release files

0.76.0

2 release files

0.75.0

2 release files

0.74.1

2 release files

0.74.0

2 release files

0.73.0

2 release files

0.72.9

2 release files

0.72.8

2 release files

0.72.7

2 release files

0.72.6

2 release files

0.72.5

2 release files

0.72.4

2 release files

0.72.2

2 release files

0.72.1

2 release files

0.72.0

2 release files

0.71.3

2 release files

0.71.2

2 release files

0.71.1

2 release files

0.71.0

2 release files

0.70.2

2 release files

0.70.1

2 release files

0.70.0

2 release files

0.69.1

2 release files

0.69.0

2 release files

0.68.0

2 release files

0.67.0

2 release files

0.66.0

2 release files

0.65.0

2 release files

0.64.1

2 release files

0.64.0

2 release files

0.63.0

2 release files

0.62.0

2 release files

0.61.0

2 release files

0.60.0

2 release files

0.59.3

2 release files

0.59.2

2 release files

0.59.1

2 release files

0.59.0

2 release files

0.58.1

2 release files

0.58.0

2 release files

0.57.0

2 release files

0.56.0

2 release files

0.55.0

2 release files

0.54.0

2 release files

0.53.0

2 release files

0.52.0

2 release files

0.51.1

2 release files

0.51.0

2 release files

0.50.1

2 release files

0.50.0

2 release files

0.49.0

2 release files

0.48.5

2 release files

0.48.4

2 release files

0.48.3

2 release files

0.48.2

2 release files

0.48.1

2 release files

0.48.0

2 release files

0.47.4

2 release files

0.47.3

2 release files

0.47.2

2 release files

0.47.1

2 release files

0.47.0

2 release files

0.46.0

2 release files

0.45.1

2 release files

0.45.0

2 release files

0.44.0

2 release files

0.43.0

2 release files

0.42.1

2 release files

0.42.0

2 release files

0.41.0

2 release files

0.40.0

2 release files

0.39.0

2 release files

0.38.1

2 release files

0.38.0

2 release files

0.37.0

2 release files

0.36.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page