Skip to main content

openadapt-flow

CI PyPI Python License: MIT

Record a GUI workflow once. Replay it deterministically, locally, for free. A model only touches the script to repair it.

One demonstration, two UIs, same compiled workflow — the right side self-heals under a theme it has never seen

Real screenshots from the two runs in docs/showcase/. Left: the UI the demo was recorded on. Right: a theme it had never seen — each step re-resolves through OCR or geometry, and each fix is written back to the script as a reviewable diff. Zero model calls on either side.

Safety, stated honestly. It halts instead of guessing, and we measure how often it could still resolve the wrong target under UI drift — then publish it. Read what it doesn't do yet and how we test it, including five adversarial rounds against our own wrong-target check.

Try it

pip install openadapt-flow

openadapt-flow demo-record --out rec                     # record a demonstration
openadapt-flow compile rec --out bundle --name my-task   # compile it
openadapt-flow lint bundle                               # report coverage gaps
openadapt-flow certify bundle --policy clinical-write    # refuse it if unsafe
openadapt-flow replay bundle                             # replay: local, $0
openadapt-flow replay bundle --drift theme               # drift the UI, watch it heal

On the first command that needs a browser, openadapt-flow downloads the Chromium build Playwright needs (a one-time ~150MB fetch) — no separate playwright install chromium step. Prefer the fast, isolated installs uvx openadapt-flow … or uv tool install openadapt-flow. In air-gapped or CI environments that pre-provision the browser, set OPENADAPT_FLOW_NO_AUTO_INSTALL=1 to disable the auto-download.

The last two commands serve the bundled MockMed demo app and write an illustrated REPORT.md per run.

Record your own app

record --url opens a headed browser on YOUR app and watches what you do — real clicks, typing, key presses and scrolls — writing the same recording format compile consumes. Perform the workflow, then press Ctrl-C (or close the window) to finish:

openadapt-flow record --url https://your.app --out rec   # do the task, Ctrl-C
openadapt-flow compile rec --out bundle --name my-task
openadapt-flow replay bundle --url https://your.app       # replay it

Pass --url to replay to run against your own app; recorded parameter values are the defaults and --param overrides them.

Secrets never get recorded. A input[type=password] field (or any field named with --secret <name>) is a secret parameter: its value is never written to the recording, the events log, the compiled bundle, or the saved frames (its region is redacted). At replay it is injected from the environment and a missing one fails fast:

openadapt-flow record --url https://your.app --out rec --secret password
export OPENADAPT_FLOW_SECRET_PASSWORD='…'                 # supplied at replay
openadapt-flow replay bundle --url https://your.app

Compiled is not the same as certified safe. lint reports a bundle's coverage gaps (clicks that act with no identity check, steps that assert nothing, write steps left mis-classified) with a severity each; certify enforces a policy and exits nonzero — refusing the bundle before it deploys — when it fails. Risk is auto-classified at compile time (write-shaped clicks — save/submit/create/delete/... — become irreversible, which arms the low-confidence refusal), and two example policies ship: a permissive default and a strict clinical-write.yaml. See docs/LIMITS.md for what the heuristic does and does not catch.

How it works

Computer-use agents re-reason through your task with a large model on every run. That's the right shape for a task nobody has automated before, and the wrong one for the 500th referral this month. openadapt-flow compiles the demonstration instead.

Each compiled step carries a template crop, an OCR label, geometry landmarks, and postconditions derived from what the demo actually changed on screen. At replay time a resolution ladder tries them in order: local template match, global template match, OCR, landmark geometry, then (optionally) a grounding model. Healthy scripts never leave the first rung. Milliseconds, no model calls, no per-run cost.

When the UI drifts, a lower rung still finds the target and the fix lands in the bundle as a diff you can review. When the screen stops matching expectations entirely, the run halts with a report instead of guessing, and steps tagged irreversible won't act on a low-confidence match at all.

The runtime is vision-only (PNG in, clicks and keys out) behind a small Backend protocol. The reference backend is a headless browser, which is why the whole loop runs in CI with no OS permissions. Desktop and RDP backends are adapters to come, not rewrites.

Proof

Every CI run records a demonstration, compiles it, and checks:

Scenario Outcome
Baseline replay ×3 all steps template rung, 0 heals, 0 model calls
Theme drift succeeds; 8/8 anchors healed; healed bundle replays clean
Moved buttons succeeds via global template search
Renamed buttons succeeds via landmark geometry
Surprise modal fails loudly, naming the violated postcondition
Non-recorded parameter substituted and verified by OCR of the final screen

Artifacts: baseline run report · theme-drift run report.

Compiled workflows can also be emitted as Agent Skills or MCP servers (emit-skill / emit-mcp), so other agents can invoke them.

Benchmark

OpenEMR: compiled replay vs computer-use agent, latency and cost

The lead result is on a real third-party app: the official OpenEMR public demo (fake patients only, resets daily). We ran an 18-step add-patient-note workflow both ways — log in, find a patient, scroll a dense dashboard, add a note — with a distinct note value each run and the same OCR success check on both arms: 20 compiled replays against 10 runs of a claude-sonnet-5 computer-use agent. Compiled went 20/20 at 39.2s (p50) with zero model calls; the agent went 10/10 at 70.4s (p50), about $0.55 per run at list price ($5.52 total for the 10 runs, with prompt caching and hard cost caps enforced in the harness). It's a shared public demo that other users mutate and that resets daily — not CI-reproducible, and the sample is small. Correctness alone (no agent arm, 5/5 fresh browsers, zero model calls, closed-loop scrolling) is in docs/showcase-openemr/FINDINGS.md. Full numbers, methodology, and caveats: benchmark/openemr/BENCHMARK.md.

For a controlled, CI-reproducible comparison — the methodology anchor — we ran the bundled MockMed task both ways on 2026-07-08 with the same OCR success check: 100 compiled replays against 20 runs of the same agent. Both arms went 100 for 100 and 20 for 20, so on an app this simple the story isn't success rate. It's that a compiled replay finishes in 4.9s (p50; 5.1s p95) with zero model calls, while the agent takes 37.5s (p50; 43.4s p95) at about $0.27 per run at list price, every run, forever. Full numbers, methodology, and caveats: benchmark/BENCHMARK.md.

Status

v0: 864 tests, drift matrix in CI. Solid for the reference browser backend. DESIGN.md has the module contracts; docs/L1_INTEGRATION.md covers feeding layered clinical-data platforms.

Privacy (PHI)

For regulated deployments, PHI scrubbing on the persist/log paths is provided by the optional privacy extra (Presidio-backed openadapt-privacy):

pip install 'openadapt-flow[privacy]' && python -m spacy download en_core_web_trf
export OPENADAPT_FLOW_SCRUB=on          # scrub REPORT.md + logs, fail closed

The shareable REPORT.md and console logs are scrubbed; the compiled bundle and report.json keep literal identifiers on purpose (identity check + audit trail) and are protected by a documented boundary. Identity crops sent to the on-prem VLM appliance are deliberately not scrubbed — the control there is on-prem-only + no-retention. Full map: docs/PRIVACY.md.

Development

git clone https://github.com/OpenAdaptAI/openadapt-flow && cd openadapt-flow
pip install -e '.[dev]'
playwright install chromium  # optional: else auto-downloads on first launch
pytest -q

The demo GIF is generated from real run artifacts by scripts/make_demo_gif.py. MIT license.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

openadapt_flow-0.19.1.tar.gz (15.3 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

openadapt_flow-0.19.1-py3-none-any.whl (594.7 kB view details)

Uploaded Python 3

File details

Details for the file openadapt_flow-0.19.1.tar.gz.

File metadata

  • Download URL: openadapt_flow-0.19.1.tar.gz
  • Upload date:
  • Size: 15.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for openadapt_flow-0.19.1.tar.gz
Algorithm Hash digest
SHA256 94bd2598a535211f23d411c2cbc4ef92b8cd6881f235c9712ab9dacc60ebc29d
MD5 c16df60af10f1cf657e3e988308aaf29
BLAKE2b-256 8e307134256ef23160acc0fefd197ade38032a8fb059f2b7d5bd733e4378adfa

See more details on using hashes here.

Provenance

The following attestation bundles were made for openadapt_flow-0.19.1.tar.gz:

Publisher: release.yml on OpenAdaptAI/openadapt-flow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file openadapt_flow-0.19.1-py3-none-any.whl.

File metadata

  • Download URL: openadapt_flow-0.19.1-py3-none-any.whl
  • Upload date:
  • Size: 594.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for openadapt_flow-0.19.1-py3-none-any.whl
Algorithm Hash digest
SHA256 273d393041cae471450caac139aa65133a58d4523f67da30dcef61aad4f26e86
MD5 dc7713f03ccb7a7aaae10c4eaba008d8
BLAKE2b-256 037a3a6d7c99bf197c18c5d4418cf6cc0b9c30035be1dd0b3bb8a31cb790dc1e

See more details on using hashes here.

Provenance

The following attestation bundles were made for openadapt_flow-0.19.1-py3-none-any.whl:

Publisher: release.yml on OpenAdaptAI/openadapt-flow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.32.0

2 files

1.31.0

2 files

1.30.0

2 files

1.29.0

2 files

1.28.0

2 files

1.27.1

2 files

1.27.0

2 files

1.26.0

2 files

1.25.1

2 files

1.25.0

2 files

1.24.0

2 files

1.23.0

2 files

1.22.0

2 files

1.21.0

2 files

1.20.2

2 files

1.20.1

2 files

1.20.0

2 files

1.19.0

2 files

1.18.1

2 files

1.18.0

2 files

1.17.2

2 files

1.17.1

2 files

1.17.0

2 files

1.16.1

2 files

1.16.0

2 files

1.15.0

2 files

1.14.1

2 files

1.14.0

2 files

1.13.0

2 files

1.12.2

2 files

1.12.1

2 files

1.12.0

2 files

1.11.0

2 files

1.10.1

2 files

1.10.0

2 files

1.9.1

2 files

1.9.0

2 files

1.8.1

2 files

1.8.0

2 files

1.7.3

2 files

1.7.2

2 files

1.7.1

2 files

1.7.0

2 files

1.6.0

2 files

1.5.1

2 files

1.5.0

2 files

1.4.0

2 files

1.3.0

2 files

1.2.0

2 files

1.1.0

2 files

1.0.0

2 files

0.26.0

2 files

0.25.0

2 files

0.24.0

2 files

0.23.0

2 files

0.22.0

2 files

0.21.2

2 files

0.21.1

2 files

0.21.0

2 files

0.20.1

2 files

0.20.0

2 files

This release

0.19.1 This release

2 files

0.19.0

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page