WebPilot
Browser agent CLI. Say what you want in plain language: a goal to carry out in a browser, test cases to run from Jira, Azure DevOps, Xray or a file, a test to write, or a CI pipeline to set up. WebPilot's agent (opencode) works out what that needs and calls WebPilot's tools. browser-use 0.13.10 drives Chrome, and every run leaves a replayable Playwright spec.
The engine is the browser-use service from test-agent-nexus, copied into this package (py/webpilot_engine). It keeps the nexus behaviour: error-recovery system prompt, fast-mode agent tuning, Chrome launch flags, the search/select loop breaker, the 600 s run timeout, browser-use-v2-compact workflow YAML with semantic locators, and the Playwright codegen. There is no dependency on test-agent-nexus, FastAPI, or a database.
The shell is OpenTUI. webpilot run does the same without the shell, and webpilot suite is the fixed command pipelines call. See docs/ARCHITECTURE.md and HOW_TO_USE.md.
Install
OpenTUI needs Bun 1.3+. Node cannot load the native renderer. Live runs also need uv or Python 3.11+ for the browser-use engine.
The CLI is published on npm as @capagents/webpilot. PyPI package webpilot-cli is a pip launcher for the same webpilot command.
curl -fsSL https://bun.sh/install | bash
curl -LsSf https://astral.sh/uv/install.sh | sh # recommended for the engine
npm install -g @capagents/webpilot
# or
pip install webpilot-cli
webpilot setup # optional: the first live run does this too
webpilot setup creates ~/.webpilot/engine/.venv with browser-use 0.13.10. It uses installed Google Chrome, or installs Chromium when Chrome is missing. WEBPILOT_PYTHON points WebPilot at an existing Python that already has browser-use.
From this repo:
cd WebPilot
bun install
bun src/index.ts init
bun src/index.ts start
One-shot release (same version on both registries; assumes npm and twine are already logged in):
./scripts/publish.sh # current version
./scripts/publish.sh 0.2.0 # bump + publish
./scripts/publish.sh --dry-run # pack only
Privacy
WebPilot sends no telemetry and turns it off in everything it starts:
- browser-use: PostHog telemetry, cloud sync, and the price-list download are off, and its blank-tab logo (from cf.browser-use.com) is replaced by WebPilot's start page, which loads nothing.
- opencode: auto-update, session sharing, OpenTelemetry spans, and the models.dev download are off; opencode uses its bundled model list.
- The repo's test command:
DO_NOT_TRACK=1, plus the opt-outs for Next.js, Nuxt, Astro, Gatsby, Storybook, Turborepo, Angular, .NET, and Cypress.
OTEL_EXPORTER_OTLP_* variables are removed from those processes, and your environment cannot turn any of this back on. The only traffic is to your model endpoint and the sites the agent visits. browser-use also downloads its ad-block and cookie-banner extensions from the Chrome Web Store once.
Configure
init writes:
| File | Purpose |
|---|---|
webpilot.yaml |
Browser, prompts, export, active profile |
llms.json |
Named LLM profiles (azure, openai, ollama, openai_compatible, mock) |
.env.example |
API key names |
webpilot models
webpilot run "Read the homepage" --url https://example.com --profile mock --plain
mock is an offline demo on a synthetic page: no engine, browser, or model. Pick a live profile (Tab in the shell, or llm.active) for a real run.
CLI
webpilot start
webpilot start -g "Go to booking.com and search hotels in Mumbai" -p azure-gpt4o
webpilot run "Get a quote on example.com for a family of four" --profile azure-gpt4o --headed
webpilot run "Run SHOP-12 and SHOP-14, 2 browsers at once"
webpilot run "Write a Playwright test for the login flow on staging.app.test" --repo ../app
webpilot run "Test the refund rules on https://acme.atlassian.net/wiki/spaces/QA/pages/123"
webpilot login shop # sign in by hand once; runs reuse the session
webpilot run --direct "Read the homepage" --url https://example.com --out ./out/demo # one browser goal, no agent
There are no commands to pick a source or a mode. WebPilot's agent reads the request and decides: a browser goal, test cases (Jira keys, Azure DevOps ids or test plans, Xray, TestRail, Zephyr Scale or qTest cases, files in the working directory, or cases pasted into the message), requirements to turn into cases (a Confluence page, a story), a test, or a pipeline. When a reference could mean more than one thing, it asks. The agent picks the site to open; --url only pins a start page.
Inside the shell, type the request and press Enter. The only commands are for the session: /stop, /headed, /profile, /clear, /help, /exit. Everything the other commands do can be asked for in the shell too, since each one is a tool the agent calls: "heal the checkout tests", "check the smoke cases once", "monitor them every 10 minutes", "sign in to shop", "forget the shop session", "which profiles do I have?". Results appear as cards in the transcript, and when a step needs you (signing in by hand), a Your turn prompt waits until you pick Done, continue or Cancel.
The agent needs opencode (curl -fsSL https://opencode.ai/install | bash) and a live profile. It uses the active llms.json profile, or code.model: provider/model. With the mock profile, or without opencode, the request runs as one browser goal.
webpilot run exits 0 when the job is done, 1 when it failed, 2 when it was stopped, and 3 when the agent needs more information.
Page memory
Every live run records each page it visits in .webpilot/memory next to webpilot.yaml (~/.webpilot/memory when there is no config file): every interactive element, with all of its locators (test id, id, role and name, label, placeholder, alt, title, href, text, CSS, XPath), a confidence % for each, and when it was first seen, last seen, and last used. Confidence rises when a locator keeps matching exactly one element and works when used, and falls when it goes missing or fails.
A goal that passed before is replayed from memory with no model calls. Each step finds its element by the highest-confidence locator that still matches, so a renamed button or a changed id heals itself. The agent then checks the result with one model call (replay: verify), or the run finishes without the model (replay: trust). If a step cannot be found, the agent takes over from that point. Generated specs use the best locator from memory.
webpilot memory # hosts, pages, elements, flows
webpilot memory show booking.com # a host's pages
webpilot memory show https://app.test/login --locators # every locator with confidence
webpilot memory flows --steps # recorded goals and their steps
webpilot memory clear booking.com --yes
webpilot run "..." --replay trust # or --replay off, --no-memory
Tests
Point WebPilot at a repo and a run that passes becomes a test in that repo. The agent gets the run's steps, the locators and confidence from page memory, and the generated spec. It studies the repo's framework, page objects, fixtures, and helpers, reuses them, and writes the test. WebPilot then runs the test command itself and hands any failure back, until the test passes or the attempts run out. The test counts only when WebPilot's own run of it passes.
curl -fsSL https://opencode.ai/install | bash # or: npm i -g opencode-ai
webpilot run "Log in and open the invoices page on staging.app.test" --repo ../app
webpilot code --repo ../app # write a test for the latest run
webpilot code out/20260928-172700 --repo ../app --attempts 3
With --repo (or code.repo in webpilot.yaml) every passed run becomes a test unless you say otherwise. Without one, the agent writes a test only when you ask for one; --no-code turns tests off.
Fixing broken tests
When the app changes, tests break. webpilot heal runs the repo's tests, and for each failing test the agent re-runs its flow in the real browser. If the flow still works, the test is out of date: the agent updates its steps and locators (in the page object, when that's where they live) from what the browser just did and from page memory. If the browser fails at the same step, the app is broken there: the test is left alone and the failure is reported as an app bug. Tests are never deleted, skipped, loosened, or given longer timeouts to make them pass.
WebPilot re-runs every healed test and then the whole command itself; a test counts as healed only when WebPilot's own run of it passes.
webpilot heal --repo ../app # code.test_command, or the agent finds the command
webpilot heal "npx playwright test tests/checkout" --repo ../app
webpilot heal -- npx playwright test --project=chromium
webpilot run "The login tests broke after the redesign, fix them" --repo ../app # the agent heals them the same way
The report is heal.md (and heal.json) in out/heal/<time>/, with each test's verdict, the files changed, the browser runs, and the test output before and after. webpilot heal exits 0 when the tests pass (or nothing needed healing), 1 when some still fail, and 2 when stopped.
Test cases and suites
Run existing test cases instead of typing goals: ask for them by name ("run SHOP-12", "run test plan 12 suite 34", "run the cases in smoke.yaml") or paste them into the message. Each case becomes one goal: its steps are what the browser agent does and its expected results are what it checks, and the agent's own verdict decides pass or fail. With a repo set, every passed case also gets a test.
webpilot suite runs cases the same way every time, with no agent choosing anything, which is what pipelines need. It takes explicit references:
webpilot cases cases.yaml # what a source gives, without running anything
webpilot suite cases.yaml --parallel 3 # run them, 3 browsers at once
webpilot suite ado:plan=12/suite=34 --repo ../app
webpilot suite jira:jql="project = SHOP AND labels = smoke"
webpilot suite xray:plan=SHOP-100
webpilot suite "Open example.com and check the title says Example Domain"
| Source | Reference |
|---|---|
| Plain text | the text itself, quoted |
| Text or Markdown file | .txt / .md, one case per block between --- lines |
| CSV | .csv with id, title, url, description, preconditions, tags, action, data, expected columns |
| YAML / JSON | .yaml / .json, a list of {id, title, url, preconditions, steps: [{action, data, expected}]} |
| Gherkin | .feature, one case per scenario (and per Examples row) |
| Azure DevOps | ado:123,456 (test case ids), ado:plan=12, ado:plan=12/suite=34 |
| Jira | jira:SHOP-1,SHOP-2, jira:jql=<query> (Cloud and Server/Data Center) |
| Xray | xray:SHOP-1, xray:jql=<query>, xray:plan=SHOP-100 (Cloud and Server/Data Center) |
| TestRail | testrail:C12,C13, testrail:run=45, testrail:plan=7, testrail:suite=3 |
| Zephyr Scale (Cloud) | zephyr:SHOP-T1,SHOP-T2, zephyr:cycle=SHOP-R5, zephyr:folder=12 |
| qTest | qtest:1234 (test case ids), qtest:cycle=88, qtest:suite=9 |
For a case with numbered steps the browser agent reports a verdict for every step (passed, failed, or not run, with what it actually saw), not only for the case.
Results go back where the cases came from, with the step verdicts: an Azure DevOps test run for a plan (or a comment on the work item), a Jira comment, an Xray Test Execution with the failure screenshot, a TestRail run (the run the cases came from, or a new one) with step results and the screenshot, a Zephyr Scale test cycle, or qTest test logs. --no-write-back or sources.write_back: false turns it off.
Bugs
With sources.bugs, each failed case gets a bug in Jira or Azure DevOps with the steps and verdicts, where it ended, the build link, and the last screenshot, linked to the test case. WebPilot tags the bug with a fingerprint of the case, so the next failure adds a "still failing" comment to the open bug instead of filing another, and a pass adds a note that it passes again.
sources:
bugs: { system: jira, project: SHOP } # or system: azure_devops (uses azure_devops.project)
Reports
Every suite writes report.html, junit.xml, and suite.json to out/suites/<time>/, with each case's run folder beside them; every single run writes its own report.html too. The report is one HTML file with the screenshots embedded, so you can open it straight from disk, attach it to a build, or mail it; no server is needed. It has an overview (verdict, totals, step checks, tokens, cost, page memory savings, browser errors, accessibility issues, bugs filed, what changed since the last run, results over earlier runs, the cases that need attention, results by source and tag, the slowest cases, and what was posted where), a test results page (each case's step table with expected and actual results, why it failed, its last passing screenshot beside this run's, browser errors per step, page memory with healed locators, the browser agent's steps with a screenshot each and the clicked element outlined, a timing waterfall, accessibility issues, and the trace, spec, and workflow to read in place), a failures page, since last run, trends and flaky tests from the earlier suite runs in the same out/suites/ folder, time and cost (model vs browser time, a parallel worker timeline, cost per case from the model's price in llms.json), coverage by requirement and group from the test system, accessibility checks for every page visited, and the environment (model, browser, WebPilot version, CI build, branch, commit). It exports Markdown and JSON and prints cleanly. See samples/reports/nightly/20260929-020005/report.html for an example. report.screenshots: failures keeps screenshots for failed runs only, off drops them; report.accessibility: false skips the page checks. On GitHub Actions the suite also writes a job summary.
Signing in
Sites that need a login go under logins:. The browser agent types credentials from the environment without ever seeing them (it only sees placeholders such as shop_password), fills authenticator codes from a TOTP secret, and keeps the signed-in session in .webpilot/sessions/ so the next run starts signed in. Replays and generated Playwright specs read the same environment variables, and specs compute TOTP codes themselves.
logins:
shop:
url: https://staging.shop.test/login
username_env: SHOP_USER
password_env: SHOP_PASSWORD
totp_env: SHOP_TOTP_SECRET # optional
For single sign-on, captchas, or security keys, sign in by hand once: webpilot login shop opens a browser, you sign in, press Enter, and the session is saved. webpilot logins lists logins and sessions; webpilot logins clear shop forgets one. Session files are private to your user and have their own .gitignore.
Azure DevOps, Jira, and other systems
WebPilot's agent connects to each system's MCP server, configured from the same sources: settings in webpilot.yaml: organization, project, site URL, and the names of the environment variables that hold the tokens. You don't write any MCP config yourself.
| System | MCP server | Needs |
|---|---|---|
| Azure DevOps Services | Microsoft's Azure DevOps MCP server (npx @azure-devops/mcp), with work items, boards, test plans, pipelines, and repos |
Node.js, and ADO_PAT or az login |
| Jira Cloud, Server/Data Center | mcp-atlassian (uvx), with issues, JQL search, comments, and transitions. mcp: rovo uses Atlassian's hosted server instead (Cloud, and an admin must allow API tokens) |
uv; JIRA_EMAIL and JIRA_API_TOKEN (Cloud) or a personal access token (Server/DC) |
| Confluence | mcp-atlassian, shared with Jira (or its own when Jira isn't set up); Atlassian's hosted server covers it on the same site | confluence: true on a Jira Cloud site, or confluence.url + tokens |
| Xray, TestRail, Zephyr Scale, qTest | none; WebPilot's API | XRAY_CLIENT_ID / XRAY_CLIENT_SECRET, TESTRAIL_EMAIL / TESTRAIL_API_KEY, ZEPHYR_API_TOKEN, QTEST_TOKEN |
| Anything else | servers you add under mcp: |
sources:
azure_devops: { org_url: https://dev.azure.com/acme, project: Shop } # token in ADO_PAT
jira: { url: https://acme.atlassian.net } # JIRA_EMAIL + JIRA_API_TOKEN
mcp:
github:
url: https://api.githubcopilot.com/mcp/
headers: { Authorization: "Bearer {env:GITHUB_TOKEN}" }
Tokens never go into the config: servers get them through {env:NAME} references that opencode fills in, and the agent never sees them. webpilot mcp list shows what your settings give. When the agent starts, the shell and webpilot run print which servers connected.
Whatever has no MCP server, or whose server didn't connect, goes through WebPilot's own tools, which call the REST APIs directly: find_test_cases and run_test_cases for test cases and posting results back, read_confluence for Confluence pages, and query_api to read anything else from Azure DevOps, Jira, Xray, or Confluence. query_api only reads (GET requests and Xray GraphQL queries) and only reaches the configured hosts. sources.azure_devops.mcp: false or sources.jira.mcp: false uses the API only.
Requirements work too: "write and run test cases for the checkout rules page in Confluence" makes the agent read the page, write one case per rule or acceptance criterion (with the main negative paths), show them, and run them. It saves them into a test system only when you ask.
The agent creates or changes items (a bug, a comment, a status change) only when you ask it to; bugs for failed cases come from sources.bugs, not from the agent.
Pipelines
Ask for the pipeline you want and the agent builds it:
webpilot run "Add a GitHub Actions workflow that runs the smoke cases in cases.yaml on every pull request"
webpilot run "Set up an Azure DevOps pipeline that runs test plan 12 every night and posts results back"
webpilot run "The WebPilot pipeline on main is failing, fix it"
It works in the repo (the working directory, or --repo): it reads the repo's existing CI, writes or updates the workflow so it installs Bun, uv and WebPilot, calls webpilot suite <source...> --no-code --plain, and publishes junit.xml and the report, and validates it (actionlint when it is installed). Then it commits to a new branch named webpilot/ci-<topic>, stores the keys the pipeline needs as secrets, triggers the run with gh or az, watches it, and fixes failures until the run passes (up to code.max_attempts failed runs).
Guard rails:
- Pushing goes through WebPilot's
push_branchtool, which refuses the default branch,main, andmaster, and never forces.git pushitself is blocked for the agent. - Secrets go through
set_ci_secret: the agent names an environment variable and WebPilot passes its value straight togh secret setoraz pipelines variable. The agent never sees the value. - It needs
gh(GitHub, logged in) orazwith theazure-devopsextension (Azure DevOps, logged in). Without them it still writes, validates, and commits the pipeline, then tells you what to run.
Monitoring
webpilot monitor runs checks on a schedule and tells your team when one breaks. A check is any test case webpilot suite takes: a YAML file, a Jira key, an Azure DevOps plan. Each round replays the flows from page memory, so a flow that passed before needs one model call to confirm the result (or none with replay: trust).
monitor:
checks: [checks/smoke.yaml, "jira:jql=labels = monitor"]
every: 15m
fail_after: 2 # two failing rounds in a row before the first alert, to ride out a flaky run
remind_every: 2h # repeat while it keeps failing
alerts:
teams: { webhook_env: TEAMS_WEBHOOK_URL }
email: { host: smtp.office365.com, port: 587, to: [qa-team@example.com] } # SMTP_USERNAME, SMTP_PASSWORD
webpilot monitor --test-alerts # a sample alert, to check the webhook and SMTP settings
webpilot monitor # every 15 minutes until Ctrl+C
webpilot monitor --once # one round, for cron or a scheduled pipeline
An alert goes out when a check starts failing, again every remind_every while it fails, and once when it recovers. One round sends one message covering every check that changed. Teams gets an Adaptive Card (both incoming webhooks and Workflows webhooks accept it) and email gets the failing checks' screenshots attached. Each alert links to the round's report: the local file, or <report_url>/<round>/report.html when you publish out/monitor/ somewhere. A round that cannot run at all (a test system is down, the engine will not start) alerts as the check "WebPilot monitor".
Rounds go to out/monitor/<time>/, the same as suites, so the report shows trends and flaky checks over the rounds; the last keep rounds (100) are kept. state.json remembers which checks are failing and since when, so with --once in a pipeline, cache that folder between runs. Only one monitor can use a folder at a time. --once exits 0 when every check passed and 1 otherwise. Results are posted back to the test system only with write_back: true. In the shell, "monitor the smoke cases every 10 minutes" runs the rounds in the background while you keep working, and each round shows as a card; they stop when the shell closes.
MCP
webpilot mcp serves WebPilot's tools over stdio for other agents such as Claude Code, Cursor, or Codex:
{ "mcpServers": { "webpilot": { "command": "webpilot", "args": ["mcp", "--repo", "/path/to/app"] } } }
Tools: browser_run, find_test_cases, run_test_cases, run_details, run_test, page_memory, read_confluence, query_api, push_branch, set_ci_secret, and one tool per command: write_test, heal_tests, monitor, logins, webpilot_setup. Browser runs and suites log their progress to stderr, and long calls send progress notifications.
Metadata
Release files for webpilot-cli 0.1.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| webpilot_cli-0.1.8.tar.gz | 20.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| webpilot_cli-0.1.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.7 kB
Release files / webpilot_cli-0.1.8.tar.gz
| Download URL | webpilot_cli-0.1.8.tar.gz |
|---|---|
| Size | 20.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
005c8e9b47b61ae1bdcf3b5a3d542310af313a2fdbb6c6afc224543896ee14d4
|
|
BLAKE2b-256 checksum How to use checksums |
d5f5aa3adce8feb4cf75291888b829badbaaf8d4f7ac5ca97b0ce712702c3a09
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.13
|
Release files / webpilot_cli-0.1.8-py3-none-any.whl
| Download URL | webpilot_cli-0.1.8-py3-none-any.whl |
|---|---|
| Size | 12.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5c914729e27145ca5e8b77cc05065aee09313b286325574e4ce6978bfcbd6145
|
|
BLAKE2b-256 checksum How to use checksums |
d8890347cd1780bffb08af9045358968eee8df1bd298eab5fcf49c9f0f42dbe5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.13
|