WebPilot
Browser agent CLI. Type a goal; browser-use 0.13.10 drives Chrome, then WebPilot writes a replayable Playwright spec.
The engine is the browser-use service from test-agent-nexus, copied into this package (py/webpilot_engine). It keeps the nexus behaviour: error-recovery system prompt, fast-mode agent tuning, Chrome launch flags, the search/select loop breaker, the 600 s run timeout, browser-use-v2-compact workflow YAML with semantic locators, and the Playwright codegen. There is no dependency on test-agent-nexus, FastAPI, or a database.
The shell is OpenTUI. Headless run is for CI. See docs/ARCHITECTURE.md and HOW_TO_USE.md.
Install
OpenTUI needs Bun 1.3+. Node cannot load the native renderer. Live runs also need uv or Python 3.11+ for the browser-use engine.
The CLI is published on npm as @capagents/webpilot. PyPI package webpilot-cli is a pip launcher for the same webpilot command.
curl -fsSL https://bun.sh/install | bash
curl -LsSf https://astral.sh/uv/install.sh | sh # recommended for the engine
npm install -g @capagents/webpilot
# or
pip install webpilot-cli
webpilot setup # optional: the first live run does this too
webpilot setup creates ~/.webpilot/engine/.venv with browser-use 0.13.10. It uses installed Google Chrome, or installs Chromium when Chrome is missing. WEBPILOT_PYTHON points WebPilot at an existing Python that already has browser-use.
From this repo:
cd WebPilot
bun install
bun src/index.ts init
bun src/index.ts start
One-shot release (same version on both registries; assumes npm and twine are already logged in):
./scripts/publish.sh # current version
./scripts/publish.sh 0.2.0 # bump + publish
./scripts/publish.sh --dry-run # pack only
Privacy
WebPilot sends no telemetry and turns it off in everything it starts:
- browser-use: PostHog telemetry, cloud sync, and the price-list download are off, and the about:blank logo isn't fetched.
- opencode: auto-update, session sharing, OpenTelemetry spans, and the models.dev download are off; opencode uses its bundled model list.
- The repo's test command:
DO_NOT_TRACK=1, plus the opt-outs for Next.js, Nuxt, Astro, Gatsby, Storybook, Turborepo, Angular, .NET, and Cypress.
OTEL_EXPORTER_OTLP_* variables are removed from those processes, and your environment cannot turn any of this back on. The only traffic is to your model endpoint and the sites the agent visits. browser-use also downloads its ad-block and cookie-banner extensions from the Chrome Web Store once.
Configure
init writes:
| File | Purpose |
|---|---|
webpilot.yaml |
Browser, prompts, export, active profile |
llms.json |
Named LLM profiles (azure, openai, ollama, openai_compatible, mock) |
.env.example |
API key names |
webpilot models
webpilot run --goal "Read the homepage" --url https://example.com --profile mock --plain
mock is an offline demo on a synthetic page: no engine, browser, or model. Pick a live profile (Tab in the shell, or llm.active) for a real run.
CLI
webpilot start
webpilot start -g "Go to booking.com and search hotels in Mumbai" -p azure-gpt4o
webpilot run --goal "Get a quote" --url https://example.com --profile azure-gpt4o --headed
webpilot run --goal "..." --url https://example.com --profile mock --plain --out ./out/demo
Inside the shell: /url, /goal, /run, /stop, /headed, /steps, /profile, /export, /memory, /repo, /code, /exit. A line without a slash is the goal and starts the run; the agent chooses which site to open. /url only pins a start page.
Page memory
Every live run records each page it visits in .webpilot/memory next to webpilot.yaml (~/.webpilot/memory when there is no config file): every interactive element, with all of its locators (test id, id, role and name, label, placeholder, alt, title, href, text, CSS, XPath), a confidence % for each, and when it was first seen, last seen, and last used. Confidence rises when a locator keeps matching exactly one element and works when used, and falls when it goes missing or fails.
A goal that passed before is replayed from memory with no model calls. Each step finds its element by the highest-confidence locator that still matches, so a renamed button or a changed id heals itself. The agent then checks the result with one model call (replay: verify), or the run finishes without the model (replay: trust). If a step cannot be found, the agent takes over from that point. Generated specs use the best locator from memory.
webpilot memory # hosts, pages, elements, flows
webpilot memory show booking.com # a host's pages
webpilot memory show https://app.test/login --locators # every locator with confidence
webpilot memory flows --steps # recorded goals and their steps
webpilot memory clear booking.com --yes
webpilot run --goal "..." --replay trust # or --replay off, --no-memory
Coding agent
Point WebPilot at a repo and a run that passes becomes a test in that repo. opencode gets the run's steps, the locators and confidence from page memory, and the generated spec. It studies the repo's framework, page objects, fixtures, and helpers, reuses them, and writes the test. WebPilot then runs the test command itself and sends any failure back to opencode, until the test passes or the attempts run out. The job is done when WebPilot's own run of the test passes.
curl -fsSL https://opencode.ai/install | bash # or: npm i -g opencode-ai
webpilot run --goal "Log in and open the invoices page" --url https://staging.app.test --repo ../app
webpilot code --repo ../app # write a test for the latest run
webpilot code out/20260928-172700 --repo ../app --attempts 3
opencode uses the same llms.json profile as the browser agent (code.model: provider/model picks an opencode model instead). In the shell, /repo <dir> turns it on and /code reruns it on the last run. Set code.repo in webpilot.yaml to make it the default.
Metadata
Release files for webpilot-cli 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| webpilot_cli-0.1.5.tar.gz | 5.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| webpilot_cli-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 11.2 kB
Release files / webpilot_cli-0.1.5.tar.gz
| Download URL | webpilot_cli-0.1.5.tar.gz |
|---|---|
| Size | 5.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c998c975453cfeeb3417239eea2165ca9ba160f3e45b97fa915a2b31a63d6e2c
|
|
BLAKE2b-256 checksum How to use checksums |
5c095225dadcdfdd644b36678a7040d3495dd928af8a5fa813f81e2693dc513c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.13
|
Release files / webpilot_cli-0.1.5-py3-none-any.whl
| Download URL | webpilot_cli-0.1.5-py3-none-any.whl |
|---|---|
| Size | 5.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ffb60534fe26cb8fb2a831e69dcb88797a6a23fb95914801fe8e27f7e5306613
|
|
BLAKE2b-256 checksum How to use checksums |
08134a723f54b09d28f9695453b46bda6ffce35f26f08925d0068ca2fcce18aa
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.13
|