Skip to main content

Sherpa

Sherpa is a hybrid web agent: a planner reads the page (screenshot + compact controls/content), a screenshot-only grounder picks click coordinates, and Playwright acts. Default models are Qwen3.5-35B-A3B (planner/verifier) and UI-TARS-1.5-7B (grounder) via OpenRouter.

Install (Browser Use–style)

With uv (recommended):

# From GitHub (until published on PyPI):
uv tool install git+https://github.com/AnilPuram/Sherpa.git

# Or from a local checkout:
uv tool install .

# After PyPI publish:
# uv tool install sherpa-agent
# uvx --from sherpa-agent sherpa --help

# One-shot without a permanent install:
uvx --from git+https://github.com/AnilPuram/Sherpa.git sherpa --help

Then:

sherpa install          # Playwright Chromium (+ Linux system deps)
sherpa init             # creates .env if missing
# edit .env → OPENROUTER_API_KEY=
sherpa "Confirm the page heading on https://example.com, then finish."

sherpa install runs the same pattern as Browser Use (uvx playwright install chromium, with --with-deps on Linux). If Chromium is missing when you run a task, Sherpa installs it automatically.

Dev checkout

uv sync
uv run sherpa install
uv run sherpa init
uv run sherpa "Confirm the page heading on https://example.com"

Usage

Config is one required key. Sherpa loads .env automatically (without overriding variables already set in your shell). Optional model/budget overrides are commented in .env.example.

# Start URL can be in the task, or passed explicitly:
sherpa "Confirm the page heading" --url https://example.com

# Headless + machine-readable result:
sherpa "…" --url https://example.com --headless --json

By default the browser is headed, steps print as they run, and artifacts go to artifacts/runs/<timestamp>/.

Verify offline

uv run ruff check .
uv run pytest
uv run sherpa eval

Use --real-model only when you intend paid OpenRouter calls for grounding eval:

uv run sherpa eval --real-model --output artifacts/eval.json

Benchmarks / research

WebVoyager subset tooling, judgment rubrics, run history, and the cross-agent bakeoff live under eval/ and the docs below. Live WebVoyager still requires --real-model so offline scoring stays safe and free.

uv run sherpa webvoyager --output artifacts/webvoyager-offline.json
uv run sherpa webvoyager --real-model --max-cost-usd 1.00 \
  --artifacts artifacts/webvoyager --output artifacts/webvoyager-live.json

See WEBVOYAGER_TESTS_AND_RESULTS.md, WEBVOYAGER_RUN_HISTORY.md, and eval/bakeoff-round2-comparison.md.

How it works

The loop keeps a bounded progress ledger and memories, recovers from grounding/execution errors, and requires a separate verifier to accept every done answer. Defaults: 20 planner steps, 5 consecutive corrections, high planner reasoning effort. Details: ARCHITECTURE.md.

Files

  • src/sherpa/cli.py: simple task CLI + research subcommands
  • src/sherpa/install_browser.py: sherpa install / auto Chromium setup
  • src/sherpa/agent.py: bounded progress, recovery, and hybrid-observation agent loop
  • src/sherpa/models.py: planner, grounder, and verifier OpenRouter boundary
  • src/sherpa/eval.py: point-in-box grounding evaluation
  • src/sherpa/webvoyager.py: offline validation and bounded live WebVoyager subset runner
  • ARCHITECTURE.md: agent flow and component diagrams
  • eval/bakeoff-round2-comparison.md: Sherpa vs Browser Use vs Magnitude bakeoff
  • scripts/bakeoff/: external-agent bakeoff runners

Metadata

Release files for sherpa-agent 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sherpa-agent 0.1.0
File Size Uploaded
sherpa_agent-0.1.0.tar.gz 133.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sherpa-agent 0.1.0
File Interpreter ABI Platform
sherpa_agent-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 173.8 kB

Release files / sherpa_agent-0.1.0.tar.gz

Download URL sherpa_agent-0.1.0.tar.gz
Size 133.1 kB
Tags Source
SHA-256 checksum
How to use checksums
5bb6ab826f2e5731239c71198c2f786941183b47183f157997cc91260ff00ea1
BLAKE2b-256 checksum
How to use checksums
13027b07a7cb3adf1f14cec6d43681d685f024afacb6b08b0aae79cedc7266e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.11

Release files / sherpa_agent-0.1.0-py3-none-any.whl

Download URL sherpa_agent-0.1.0-py3-none-any.whl
Size 40.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e2cad5b9bce1fc2ffddb9e9af30066d65bebf4defcff1a432b3f0d48bbef5643
BLAKE2b-256 checksum
How to use checksums
f8390c63d722c23c324a2d4bf10d9a178a2b59160faf4df0e827e9daa3735d0b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.11

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page