Semantic Browser
Version 1.6.0 (Beta) · PyPI · Changelog · License: MIT
Semantic Browser turns live Chromium pages into compact, numbered text views for LLM agents. The agent reads one short view
(content and its options), replies with one command like click 12, and the runtime executes it on that exact element.
$ sb goto https://shop.example/search?q=anvil
navigated to … -> 'Search: anvil'
@ Search: anvil — https://shop.example/search?q=anvil [screen 1/1]
# Catalogue
[1 input:search "Search catalogue"="anvil"] [2 button]Search
- [3]Anvil Classic £49.99 In stock
- [4]Anvil Pro £89.50 Low stock
$ sb click 4
clicked [4] 'Anvil Pro' -> now at 'Anvil Pro' (https://shop.example/p/anvil-pro)
Less confusion, less hallucination, dramatically less cost: on 8 live sites the model reads a median ~2.5k tokens per step (1.3k with
--budget 2500) instead of ~15k for a raw accessibility snapshot, at ~1.4 s per two-step task with 16/16 success
(benchmark report).
Give it to your agent (30 seconds)
pip install "semantic-browser[managed]" && semantic-browser install-browser
sb guide # the playbook an agent should read: loop, view syntax, recovery, bot walls, safety (~1.5k tokens)
- Any LLM agent: paste the output of
sb guideinto its system prompt, or tell it: "To use the web, runsb guide, then drivesb." - Claude / Cursor skills:
mkdir -p ~/.claude/skills/semantic-browser && sb guide --skill > ~/.claude/skills/semantic-browser/SKILL.md - Rules files (
AGENTS.md,.cursorrules): add the same one-liner. - The guide is tested against the code (every verb covered, size-capped) and also lives at docs/agent_guide.md.
What the agent gets: one outcome line plus a numbered page per command, batching with sb do, whole-page find, honest failure
messages, bot-wall and CAPTCHA signalling, and rules that treat page text as untrusted input and default to read-only.
Why Semantic Browser
- One compact view — page content with inline
[n]refs; shadow DOM, iframes (cross-origin too), custom div widgets and duplicate labels handled. - Exact-element execution — a ref is bound to one element handle; stale refs fail loudly, never retarget (no silent
<body>clicks). - Fast — event-driven settling: ~0.3–0.5 s to a usable page; whole 3–5-step tasks in ~0.3–1 s on the local hard-pattern suite (v1.3.2: ~4 s).
- Built-in blockers — cookie banners and modals are detected with a ranked dismiss hint; bot walls and rate limits are flagged (
! This looks like a bot/verification page) so an agent stops instead of hammering. - CAPTCHA assist —
sb captchaproduces a numbered, annotated image (or PDF) for a vision model and clicks the tiles it names. - Four interfaces —
sbagent CLI (persistent daemon), Python API, legacysemantic-browserCLI, and HTTP service.
Install
pip install "semantic-browser[managed]"
semantic-browser install-browser
For service mode: pip install "semantic-browser[server]"
Quickstart
Agent CLI (sb, new in 1.4; do batches in 1.5)
sb goto news.ycombinator.com # first call starts a background Chromium (headful); later calls reuse it
sb click 12 # act on the [12] you see in the view (or: sb click "comments")
sb find "price" # search the whole page (table rows come with their column headers)
sb do "type 3 boots --enter" "click Football" view # several steps, ONE call (stops at the first failure)
sb captcha --pdf # annotated image of a CAPTCHA for a vision model
sb stop
Full verb list, view syntax, sessions, --cdp attach and security model: docs/agent_cli.md.
Interactive portal
semantic-browser portal --url https://example.com --headless
Python
import asyncio
from semantic_browser import ManagedSession
from semantic_browser.models import ActionRequest
async def main() -> None:
session = await ManagedSession.launch(headful=False)
runtime = session.runtime
await runtime.navigate("https://example.com")
obs = await runtime.observe(mode="summary")
print(obs.planner.room_text)
first_link = next((a for a in obs.available_actions if a.op == "open"), None)
if first_link:
result = await runtime.act(ActionRequest(action_id=first_link.id))
print(result.status, result.observation.page.url)
await session.close()
asyncio.run(main())
LLM Agent Loop (Minimal)
async def agent_loop(url: str, task: str) -> None:
session = await ManagedSession.launch(headful=False)
runtime = session.runtime
await runtime.navigate(url)
obs = await runtime.observe(mode="summary")
for step in range(25):
action_id = call_your_llm(obs.planner.room_text, task) # returns one action ID
if action_id == "done":
break
result = await runtime.act(ActionRequest(action_id=action_id))
obs = result.observation
await session.close()
Full worked examples for OpenAI, Anthropic, and more: Integration Examples
Documentation
| Document | What it covers |
|---|---|
Agent guide (sb guide) |
The playbook to give an AI: how to drive sb well, recover, handle bot walls, stay safe |
Agent CLI (sb) |
v1.4: the verbs, how to read the view, sessions, attach to a running Chrome, security model |
| CAPTCHA assist | Detect → annotate (PNG/PDF) → answer by tile number; what was verified and honest limits |
| Dogfood benchmark | v1.3.2 vs v1.4 vs raw Playwright vs httpx, hard-pattern suite, 8 live sites, gated sites, CAPTCHA demo |
| System architecture | Data flow, invariants, extension points |
| Getting Started | Install, first run, interactive portal, Python/CLI/service quickstarts |
| Planner Contract | The exact interface between Semantic Browser and an LLM planner — what the planner receives, what it should reply, how to handle blockers, failures, and stopping |
| Integration Examples | End-to-end examples: OpenAI chat, OpenAI function-calling, Anthropic tool use, HTTP service, CDP attach, error handling patterns |
| API Reference | Every public class, method, model, and field — ManagedSession, SemanticBrowserRuntime, Observation, StepResult, ActionDescriptor, configuration, errors |
| Runtime Modes | Decision table for ephemeral/persistent/clone/attach/service modes, headful vs headless, ownership semantics |
| Real Profiles | Using real Chromium profiles for login persistence, SSO, clone mode, safety guarantees, common pitfalls |
| Benchmark Protocol | How benchmark numbers are produced and validated |
| Versioning | Version numbering scheme |
| Publishing | PyPI publish checklist |
| Changelog | Full release history |
How It Works
Chromium page ──one JS walk per frame──▶ semantic nodes + reading-order flow ──▶ text view with inline [n] refs
▲ │
│ agent reads the view, replies with one verb ▼
└── runtime executes on the exact element handle, settles, reports the outcome ◀── `click 12`
- Observe — one DOM pass per frame (shadow DOM, same- and cross-origin iframes) builds nodes and a reading-order flow; occlusion and overlay detection mark what is really clickable.
- Render — the view keeps content and controls together, collapses navigation, windows long pages and ranks overlay dismiss controls.
- Act — a ref is bound to one element handle for the whole session, so a stale ref fails loudly instead of clicking something else. The runtime waits for the page to settle and prints one outcome line (
-> now at …,page changed (+3/-1 lines),no visible change). - Repeat — or batch known steps with
sb do. Architecture: docs/system_arch.md. The older planner/observe/actAPI (room text, action IDs) is still available for Python and service use: Planner Contract.
Benchmarks
Latest dogfood run (details, protocol, raw data and caveats: docs/benchmarks/2026-10-06-dogfood-v1.4.md):
| Method | Local hard-pattern suite | 8 live sites | Tokens read / step (live, median) |
|---|---|---|---|
| v1.3.2 | 12/27 | 14/16 | ~0.8k |
| Raw Playwright accessibility snapshot | 21/27 | 11/16 | ~15k (max 143k) |
| v1.4 | 27/27 | 16/16 | ~2.5k (1.3k with --budget 2500) |
Real model, live sites (me, Claude Sonnet 5.5, one command per call, journaled with scripts/dogfood/sbj.py;
report): round 1 on 1.4.0 completed 11 of 13 tasks (median 4 calls; it found a wrong-price bug and several missing-control bugs);
after the fixes the final round completed 7 of 7 (median 1 call, ~0.26k tokens read; part of that is the new do batching). Reddit and Hacker News blocked this IP in every mode and are reported as such.
The first table is a scripted stand-in for the model; the real-model numbers above are a handful of tasks on one machine. Neither is a universal guarantee. Protocol: docs/benchmark_protocol.md. Manifest: benchmarks/manifest.json.
CLI Reference
semantic-browser version # Show version
semantic-browser doctor # Verify installation
semantic-browser install-browser # Download Chromium
semantic-browser launch --headless # Start a session
semantic-browser attach --cdp <ws-url> # Attach to running Chrome
semantic-browser portal --url <url> # Interactive exploration REPL
semantic-browser observe --session <id> --mode summary
semantic-browser act --session <id> --action <action_id>
semantic-browser inspect --session <id> --target <target_id>
semantic-browser navigate --session <id> --url <url>
semantic-browser back --session <id>
semantic-browser forward --session <id>
semantic-browser reload --session <id>
semantic-browser diagnostics --session <id>
semantic-browser export-trace --session <id> --out trace.json
semantic-browser serve --host 127.0.0.1 --port 8765 --api-token <token>
What's New in v1.6.0
sb guide— the agent playbook ships inside the package (sb guide,sb guide --skillfor a drop-in SKILL.md). Tested for verb coverage and size.- Bot walls are named — Fastly/Cloudflare/DuckDuckGo-style gates and rate limits get a
! This looks like a bot/verification pageline with the right next step, also on long pages when the controls are covered. Found live: PyPI "Client Challenge", DuckDuckGo "bots use DuckDuckGo too". - Search boxes in headers stay search boxes — collapsed header/nav lines used to print an
<input>like a link ([4]Search with DuckDuckGo); fillable controls now keep their full form and are never cut by the "(+N more)" cap. - Better
find— whole-word hits first (MITno longer drowns in "commit"), and a short hit (a bare price) carries its card title. - Unknown-verb errors list the verbs from the code and point at
sb guide.
What's New in v1.5.0
Everything here came from driving sb as a real model (Claude Sonnet 5.5) on live sites, then fixing what hurt
(report):
sb do "step" "step" …— a whole GOV.UK visa wizard (8 steps) is one call; median task went from 4 calls to 1.- Correct text on hard markup — Amazon prices no longer lose their decimal (
£569→£5.69); transparent radios/checkboxes (GOV.UK) have refs; delegated-handler widgets (jQuery UI datepicker Prev/Next) are controls. - Honest overlays — the dismiss hint ranks close/reject/short-accept, never "Continue", sign-in or pay-to-reject;
view --allshows the page behind. - Better
find— centred snippets, real<th>column headers, works behind an overlay. Labels you type (click Save settings) resolve, and ambiguous ones show where each candidate goes. - CAPTCHA — the image is captured once the page is still (dynamic reCAPTCHA grids replace tiles after Verify).
- Stale background sessions restart themselves after an upgrade. Release tooling:
scripts/publish.sh.
Full list: CHANGELOG.md. Note: 1.4.0 was never published to PyPI; 1.5.0 includes it.
Earlier releases
1.4.0 introduced the sb agent CLI, the one-pass view engine, fast settling, CAPTCHA assist and safe --cdp attach (never published on its own; 1.5.0 was the first release to include it). Full history: CHANGELOG.md.
Contributing
See CONTRIBUTING.md for development setup and PR expectations.
License
MIT — see LICENSE.
Metadata
Release files for semantic-browser 1.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| semantic_browser-1.6.0.tar.gz | 228.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| semantic_browser-1.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 352.4 kB
Release files / semantic_browser-1.6.0.tar.gz
| Download URL | semantic_browser-1.6.0.tar.gz |
|---|---|
| Size | 228.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8853ed868f4e9afee0a5519195c124820ad711d384f0bd6b893fc2c3e367bd21
|
|
BLAKE2b-256 checksum How to use checksums |
1002916820e7d638422a87f1bc5052af21d29717d75fed141d7f5040c0dfbd8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|
Release files / semantic_browser-1.6.0-py3-none-any.whl
| Download URL | semantic_browser-1.6.0-py3-none-any.whl |
|---|---|
| Size | 123.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
688f075a2e44eb81b5a069ddf36a8a5ed2d86bf9179b2ea373acc6dc84d6d141
|
|
BLAKE2b-256 checksum How to use checksums |
272ef569f563f5afcca5aaaf454d8cd9388696bf4d11986396eab3fdd4124c99
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|