An auditable, low-cost, DOM-only browser agent built on Playwright
Project description
GlassBrowser
An auditable, low-cost, DOM-only browser agent you can actually read and debug.
GlassBrowser drives a real Chromium (via Playwright) to complete natural-language
web tasks. It is built on a small (~5k line) provider-agnostic kernel: a
think → act → observe loop, three-layer context compaction, and an append-only
audit trace. It reasons over a compact DOM snapshot — no screenshots
required — so it stays cheap enough to run well on small, inexpensive models.
Python 3.11+ · Playwright · MIT · providers: Anthropic / OpenAI / any
Anthropic- or OpenAI-compatible endpoint (DeepSeek, local gateways, …)
Why another browser agent?
Most capable browser agents lean on the biggest model + vision + a hosted cloud. That maximizes raw benchmark scores, but it's expensive, opaque, and hard to run yourself. GlassBrowser makes the opposite bet:
| GlassBrowser | Typical "max model + vision + cloud" agents | |
|---|---|---|
| Perception | Compact DOM snapshot (no vision) | DOM + screenshots |
| Cost | Runs on cheap models (long-task context compaction) | Needs a top-tier model |
| Transparency | Full audit trace of every step + exact LLM request | Usually a black box |
| Reproducibility | ba eval mind2web — one command, no API key needed |
Heavy, often cloud-gated |
| Deploy | pip install, runs locally, or embed as a library |
Hosted service / heavy infra |
| Provider | Swap models without code changes | Often locked to one vendor |
It won't top a vision-driven leaderboard. It's for people who want an agent they can run cheaply, embed, and debug — and researchers who need to see exactly what the model saw and why it acted. The name is the promise: a glass box, not a black box.
Highlights
- Auditable & resumable. Every run writes
trace.jsonl(full step log),llm_context.jsonl(the exact request sent each turn) andmessages.jsonl(resume snapshot). You can always answer "what did the model see, and why did it do that?" — and continue a run with--resume. - Low-cost by design. Three-layer context compaction (recent-snapshot elision → lossless Tier-1 shrink → model summary) keeps long tasks inside a small context window, so cheap models stay viable.
- DOM-only perception. A single in-page pass builds an LLM-friendly tree,
stamps interactive elements with
data-ba-ref, pierces open shadow DOM, descends into iframes, and lists<select>options — no pixels needed. - Stale-ref safety. Refs are valid only for the current snapshot; after the page changes the agent is forced to re-observe instead of clicking a ghost.
- Provider-native history. Anthropic
thinkingsignatures and OpenAI reasoningencrypted_contentare replayed verbatim across turns — swap providers without touching the loop. - Skills. Drop a
SKILL.mdin to teach the agent a site's conventions; activate with a/slash-command.
Quickstart
pip install -e . # core: Anthropic + any Anthropic-compatible endpoint
# pip install -e ".[openai]" # optional: OpenAI provider
# pip install -e ".[dev]" # optional: run the tests
playwright install chromium
Run it with whatever model you like — including cheap ones via an Anthropic-compatible endpoint:
# Cheap: DeepSeek (Anthropic-compatible endpoint)
export ANTHROPIC_API_KEY=sk-...
export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic
export BROWSER_AGENT_LLM_MODEL=deepseek-v4-flash
ba chat -m "search Hacker News for 'playwright' and give me the top 3 titles"
# Or Anthropic / OpenAI
export ANTHROPIC_API_KEY=sk-ant-... # Claude (default provider)
ba chat
OPENAI_API_KEY=sk-... ba chat --provider openai --model gpt-4o
Useful commands:
ba chat # interactive REPL, launches a headed Chromium
ba chat -m "…" # one-shot task
ba chat --headless # no visible window
ba chat --cdp http://127.0.0.1:9222 # attach to a Chrome you already have open
ba chat --resume # continue the latest run
ba tools # list the agent's tools
ba sessions # list resumable runs
How it works
- Ref-based interaction.
browser_snapshotreturns a tree of visible salient elements; interactive ones carry aref(e.g.e12).click/fill/select_option/press_keyall target refs. Cheaper still:find_elementfetches just the element you want, andread_textextracts a value with no snapshot at all. - Three-layer compaction. Only the most recent N snapshots stay verbatim in the LLM view; older tool results are losslessly shrunk (Tier 1), and a model summary kicks in only when truly needed (Tier 2) — all KV-cache friendly.
- Provider-agnostic kernel. The loop never touches provider wire formats; each provider owns the shape of its own message history behind small hooks.
browser_agent/
├── agent/ # think → act → observe loop, 3-layer context compaction
├── llm/ # provider-agnostic LLMClient (Anthropic / OpenAI / scripted)
├── actions/ # tool framework + browser tools (snapshot/click/fill/find/read/…)
├── browser/ # Playwright driver + BrowserSession (snapshot refs)
├── eval/ # Mind2Web offline harness
├── trace/ # append-only JSONL audit log + resume snapshots
├── skills/ # SKILL.md loader, /slash-command activation
└── cli.py # `ba` REPL
Reproducible evaluation
Perception quality is measured offline against Mind2Web — no heavy infra, no 600 GB of Docker, and the diagnostic pass needs no API key:
# Perception recall only (free, no API key): does our snapshot expose the
# gold target as an actionable ref? Streams a few tasks straight off HuggingFace.
ba eval mind2web --split test_task --dry-run --limit-tasks 20
# Full scoring (needs a model): element accuracy / operation F1 / step success.
ba eval mind2web --split test_website --limit-tasks 20 --out results.json
On sampled Mind2Web splits our DOM snapshot exposes the gold target as an actionable ref for ~88% (test_task) and ~73% (test_website) of steps — the ceiling that end-to-end accuracy then builds on. This is a single-step, teacher-forced probe: it stresses perception and one-shot grounding, not the full multi-turn loop. (Multi-step benchmarks are on the roadmap.)
Full methodology, metric definitions, offline fixture caching, and reference numbers: docs/mind2web.md.
Tests
pytest # kernel regression + scripted e2e + real-Playwright smoke + eval harness
The Playwright smoke tests auto-skip when the Chromium binary is missing, so the kernel suite stays runnable on machines without a browser.
Examples
examples/ has two self-contained scripts:
zero_key_scripted.py— the full loop with no API key (scripted provider): watch tools execute and the audit trace appear, for free.embed_agent.py— embed the agent in your own Python code: five objects, oneagent.chat(...)call.
Roadmap
- Multi-step benchmarks (WebArena / Online-Mind2Web) to measure the full loop, not just single-step perception — help wanted (needs heavier infra).
- Optional vision channel — the kernel keeps a dormant hook; DOM-only stays the default.
- More site skills.
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file glassbrowser-0.1.0.tar.gz.
File metadata
- Download URL: glassbrowser-0.1.0.tar.gz
- Upload date:
- Size: 113.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
803a8558e9f925d3f1ea8482786230199c4d26c5e81324d5b6fd3f6cfdedb7e1
|
|
| MD5 |
a6d8c5b72e71bcef1153ebc404bd44c0
|
|
| BLAKE2b-256 |
a01b5897eb3f9db9dd34c7f9bedf07ed19176cf001d34a9ad519f7b9944fb957
|
File details
Details for the file glassbrowser-0.1.0-py3-none-any.whl.
File metadata
- Download URL: glassbrowser-0.1.0-py3-none-any.whl
- Upload date:
- Size: 127.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f6380706d6b658d3e3662ca7449c61da7da5410d39bf9d5fc4bfc739b47d58ef
|
|
| MD5 |
576e7665ac31bae0c29520afa8109e2a
|
|
| BLAKE2b-256 |
dcbb8a109e1eef08d8b55e6f88cf7244102f95e390c7008d0b22b75a833f0aa8
|