KNOWS Benchmark
KNOWS is a benchmark for evaluating web agents on realistic, open-ended Google Workspace tasks: writing documents, building spreadsheets, and composing slide decks that require web research, multi-step tool use, and faithful grounding in retrieved sources.
- 22 task templates × 5 instances = 110 tasks across Google Docs (5 templates), Sheets (9), and Slides (8).
- Hybrid, white-box evaluation: each task ships a programmatic evaluator (
evaluator.py) that scores the produced artifact step by step, combining deterministic checks, fuzzy/tolerance matching, geometric layout tests, document-structure tests, browsing-trace checks, and LLM/VLM judgements. - Step-level failure categories: every evaluation step records the mechanism that decided its outcome (
StepCategoryinsrc/browsergym/knows/eval/eval_utils/scoring.py), enabling quantitative failure-mode analysis of any run viaResult.get_category_summary().
Project page: alexgill321.github.io/KNOWS-benchmark · Dataset: utahnlp/knows-benchmark on Hugging Face
The Hugging Face dataset carries the prompts and the structured evaluation rubric as data, for analysis or task selection. It does not score anything — the evaluators live here.
Repository layout
src/browsergym/knows/
__init__.py # BrowserGym task registration (knows.<family>.<n>)
task.py # BrowserGym task classes (setup, prompting, grading glue)
doc_setup.py # Workspace provisioning & Drive-sharing helpers (CLI included)
eval/eval_utils/ # Shared evaluation utilities (scoring, text/image/table/chart, LLM judge)
eval/tasks/<family>/ # One directory per task template
utils.py # template-shared evaluation helpers (where present)
instance_N/ # 5 instances per template
task.md # the agent prompt (verbatim)
checkpoints.md # human-readable evaluation criteria
evaluator.py # the automated evaluator (standalone CLI)
data/ # gold reference assets used by the evaluator
analysis/ # Scripts + data to reproduce the paper's dataset statistics
RUNNING_STANDALONE.md # Run tasks on ANY agent harness (Comet, proprietary, ...)
ASSETS.md # External Drive assets some tasks depend on, and how to host your own copies
Quickstart
Option A — Run with the reference harnesses
The maintained agent harnesses live in companion repos:
- BrowserGym-Knows — a BrowserGym fork with the
knowsbackend, benchmark splits (knows_docs_1,knows_sheets_7, ...), runner scripts, and environment setup. Start here; this repository is consumed as itsbrowsergym/knowssubmodule. - AgentLab-Knows — an AgentLab fork with agent configurations used in the paper. It requires the BrowserGym-Knows setup; see its README.
git clone https://github.com/farhanishmam/BrowserGym-Knows
cd BrowserGym-Knows
git submodule update --init --recursive # pulls this repo into browsergym/knows/
# then follow BrowserGym-Knows' README for installation and runs
Option B — Run on your own harness
Every evaluator is a standalone script: give your agent the task.md prompt plus a Google file it can edit, record the URLs it visits, then run the evaluator against the resulting file ID. See RUNNING_STANDALONE.md for the complete protocol (provisioning, sharing, prompting, history capture, grading, and output parsing).
Option C — Install as a package
pip install browsergym-knows # Python >= 3.10
python -m playwright install chromium
This installs the evaluators and registers all 110 tasks as BrowserGym environments (browsergym/knows.<family>.<instance>); it is what the upstream BrowserGym knows backend installs. The ~210 MB of gold evaluation data is not in the wheel: it is downloaded once from this repository's GitHub release on first use into ~/.cache/browsergym-knows/ (override the location with KNOWS_DATA_DIR), or prefetch it explicitly:
python -c "import browsergym.knows; browsergym.knows.ensure_gold_data()"
Add the [local-models] extra to run local judge models instead of the Gemini API. The Setup section below applies to all three options.
Setup (required for all options)
Evaluators read the agent's artifact through the Google Workspace APIs and use a Gemini model as the LLM/VLM judge.
- Google Cloud project with the Drive, Docs, Sheets, and Slides APIs enabled.
- Evaluator credentials — one of:
- Service account (recommended): create a service account, download its JSON key to
auth-data/service-account.json(or setSERVICE_ACCOUNT_PATH). Every graded file must be shared with the service account's email (see RUNNING_STANDALONE.md — the harnesses do this automatically). - OAuth: place an OAuth desktop-client
credentials.jsoninauth-data/; the first run opens a browser consent flow and cachesauth-data/token.json.
- Service account (recommended): create a service account, download its JSON key to
- Judge model key: create a Google AI Studio API key and export
GOOGLE_AI_API_KEY(or configure Vertex AI application-default credentials withGOOGLE_CLOUD_PROJECT). cp .env.example .envand fill in the values (some task families use additional optional keys).- Install dependencies:
pip install -r requirements.txt
python -m playwright install chromium # only needed for harness runs / doc_setup provisioning
auth-data/ is git-ignored — never commit credentials.
6. Provision write targets (required for two task families)
Most tasks only read from Drive. Two families ask the agent to write into it:
| Family | What the agent writes |
|---|---|
sheets_10_paper_sorting |
Uploads paper PDFs and Figure 1 screenshots |
slides_17_removeimagesaddplaceholders |
Saves images extracted from a deck, and edits a copy of it |
A write destination can't be shared between users, so their prompts ship with
{{PLACEHOLDER}} tokens instead of URLs. Run this once before each benchmark pass to create the
folders in your Drive and fill the tokens in:
python src/browsergym/knows/eval/tasks/provision_run_targets.py --all
This creates a fresh run_NNNN per instance, so one pass never sees another's uploads:
KNOWS-runs/ <- created in your My Drive
sheets_10_paper_sorting/
instance_1/run_0001/
pdfs/ <- {{OUTPUT_FOLDER_URL}} points here
figures/ (the prompt names both subfolders)
slides_17_removeimagesaddplaceholders/
instance_1/run_0001/
images/ <- {{IMAGES_FOLDER_URL}}
KNOWS ... run_0001 <- {{WORKING_COPY_URL}} (a deck copy)
instance_2/run_0001/
images/ <- {{IMAGES_FOLDER_URL}}
copies/ <- {{OUTPUT_FOLDER_URL}}
Useful flags: --family <name> / --instance <N> to do a subset, --parent_folder_id <id> to
build somewhere other than a new KNOWS-runs folder, and --auth oauth|service to choose
credentials.
Use OAuth for
slides_17. Instance 1 needs a copy of a presentation, and service accounts have no Drive storage quota, so they cannot own files. Folder-only provisioning (sheets_10,slides_17instances 2–5) works with either credential.
The resulting IDs are written to run_targets.json in the repository root. The harness substitutes
them into the prompt at episode start and the evaluators read the same file when grading, so there
is nothing to edit by hand. It is git-ignored — it holds locations only you can write to.
If a task runs without provisioning, it fails immediately with the exact command to fix it. Set
KNOWS_SKIP_PROVISION=1 to reuse the existing folders (e.g. when re-grading a finished run) and
KNOWS_RUN_TARGETS=/path/to/file.json to keep the config elsewhere.
External task assets
Six task families reference source documents/folders on Google Drive from their prompts. Those sources are hosted view-only and work as-is; the two families above additionally need the write targets described in step 6. Full details — including how to rehost the sources from the released assets bundle if a link ever breaks — are in ASSETS.md.
Reproducing paper analyses
analysis/ contains the dataset-statistics script (analyze_task_instances.py) and the step-type taxonomy labels (analysis/taxonomy/). Step-level failure categories are produced natively by the evaluators (Result.get_category_summary(); see eval/eval_utils/scoring.py).
Maintenance, Versioning, and Issue Reporting
KNOWS tasks are curated to be time-agnostic and independent of any single website, but some tasks reference live web pages and Drive-hosted assets that can change over time. Our maintenance policy:
- Periodic freshness checks. We maintain an internal list of the external URLs and website-extracted gold data our evaluators rely on, and manually verify them at periodic intervals after release (automated checks will replace manual ones once validated against them).
- Updates and deprecation. If a website change or outage significantly alters a task's difficulty or solvability, we will update the affected task or replace it with a new task of similar complexity and scope.
- Versioning. Every change to tasks or evaluators is released as a new tagged version on GitHub. Prior versions remain permanently available via their release tags, so results reported against any version stay interpretable and reproducible. Always report the benchmark version (release tag) alongside your results.
Reporting outdated tasks or evaluator issues
Found a task whose referenced website changed, a dead link, or an evaluator that scores incorrectly? Please open a GitHub issue using the provided templates. Include:
- the task and instance (e.g.
sheets_7_running_analysis/instance_3) and the benchmark version (release tag); - what you observed vs. what you expected (for evaluator issues: the step name and the evaluator's printed output — scores, step details, and category summary);
- for outdated-content reports: the affected URL and what changed;
- if relevant and shareable: a link to the graded document (shared as view-only).
We triage reports against the policy above; fixes ship as new tagged releases.
Citation
If you use KNOWS, please cite:
@inproceedings{gill2026knows,
title = {The Hard Part Comes After Search: Benchmarking Web Agents on
Synthesizing, Organizing, and Displaying Knowledge},
author = {Gill, Alexander and Ishmam, Md Farhan and Nguyen, Xuyen and
Bhat, Neha and DeYoung, Parker Henry and
Hashemi Chaleshtori, Fateme and Stringham, Nathan and
Marino, Kenneth and Marasovi\'{c}, Ana},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}
Please also state which benchmark version you ran (see Maintenance, Versioning, and Issue Reporting).
License
Apache License 2.0 — see LICENSE.
Metadata
Release files for browsergym-knows 1.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| browsergym_knows-1.2.0.tar.gz | 1.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| browsergym_knows-1.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.6 MB
Release files / browsergym_knows-1.2.0.tar.gz
| Download URL | browsergym_knows-1.2.0.tar.gz |
|---|---|
| Size | 1.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b6446c90d32d8f54dc9770e89650a080775ba8fbaf205a35e7aa8fdf7af96dd2
|
|
BLAKE2b-256 checksum How to use checksums |
c74af8df05551ffd521ee00b919f75d9bafad39fb79861c576b7f7ac0feac0b0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.19
|
Release files / browsergym_knows-1.2.0-py3-none-any.whl
| Download URL | browsergym_knows-1.2.0-py3-none-any.whl |
|---|---|
| Size | 1.9 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
51ce51ee7a72b7a794d90fa432e4b87acef3a61db4751216e40001697058dfb1
|
|
BLAKE2b-256 checksum How to use checksums |
cc5ab37ed50f99153bfa90332cc9ab4bde92041f7732148fe4e93068961e71ff
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.19
|