Skip to main content

KNOWS Benchmark

KNOWS is a benchmark for evaluating web agents on realistic, open-ended Google Workspace tasks: writing documents, building spreadsheets, and composing slide decks that require web research, multi-step tool use, and faithful grounding in retrieved sources.

  • 22 task templates × 5 instances = 110 tasks across Google Docs (5 templates), Sheets (9), and Slides (8).
  • Hybrid, white-box evaluation: each task ships a programmatic evaluator (evaluator.py) that scores the produced artifact step by step, combining deterministic checks, fuzzy/tolerance matching, geometric layout tests, document-structure tests, browsing-trace checks, and LLM/VLM judgements.
  • Step-level failure categories: every evaluation step records the mechanism that decided its outcome (StepCategory in src/browsergym/knows/eval/eval_utils/scoring.py), enabling quantitative failure-mode analysis of any run via Result.get_category_summary().

Project page: alexgill321.github.io/KNOWS-benchmark · Dataset: utahnlp/knows-benchmark on Hugging Face

The Hugging Face dataset carries the prompts and the structured evaluation rubric as data, for analysis or task selection. It does not score anything — the evaluators live here.

Repository layout

src/browsergym/knows/
  __init__.py            # BrowserGym task registration (knows.<family>.<n>)
  task.py                # BrowserGym task classes (setup, prompting, grading glue)
  doc_setup.py           # Workspace provisioning & Drive-sharing helpers (CLI included)
  eval/eval_utils/       # Shared evaluation utilities (scoring, text/image/table/chart, LLM judge)
  eval/tasks/<family>/   # One directory per task template
    utils.py             #   template-shared evaluation helpers (where present)
    instance_N/          #   5 instances per template
      task.md            #     the agent prompt (verbatim)
      checkpoints.md     #     human-readable evaluation criteria
      evaluator.py       #     the automated evaluator (standalone CLI)
      data/              #     gold reference assets used by the evaluator
analysis/                # Scripts + data to reproduce the paper's dataset statistics
RUNNING_STANDALONE.md    # Run tasks on ANY agent harness (Comet, proprietary, ...)
ASSETS.md                # External Drive assets some tasks depend on, and how to host your own copies

Quickstart

Option A — Run with the reference harnesses

The maintained agent harnesses live in companion repos:

  • BrowserGym-Knows — a BrowserGym fork with the knows backend, benchmark splits (knows_docs_1, knows_sheets_7, ...), runner scripts, and environment setup. Start here; this repository is consumed as its browsergym/knows submodule.
  • AgentLab-Knows — an AgentLab fork with agent configurations used in the paper. It requires the BrowserGym-Knows setup; see its README.
git clone https://github.com/farhanishmam/BrowserGym-Knows
cd BrowserGym-Knows
git submodule update --init --recursive   # pulls this repo into browsergym/knows/
# then follow BrowserGym-Knows' README for installation and runs

Option B — Run on your own harness

Every evaluator is a standalone script: give your agent the task.md prompt plus a Google file it can edit, record the URLs it visits, then run the evaluator against the resulting file ID. See RUNNING_STANDALONE.md for the complete protocol (provisioning, sharing, prompting, history capture, grading, and output parsing).

Option C — Install as a package

pip install browsergym-knows          # Python >= 3.10
python -m playwright install chromium

This installs the evaluators and registers all 110 tasks as BrowserGym environments (browsergym/knows.<family>.<instance>); it is what the upstream BrowserGym knows backend installs. The ~210 MB of gold evaluation data is not in the wheel: it is downloaded once from this repository's GitHub release on first use into ~/.cache/browsergym-knows/ (override the location with KNOWS_DATA_DIR), or prefetch it explicitly:

python -c "import browsergym.knows; browsergym.knows.ensure_gold_data()"

Add the [local-models] extra to run local judge models instead of the Gemini API. The Setup section below applies to all three options.

Setup (required for all options)

Evaluators read the agent's artifact through the Google Workspace APIs and use a Gemini model as the LLM/VLM judge.

  1. Google Cloud project with the Drive, Docs, Sheets, and Slides APIs enabled.
  2. Evaluator credentials — one of:
    • Service account (recommended): create a service account, download its JSON key to auth-data/service-account.json (or set SERVICE_ACCOUNT_PATH). Every graded file must be shared with the service account's email (see RUNNING_STANDALONE.md — the harnesses do this automatically).
    • OAuth: place an OAuth desktop-client credentials.json in auth-data/; the first run opens a browser consent flow and caches auth-data/token.json.
  3. Judge model key: create a Google AI Studio API key and export GOOGLE_AI_API_KEY (or configure Vertex AI application-default credentials with GOOGLE_CLOUD_PROJECT).
  4. cp .env.example .env and fill in the values (some task families use additional optional keys).
  5. Install dependencies:
pip install -r requirements.txt
python -m playwright install chromium   # only needed for harness runs / doc_setup provisioning

auth-data/ is git-ignored — never commit credentials.

6. Provision write targets (required for two task families)

Most tasks only read from Drive. Two families ask the agent to write into it:

Family What the agent writes
sheets_10_paper_sorting Uploads paper PDFs and Figure 1 screenshots
slides_17_removeimagesaddplaceholders Saves images extracted from a deck, and edits a copy of it

A write destination can't be shared between users, so their prompts ship with {{PLACEHOLDER}} tokens instead of URLs. Run this once before each benchmark pass to create the folders in your Drive and fill the tokens in:

python src/browsergym/knows/eval/tasks/provision_run_targets.py --all

This creates a fresh run_NNNN per instance, so one pass never sees another's uploads:

KNOWS-runs/                                    <- created in your My Drive
  sheets_10_paper_sorting/
    instance_1/run_0001/
      pdfs/                                    <- {{OUTPUT_FOLDER_URL}} points here
      figures/                                    (the prompt names both subfolders)
  slides_17_removeimagesaddplaceholders/
    instance_1/run_0001/
      images/                                  <- {{IMAGES_FOLDER_URL}}
      KNOWS ... run_0001                       <- {{WORKING_COPY_URL}} (a deck copy)
    instance_2/run_0001/
      images/                                  <- {{IMAGES_FOLDER_URL}}
      copies/                                  <- {{OUTPUT_FOLDER_URL}}

Useful flags: --family <name> / --instance <N> to do a subset, --parent_folder_id <id> to build somewhere other than a new KNOWS-runs folder, and --auth oauth|service to choose credentials.

Use OAuth for slides_17. Instance 1 needs a copy of a presentation, and service accounts have no Drive storage quota, so they cannot own files. Folder-only provisioning (sheets_10, slides_17 instances 2–5) works with either credential.

The resulting IDs are written to run_targets.json in the repository root. The harness substitutes them into the prompt at episode start and the evaluators read the same file when grading, so there is nothing to edit by hand. It is git-ignored — it holds locations only you can write to.

If a task runs without provisioning, it fails immediately with the exact command to fix it. Set KNOWS_SKIP_PROVISION=1 to reuse the existing folders (e.g. when re-grading a finished run) and KNOWS_RUN_TARGETS=/path/to/file.json to keep the config elsewhere.

External task assets

Six task families reference source documents/folders on Google Drive from their prompts. Those sources are hosted view-only and work as-is; the two families above additionally need the write targets described in step 6. Full details — including how to rehost the sources from the released assets bundle if a link ever breaks — are in ASSETS.md.

Reproducing paper analyses

analysis/ contains the dataset-statistics script (analyze_task_instances.py) and the step-type taxonomy labels (analysis/taxonomy/). Step-level failure categories are produced natively by the evaluators (Result.get_category_summary(); see eval/eval_utils/scoring.py).

Maintenance, Versioning, and Issue Reporting

KNOWS tasks are curated to be time-agnostic and independent of any single website, but some tasks reference live web pages and Drive-hosted assets that can change over time. Our maintenance policy:

  • Periodic freshness checks. We maintain an internal list of the external URLs and website-extracted gold data our evaluators rely on, and manually verify them at periodic intervals after release (automated checks will replace manual ones once validated against them).
  • Updates and deprecation. If a website change or outage significantly alters a task's difficulty or solvability, we will update the affected task or replace it with a new task of similar complexity and scope.
  • Versioning. Every change to tasks or evaluators is released as a new tagged version on GitHub. Prior versions remain permanently available via their release tags, so results reported against any version stay interpretable and reproducible. Always report the benchmark version (release tag) alongside your results.

Reporting outdated tasks or evaluator issues

Found a task whose referenced website changed, a dead link, or an evaluator that scores incorrectly? Please open a GitHub issue using the provided templates. Include:

  1. the task and instance (e.g. sheets_7_running_analysis/instance_3) and the benchmark version (release tag);
  2. what you observed vs. what you expected (for evaluator issues: the step name and the evaluator's printed output — scores, step details, and category summary);
  3. for outdated-content reports: the affected URL and what changed;
  4. if relevant and shareable: a link to the graded document (shared as view-only).

We triage reports against the policy above; fixes ship as new tagged releases.

Citation

If you use KNOWS, please cite:

@inproceedings{gill2026knows,
  title     = {The Hard Part Comes After Search: Benchmarking Web Agents on
               Synthesizing, Organizing, and Displaying Knowledge},
  author    = {Gill, Alexander and Ishmam, Md Farhan and Nguyen, Xuyen and
               Bhat, Neha and DeYoung, Parker Henry and
               Hashemi Chaleshtori, Fateme and Stringham, Nathan and
               Marino, Kenneth and Marasovi\'{c}, Ana},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}

Please also state which benchmark version you ran (see Maintenance, Versioning, and Issue Reporting).

License

Apache License 2.0 — see LICENSE.

Metadata

Release files for browsergym-knows 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for browsergym-knows 1.2.0
File Size Uploaded
browsergym_knows-1.2.0.tar.gz 1.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for browsergym-knows 1.2.0
File Interpreter ABI Platform
browsergym_knows-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.6 MB

Release files / browsergym_knows-1.2.0.tar.gz

Download URL browsergym_knows-1.2.0.tar.gz
Size 1.6 MB
Tags Source
SHA-256 checksum
How to use checksums
b6446c90d32d8f54dc9770e89650a080775ba8fbaf205a35e7aa8fdf7af96dd2
BLAKE2b-256 checksum
How to use checksums
c74af8df05551ffd521ee00b919f75d9bafad39fb79861c576b7f7ac0feac0b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.19

Release files / browsergym_knows-1.2.0-py3-none-any.whl

Download URL browsergym_knows-1.2.0-py3-none-any.whl
Size 1.9 MB
Tags Python 3
SHA-256 checksum
How to use checksums
51ce51ee7a72b7a794d90fa432e4b87acef3a61db4751216e40001697058dfb1
BLAKE2b-256 checksum
How to use checksums
cc5ab37ed50f99153bfa90332cc9ab4bde92041f7732148fe4e93068961e71ff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.19

Release history Release notifications | RSS feed

This release

1.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page