Hermes Browserless Web Extract
Give your Hermes agent the ability to read any web page as a real browser would — powered by Browserless.io's headless Chrome infrastructure.
When your agent calls web_extract(url), this plugin opens a real Chromium instance on Browserless's servers, waits for the page to fully render (including JavaScript, SPAs, lazy-loaded content), and returns the complete HTML back to your agent. It's the difference between downloading a skeleton and seeing the actual page.
Why use this plugin?
Real browser rendering. A plain HTTP GET request misses everything JavaScript generates — single-page apps, dynamic content, lazy-loaded images, client-side navigation. Browserless runs a real headless Chrome that executes every <script> tag, applies every CSS animation, and loads every deferred resource. What you get back is what a human sees.
No browser infrastructure to manage. You don't install Chrome, Chromium, or Puppeteer on your machine. You don't manage browser processes, memory limits, or headless flags. Browserless runs the browsers on their infrastructure. You send a simple REST request and get HTML back. The plugin handles all the HTTP plumbing.
JavaScript-heavy sites work. Documentation sites, blogs, forums, and marketing pages that rely on client-side rendering (React, Vue, Svelte, Next.js, etc.) are invisible to basic HTTP fetchers. Browserless renders them fully. This is especially important for modern technical content — many API docs, changelogs, and RFC pages are SPAs.
Extract-only by design. This plugin does one thing: fetch rendered HTML from a URL. It doesn't search. It doesn't summarize. It hands the raw page content to Hermes, which can then reason over it, summarize it, or compare it against other sources. Pair it with a search provider (like the Gemini web search plugin, Brave, or SearXNG) and each tool does its job well.
Graceful error handling. Every URL gets either content or a clear error message. Network timeouts, HTTP 4xx/5xx responses, empty pages — all handled per-URL, so one bad link never kills an entire batch extraction.
Zero-code configuration. Set two environment variables, add a few lines to your config file, and restart. There's no SDK to initialize, no browser flags to tune, and no Puppeteer scripts to write.
What is Browserless?
Browserless.io is a hosted headless browser service. They run and manage Chromium instances in the cloud so you don't have to. You send HTTP requests describing what you want (a screenshot, a PDF, a page's rendered HTML) and they handle the browser lifecycle — launching, navigating, waiting for render, and tearing down.
Their /content REST endpoint (which this plugin uses) takes a URL and returns the fully-rendered HTML. It's the simplest possible interface: POST a URL, get back HTML. No WebSocket connections, no CDP protocol, no browser management code.
Key features of Browserless:
- Real Chrome/Chromium instances — not a lightweight emulation, not a DOM parser guessing at JS output
- Configurable wait strategies — wait for specific selectors, events, timeouts, or functions before returning
- Resource blocking — skip images, fonts, and media for faster extraction
- Stealth mode and proxy support — for sites with basic bot detection
- Free tier available — start without a credit card
What you'll need
- Hermes agent installed and working (any platform — CLI, desktop, Telegram, Discord, etc.)
- A Browserless.io account (free to start, takes 2 minutes)
- Your Browserless API token (from the dashboard)
- Your Browserless instance URL (provided at signup, or use the shared default)
That's it. No other accounts, no browser installs, no headless flags to debug.
Getting a Browserless account and token
- Go to browserless.io in your browser
- Click "Get Started" or "Sign Up"
- Create an account (email + password, or GitHub/Google SSO)
- Once logged in, go to the Account page
- Under API Tokens, copy your token — it will look like a long random string
- Under Service Domain, note your instance URL. It usually looks like
https://production-sfo.browserless.io
The free tier gives you a limited number of concurrent sessions and monthly usage hours. You'll know if you need to upgrade — Browserless will email you when you approach the limit.
Important: Treat your API token like a password. Anyone with it can launch browsers on your behalf and consume your usage quota. Don't commit it to git, don't paste it into chat logs, and don't share it publicly.
Installation
One command (recommended)
From inside your Hermes environment's virtual environment:
pip install hermes-browserless-web-extract
This pulls the plugin from PyPI and installs its sole dependency (httpx).
From source
git clone https://github.com/your-username/hermes-browserless-web-extract.git
cd hermes-browserless-web-extract
pip install .
Editable install (for development)
git clone https://github.com/your-username/hermes-browserless-web-extract.git
cd hermes-browserless-web-extract
pip install -e .
An editable install means changes you make to the source code are picked up immediately without reinstalling. Use this if you plan to modify or contribute to the plugin.
Configuration
Everything happens in two files inside your ~/.hermes/ directory. No code changes, no init scripts, no environment hacks.
Step 1 — Add your credentials
Create or open ~/.hermes/.env in any text editor.
Add these two lines:
BROWSERLESS_TOKEN=your-api-token-here
BROWSERLESS_URL=https://production-sfo.browserless.io (or your local URL)
What these mean:
BROWSERLESS_TOKEN— Your Browserless API token. Required.BROWSERLESS_URL— The base URL of your Browserless instance. Defaults tohttps://production-sfo.browserless.ioif not set. If you purchased a dedicated instance or signed up through a regional endpoint, use that URL instead. If you self-host Browserless, point this at your own server.
Gotchas to avoid:
- Do not wrap values in quotes (
BROWSERLESS_TOKEN="abc"— wrong, the quotes become part of the value) - Do not add spaces around the
=sign (BROWSERLESS_TOKEN = abc— wrong) - Do not reuse tokens from other services. Browserless tokens are specific to Browserless.
- Make sure the
.envfile is plain text, not rich text (use Notepad, vim, nano, or VS Code — not Word or TextEdit in rich mode)
Step 2 — Enable the plugin
Open ~/.hermes/config.yaml in any text editor. Add hermes-browserless-web-extract to the plugins.enabled list:
plugins:
enabled:
- hermes-browserless-web-extract
If you already have other plugins enabled, add this one to the existing list:
plugins:
enabled:
- some-other-plugin
- hermes-browserless-web-extract
If no plugins section exists yet in your config, create it at the top level of the file:
# ... other config options above ...
plugins:
enabled:
- hermes-browserless-web-extract
# ... other config options below ...
Step 3 — Tell Hermes to use Browserless for extraction
In the same config.yaml, under the web section, set the extract backend to browserless:
web:
extract_backend: browserless
Here's what a complete config.yaml might look like after setup (your file will have its own sections — that's fine, you only need to add the plugins and web parts):
plugins:
enabled:
- hermes-browserless-web-extract
web:
extract_backend: browserless
Step 4 — Restart Hermes
Quit your Hermes agent completely and restart it. On the next question that triggers a web_extract call, the agent will use Browserless.
Plugin discovery happens at startup. If you don't restart, Hermes won't see the new plugin.
Verifying it works
The quickest test is to ask your agent to extract a page that uses JavaScript:
> extract the content of https://news.ycombinator.com
> what does https://react.dev/learn say on the first page?
> read and summarize https://docs.python.org/3/whatsnew/3.13.html
If the agent returns current, rendered content with clear structure, the plugin is working.
You can also check from the Hermes tools picker:
hermes tools
Navigate to Web Extract — you should see Browserless listed with a green checkmark and the "API key required" badge.
What success looks like
The agent should be able to read and reason about the actual page content — headings, paragraphs, code blocks, links — not just metadata or title tags. If you ask "what are the top 5 stories on Hacker News?" and the agent can list them, it's working.
What to expect from JavaScript-heavy pages
Some pages take a few seconds to fully render. This is normal. The plugin waits up to 30 seconds (configurable) for the page to settle before returning the HTML. If a page is particularly slow, you might notice a small delay, but the content should be complete.
How it works
When your agent calls web_extract("https://example.com"), here's what happens:
- Hermes dispatches to this plugin. It looks up the active extract backend (
browserless) and callsprovider.extract(["https://example.com"]). - The plugin builds a request. It takes your token and URL, constructs a POST to
$BROWSERLESS_URL/content?token=$BROWSERLESS_TOKEN, and sends a JSON body with the target URL plus browser configuration (timeout, resource blocking). - Browserless launches a headless Chrome instance. On their infrastructure, a real Chromium process starts. It navigates to the URL, executes all JavaScript, waits for network requests to settle, and captures the final DOM state.
- The rendered HTML is returned. Browserless sends back the complete
<html>...</html>string — not just the initial server response, but the fully rendered page including all JS-generated content. - The plugin passes it to Hermes. The HTML is returned in the standard extract result format, with the URL, title, raw content, and metadata. Hermes can then feed it to the LLM for summarization, comparison, or any other reasoning task.
The raw HTTP request looks like this:
POST https://production-sfo.browserless.io/content?token=YOUR_TOKEN
Content-Type: application/json
{
"url": "https://example.com",
"waitForTimeout": 30000,
"bestAttempt": true,
"rejectResourceTypes": ["image", "media", "font"],
"rejectRequestPattern": [".*\\.(png|jpg|jpeg|gif|webp|svg|ico|css|woff2?).*"]
}
Why block images? Text extraction doesn't need images, fonts, CSS, or media files. Blocking them makes pages load faster and reduces bandwidth — you get the same HTML content in less time. If you ever need to extract images (e.g., for vision-capable models), you can override this behavior by setting rejectResourceTypes to an empty list (see Advanced configuration).
Using alongside other web plugins
This plugin is extract-only. It does not search the web — it only fetches content from URLs you already have.
This is intentional. You can pair it with a search provider for a complete web pipeline:
Example: Gemini for search, Browserless for extract
plugins:
enabled:
- hermes-gemini-web-search
- hermes-browserless-web-extract
web:
search_backend: gemini
extract_backend: browserless
With this setup, when your agent needs to find something, Gemini searches Google and returns results. When it wants to read a specific page, Browserless fetches the fully-rendered HTML.
Example: Brave for search, Browserless for extract
plugins:
enabled:
- hermes-brave-web-search
- hermes-browserless-web-extract
web:
search_backend: brave-free
extract_backend: browserless
Any search provider (Brave, SearXNG, Tavily, Firecrawl, Gemini) can be paired with Browserless for extraction. The plugins are independent and Hermes routes each tool call to the right provider.
Using a single backend for both
If you have a provider that does both search and extract (like Firecrawl or Gemini), you can use web.backend as a shared default and override just one capability. For example, use Firecrawl for everything but Browserless for extract on JS-heavy pages:
web:
backend: firecrawl
extract_backend: browserless
Advanced configuration
Extraction timeout
Control how long Browserless waits for the page to fully render before returning whatever HTML is available. Default is 30000 ms (30 seconds):
web:
extract_backend: browserless
browserless_timeout: 60000
Increase this for slow-loading pages (e.g., SPAs with heavy data fetching). Decrease it if you prefer faster-but-maybe-incomplete results.
Note: The value is in milliseconds. Common settings:
10000— 10 seconds, for fast static pages30000— 30 seconds, the default, works for most pages60000— 60 seconds, for heavy SPAs or rate-limited APIs120000— 2 minutes, rarely needed, for very slow backends
Custom Browserless instance URL via config
Instead of setting BROWSERLESS_URL in .env, you can set it in config.yaml:
web:
extract_backend: browserless
browserless_url: https://my-custom-instance.browserless.io
The environment variable takes precedence over the config file value, which makes .env the preferred place for secrets and the config file the right place for instance-specific overrides.
Running browserless locally (self-hosted)
If you run Browserless via Docker on your own machine:
web:
extract_backend: browserless
browserless_url: http://localhost:3000
With matching .env:
BROWSERLESS_TOKEN= # can be empty for local instances without auth
BROWSERLESS_URL=http://localhost:3000
Per-capability routing
Browserless can be used as the shared backend for both search and extract, even though it only supports extract. Hermes will fall back gracefully for search:
web:
backend: browserless
This works: search calls will fail with a clear error, extract calls will work normally. For better UX, explicitly set search_backend to a provider that supports search and extract_backend to browserless.
Troubleshooting
"BROWSERLESS_TOKEN environment variable not set"
The plugin can't find your API token. Check that:
~/.hermes/.envexists and contains a line starting withBROWSERLESS_TOKEN=- There are no extra spaces or quotes around the value (correct:
BROWSERLESS_TOKEN=abc123, incorrect:BROWSERLESS_TOKEN="abc123"orBROWSERLESS_TOKEN = abc123) - You restarted Hermes after adding the key (environment variables are read at startup)
- The file has a trailing newline (some editors strip it; add a blank line at the end to be safe)
If you've verified all of the above and it still fails, try setting the variable directly in your shell before launching Hermes:
export BROWSERLESS_TOKEN=your-token-here
hermes
This bypasses the .env file entirely and confirms the variable is set in the process environment.
"BROWSERLESS_URL not set" or connection failures
If you got the token right but requests fail to connect:
- Verify you set
BROWSERLESS_URLin.env(orbrowserless_urlinconfig.yaml) - If using the default (
https://production-sfo.browserless.io), make sure that domain is reachable from your network:curl -I https://production-sfo.browserless.io - If using a custom instance URL, verify it starts with
https://orhttp://
"Browserless HTTP 401: Unauthorized"
Your token is invalid, expired, or missing. Get a new one from the Browserless account page. Tokens can be rotated from the dashboard — if you previously rotated your token, the old one stops working immediately.
Also check: the token is passed as a query parameter (?token=...), not as an Authorization header. If you're inspecting the request, you should see the token in the URL, not in the headers.
"Browserless HTTP 403: Forbidden"
Your account doesn't have access to the requested resource. Common causes:
- You're on the free tier and have exceeded your concurrent session limit. Wait a moment and retry.
- Your account is in a trial period that has expired. Check the Browserless dashboard for account status.
- You're using a dedicated instance URL but sending requests through the shared endpoint, or vice versa.
"Browserless returned empty content"
The page may have blocked the headless browser, or the page finished loading before any content was rendered. Try:
- Increase the timeout. Add
browserless_timeout: 60000to yourconfig.yamlto give the page more time to render. - Check the page manually. Open the URL in your own browser. Does it load? Does it require login? Is it behind a paywall or CAPTCHA?
- Try the /unblock endpoint. Browserless has an unblock API for sites with bot detection. This plugin uses
/contentby default. For heavily-protected sites, you may need a different Browserless endpoint or a more sophisticated extraction tool. - Check if the site blocks headless browsers. Some sites (major social media platforms, banks, ticket vendors) aggressively detect and block automated browsers. Browserless handles basic protections, but sophisticated anti-bot systems may still succeed. For these cases, consider using the official API of the site or accepting that extraction won't work.
The agent is not using the plugin
- Did you restart Hermes after making configuration changes? Plugin discovery happens at startup — changes to
config.yamlare not picked up dynamically. - Check that
hermes-browserless-web-extractis listed underplugins.enabledinconfig.yaml. A missing entry means the plugin is installed but inactive. - Check that
web.extract_backendis set tobrowserless. If it's set to something else, that provider handles extraction instead. - Run
hermes toolsand look under the Web Extract section. If Browserless isn't listed, the plugin isn't being discovered — check the installation.
"hermes-browserless-web-extract is not enabled"
Hermes found the plugin package but hasn't been told to activate it. Add it to your config:
plugins:
enabled:
- hermes-browserless-web-extract
Then restart Hermes. The message should disappear.
"Could not import hermes_browserless_web_extract"
The plugin isn't installed in your active Python environment. Verify:
pip show hermes-browserless-web-extract
If nothing shows up, install it:
pip install hermes-browserless-web-extract
If you're using a virtual environment for Hermes, make sure you're installing into that same environment. Run which python or which pip to confirm you're in the right one.
Development
Setup
git clone https://github.com/your-username/hermes-browserless-web-extract.git
cd hermes-browserless-web-extract
pip install -e .
Running tests
pytest
Tests use mocks — no API token or network access is needed to run the test suite. The tests verify:
- Provider availability detection (both env vars present, only one, neither)
- Capability flags (search disabled, extract enabled)
- Successful extraction (mocked HTTP response)
- Multi-URL extraction (batch correctness)
- Error handling (missing token, HTTP errors, empty responses)
Project structure
hermes-browserless-web-extract/
├── pyproject.toml # Package metadata and build config
├── README.md
├── LICENSE
├── .gitignore
├── hermes_browserless_web_extract/
│ ├── __init__.py # Plugin entry point (calls register)
│ └── provider.py # BrowserlessWebExtractProvider class
├── tests/
│ ├── __init__.py
│ └── test_browserless_provider.py # Unit tests
How the plugin integrates with Hermes
- Entry point:
pyproject.tomldeclares[project.entry-points."hermes_agent.plugins"]pointing to the package. Hermes discovers plugins via setuptools entry points. - Registration:
__init__.pycallsregister(ctx)which callsctx.register_web_search_provider(...)with an instance ofBrowserlessWebExtractProvider. - Provider class:
provider.pysubclassesWebSearchProviderfromagent.web_search_providerand implements the required interface:name,display_name,is_available(),supports_extract(),extract(), andget_setup_schema(). - Dispatching: When Hermes needs to call
web_extract, it looks up the active extract backend, finds this provider, and callsextract(urls).
License
MIT — see the LICENSE file for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hermes_browserless_web_extract-1.0.0.tar.gz.
File metadata
- Download URL: hermes_browserless_web_extract-1.0.0.tar.gz
- Upload date:
- Size: 20.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4a4d51cfce2173fff6861e25b6aa6097c4cbc9276dd1c900d4732318e1fa14d9
|
|
| MD5 |
f0ae9ef107992f54765fb5a69b78ae51
|
|
| BLAKE2b-256 |
876f7350626eac185d692761652aa2518008d6ac351161ce8394e7d866cef6d8
|
Provenance
The following attestation bundles were made for hermes_browserless_web_extract-1.0.0.tar.gz:
Publisher:
publish.yml on donbowman/hermes-browserless-web-extract
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hermes_browserless_web_extract-1.0.0.tar.gz -
Subject digest:
4a4d51cfce2173fff6861e25b6aa6097c4cbc9276dd1c900d4732318e1fa14d9 - Sigstore transparency entry: 2204736568
- Sigstore integration time:
-
Permalink:
donbowman/hermes-browserless-web-extract@94da54ad4fc2706bb2b0ae7d23779e7746113c86 -
Branch / Tag:
refs/tags/v0.0.1 - Owner: https://github.com/donbowman
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@94da54ad4fc2706bb2b0ae7d23779e7746113c86 -
Trigger Event:
push
-
Statement type:
File details
Details for the file hermes_browserless_web_extract-1.0.0-py3-none-any.whl.
File metadata
- Download URL: hermes_browserless_web_extract-1.0.0-py3-none-any.whl
- Upload date:
- Size: 13.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c8325fee483c6dfd3f3bffb8ebf76c6d695beea5cbb0ca48ca7db5bc273d190c
|
|
| MD5 |
93a08c9a44b1a57f25dd9466c418217e
|
|
| BLAKE2b-256 |
fba94f17821490e0245dbfc1bfdfdd23fbed9ff281cdb223fbc94fb227ad2bb3
|
Provenance
The following attestation bundles were made for hermes_browserless_web_extract-1.0.0-py3-none-any.whl:
Publisher:
publish.yml on donbowman/hermes-browserless-web-extract
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hermes_browserless_web_extract-1.0.0-py3-none-any.whl -
Subject digest:
c8325fee483c6dfd3f3bffb8ebf76c6d695beea5cbb0ca48ca7db5bc273d190c - Sigstore transparency entry: 2204736588
- Sigstore integration time:
-
Permalink:
donbowman/hermes-browserless-web-extract@94da54ad4fc2706bb2b0ae7d23779e7746113c86 -
Branch / Tag:
refs/tags/v0.0.1 - Owner: https://github.com/donbowman
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@94da54ad4fc2706bb2b0ae7d23779e7746113c86 -
Trigger Event:
push
-
Statement type: