Skip to main content

plasmate-browser-use

SOM-based content extraction for Browser Use. Drop-in alternative to Browser Use's default DOM serializer that uses Plasmate's Semantic Object Model (SOM) to reduce token costs by 10x or more.

Instead of sending the full DOM tree to your LLM, Plasmate compresses web pages into a compact semantic representation. Same information, 90% fewer tokens, lower costs, faster responses.

Install

pip install plasmate-browser-use

Prerequisites

You need the plasmate binary installed:

# Via cargo
cargo install plasmate

# Or via install script
curl -fsSL https://plasmate.app/install.sh | sh

Verify it works:

plasmate --version

Quick Start

Basic extraction

from plasmate_browser_use import PlasmateExtractor

extractor = PlasmateExtractor()

# Get raw SOM data as a dict
som = extractor.extract("https://news.ycombinator.com")
print(f"Elements: {som['meta']['element_count']}")
print(f"Compression: {som['meta']['html_bytes'] / som['meta']['som_bytes']:.1f}x")

Get page context for an LLM

The get_page_context() method returns a formatted string optimized for LLM consumption, with interactive elements, links, content, and compression stats:

context = extractor.get_page_context("https://example.com")
print(context)

Output:

# Example Domain
URL: https://example.com
Language: en

## Interactive Elements (1)
  [e1] link "More information..." (click)

## Content
This domain is for use in illustrative examples in documents...

---
Compression: 15.2x (1256 HTML bytes -> 83 SOM bytes)
Elements: 5 (1 interactive)

Markdown extraction

md = extractor.extract_markdown("https://example.com")
print(md)

Async support

All methods have async variants:

import asyncio

async def main():
    extractor = PlasmateExtractor()
    context = await extractor.get_page_context_async("https://example.com")
    som = await extractor.extract_async("https://example.com")
    md = await extractor.extract_markdown_async("https://example.com")

asyncio.run(main())

Using with a Browser Use agent

from browser_use import Agent
from plasmate_browser_use import PlasmateExtractor

extractor = PlasmateExtractor()

# Get compact page context instead of full DOM
context = extractor.get_page_context("https://example.com/products")

# Feed to your Browser Use agent with 10x fewer tokens
agent = Agent(task="Find the cheapest product", page_context=context)
result = await agent.run()

Token savings comparison

from plasmate_browser_use import PlasmateExtractor, token_count_comparison

extractor = PlasmateExtractor()
som = extractor.extract("https://news.ycombinator.com")
stats = token_count_comparison(som)

print(f"HTML tokens: ~{stats['html_tokens_est']:,}")
print(f"SOM tokens:  ~{stats['som_tokens_est']:,}")
print(f"Savings:     {stats['token_savings_pct']}%")
print(f"Ratio:       {stats['token_ratio']}x fewer tokens")

Typical token savings

Site HTML tokens SOM tokens Reduction
Hacker News ~22,000 ~1,200 18x
Wikipedia article ~85,000 ~8,500 10x
Amazon product page ~120,000 ~6,000 20x
Google search results ~45,000 ~3,500 13x

Numbers vary by page. The more complex the page (ads, trackers, layout noise), the bigger the savings.

How it works

  1. Plasmate fetches the page and parses the HTML
  2. The DOM is compiled into a Semantic Object Model (SOM) that preserves meaning while stripping layout noise
  3. The SOM is serialized into a compact format with tagged interactive elements
  4. Your LLM agent sees the same page information in 10x fewer tokens

Links

License

Apache-2.0

Release files for plasmate-browser-use 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for plasmate-browser-use 0.5.0
File Size Uploaded
plasmate_browser_use-0.5.0.tar.gz 5.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for plasmate-browser-use 0.5.0
File Interpreter ABI Platform
plasmate_browser_use-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 11.8 kB

Release files / plasmate_browser_use-0.5.0.tar.gz

Download URL plasmate_browser_use-0.5.0.tar.gz
Size 5.4 kB
Tags Source
SHA-256 checksum
How to use checksums
10446c4f4969ffed94206e999856afa4c2af2ba9eac859f095e4793001448d4f
BLAKE2b-256 checksum
How to use checksums
69fbc46f9cc4014d2e35b1f8416f07b89b7294ebbff5ef639a48495f7acfb5ad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release files / plasmate_browser_use-0.5.0-py3-none-any.whl

Download URL plasmate_browser_use-0.5.0-py3-none-any.whl
Size 6.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
882edce686b2091b1e048ca9d0f35f7936a9a85133094f40978c985c1d300591
BLAKE2b-256 checksum
How to use checksums
5721f51cd7adeae5a5ef552700e5a826c0aa0f45c6523cbefb982e2ee6cd2c71
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page