Skip to main content

Capture a website's visual blueprint (screenshots + design tokens) for AI coding agents.

Project description

WebForge (Python)

Capture any website's visual blueprint — screenshots + design tokens — for AI coding agents.

WebForge screenshots a URL across desktop, tablet, and mobile viewports, extracts its design tokens (an area-weighted colour palette, top fonts, and image assets), optionally crawls a whole domain, and packages everything into a clean ZIP blueprint that tools like Claude Code, Cursor, or Codex can use to rebuild sites with pixel fidelity.

This is the Python port of the WebForge Chrome extension's capture engine, powered by Playwright.


Installation

pip install webforge-theatom

The package installs as webforge-theatom but is imported as webforge:

import webforge

WebForge drives a real Chromium browser, so after installing the package you need Playwright's browser binary once:

playwright install chromium

Installing from a local checkout during development:

cd python
pip install -e .
playwright install chromium

Quick start

import webforge

# Capture a single page across desktop, tablet and mobile.
result = webforge.capture("example.com")

print(result.title)     # "Example Domain"
print(result.colors)    # ['#ffffff', '#000000', ...]  (prominence-ordered)
print(result.fonts)     # ['Inter', 'Outfit', ...]
print(result.images)    # image URLs + inline data-URIs

# Save a ZIP blueprint (screenshots + metadata.json).
result.to_zip("example.zip")

# ...or write the raw files to a folder.
result.save("example/")

Crawl a whole domain:

site = webforge.crawl("example.com", max_pages=10)
print(len(site.pages), "pages captured")
site.to_zip("example-site.zip")

Command line

webforge example.com                    # capture -> WebForge_example.com.zip
webforge example.com -o out.zip         # choose the output path
webforge example.com --desktop          # desktop viewport only
webforge example.com --png              # PNG instead of JPEG
webforge example.com --crawl --max 10   # crawl up to 10 pages
webforge --version

API reference

webforge.capture(url, *, mode="all", viewports=None, extract_hover=True, image_format="jpeg", quality=62, timeout=15.0, executable_path=None, headless=True) -> CaptureResult

Capture a single page.

Argument Default Description
url Target URL; a missing scheme becomes https://.
mode "all" "all" = desktop + tablet + mobile; "desktop" = 1440 only.
viewports None Explicit iterable of Viewport (overrides mode).
extract_hover True Prime hover to surface sprite/hover-only artwork.
image_format "jpeg" "jpeg" (smaller) or "png".
quality 62 JPEG quality 0–100 (ignored for PNG).
timeout 15.0 Navigation timeout (seconds); on timeout it captures current state.
executable_path None Custom Chromium binary; defaults to Playwright's.
headless True Run the browser headless.

webforge.crawl(url, *, max_pages=5, custom_urls=None, mode="desktop", ...) -> CrawlResult

Breadth-first crawl of same-origin links, capturing each page (up to max_pages). Accepts the same capture options.

Result objects

  • CaptureResulturl, title, captured_at, screenshots ({viewport: bytes}), colors, fonts, images. Methods: .to_zip(path), .save(dir), .metadata().
  • CrawlResultdomain, pages (list[CapturedPage]), sitemap. Methods: .to_zip(path), .design_tokens(), .metadata().
  • Viewportkey, width, height. The defaults are in webforge.DEFAULT_VIEWPORTS.

Errors

All inherit from webforge.WebForgeError:

  • InvalidURLError — empty or non-http(s) URL.
  • BrowserNotInstalledError — Playwright/Chromium missing (run playwright install chromium).
  • CaptureError — navigation, render, or screenshot failed.

ZIP blueprint layout

Single page

WebForge_<domain>/
├── desktop.jpg
├── tablet.jpg
├── mobile.jpg
└── metadata.json

Crawl (multi-page)

WebForge_<domain>/
├── sitemap.json
├── metadata.json          # domain-wide design tokens
└── pages/
    ├── home/
    │   ├── desktop.jpg
    │   └── metadata.json
    └── <slug>/
        └── ...

This mirrors the layout produced by the WebForge browser extension.


Configuration examples

# Fast desktop-only PNG capture, no hover pass, short timeout.
webforge.capture(
    "example.com",
    mode="desktop",
    image_format="png",
    extract_hover=False,
    timeout=8.0,
)

# Custom viewport set.
from webforge import Viewport
webforge.capture("example.com", viewports=[Viewport("wide", 1920, 1080)])

# Point at a system Chrome instead of Playwright's bundled Chromium.
webforge.capture("example.com", executable_path="/Applications/Google Chrome.app/Contents/MacOS/Google Chrome")

Troubleshooting

Symptom Fix
BrowserNotInstalledError / "Executable doesn't exist" Run playwright install chromium.
Capture times out on heavy pages Increase timeout=; WebForge still returns the current state rather than failing.
A site blocks automated visits (Cloudflare, bot walls) WebForge already spoofs a real user agent and hides the WebDriver flag, but some sites can't be captured headless — try headless=False, or use the browser extension which runs in your own session.
Missing colours/fonts Some tokens live in cross-origin stylesheets the browser won't expose; this matches the extension's behaviour.
Empty/black screenshots The page likely needs longer to render — raise timeout=.

Development

cd python
pip install -e ".[dev]"
playwright install chromium

pytest -m "not integration"   # fast unit tests, no browser
pytest -m integration         # end-to-end (needs Chromium)

License

MIT © Zaki Sheriff

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webforge_theatom-0.1.1.tar.gz (19.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

webforge_theatom-0.1.1-py3-none-any.whl (20.1 kB view details)

Uploaded Python 3

File details

Details for the file webforge_theatom-0.1.1.tar.gz.

File metadata

  • Download URL: webforge_theatom-0.1.1.tar.gz
  • Upload date:
  • Size: 19.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.5

File hashes

Hashes for webforge_theatom-0.1.1.tar.gz
Algorithm Hash digest
SHA256 b13964db45e2977725fc5e84e43a1aa55551b5fdb9fcc9200793c986eeece63d
MD5 7550fe6f2af02d5c842e6ba728073617
BLAKE2b-256 efa4cad708fad0ede60429ad13a91014c763332975880ff1675fac6a6749f8ac

See more details on using hashes here.

File details

Details for the file webforge_theatom-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for webforge_theatom-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 87adcf57b1d49b2dd9587a0efb7c0da3dc9bd72ae93148da91d9dbf5583a0c3d
MD5 ae49e25fcc7e096d49daa59926c7e829
BLAKE2b-256 a6422d849827380e7f00cf643b4e7cdbbafd0833a60d93ef53b729bd33b080c9

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page