Skip to main content

scrapewright

PyPI Python License: MIT

Give it a URL. It writes the scraper.

Most e-commerce catalog scraping splits into two worlds: sites on a known platform (Shopify, WooCommerce) that expose a clean JSON feed, and everything else — bespoke HTML where you hand-write a parser per site and re-write it every time the markup shifts. scrapewright collapses both into one call:

  1. Detect the platform behind a URL.
  2. For known platforms, extract deterministically from their public catalog API — free, stable, no LLM.
  3. For custom HTML, synthesize a reusable extractor once with an LLM, cache it, and replay it deterministically forever after.

The LLM is a compiler, not a runtime. It runs once per site to produce a recipe of CSS selectors; every page after that is parsed by plain BeautifulSoup at zero marginal cost. That is the whole cost-control story — no per-page model calls, no token bill that scales with your crawl.

                    ┌─────────────┐
   store URL  ───▶  │   detect    │
                    └──────┬──────┘
        ┌──────────────────┼──────────────────┐
        ▼                  ▼                   ▼
    shopify            woocommerce         generic HTML
   products.json      wc/store/products    (page mode)
        │                  │                   │
        │  deterministic   │                   ▼
        │  (free)          │            cached recipe? ──yes──▶ replay (free)
        └────────┬─────────┘                   │ no
                 ▼                              ▼
             Product{}  ◀───── selectors ── JSON-LD? ──yes──▶ Product{} (free)
                 ▲                              │ no
                 │                              ▼
                 └──────── replay ◀── LLM synthesizes recipe ONCE ──▶ cache

Everything normalizes to one Product shape, so downstream code never knows or cares which path a record came from.

Install

pip install scrapewright               # deterministic paths (Shopify, Woo, JSON-LD)
pip install "scrapewright[llm]"        # + LLM recipe synthesis for custom HTML
pip install "scrapewright[llm,js,excel,mcp]"   # + JS rendering, XLSX, MCP server
playwright install chromium                    # only needed for --js

Use it

from scrapewright import Scrapewright

sw = Scrapewright()

# Catalog mode — a whole Shopify/WooCommerce store, deterministically
for product in sw.scrape_catalog("https://shop.example.com", max_items=200):
    print(product.brand, product.title, product.price, product.currency)

# Page mode — one custom-HTML product page.
# First call: tries JSON-LD (free); if absent, the LLM writes a recipe once.
# Every later call on that domain: replayed from the cached recipe, no LLM.
item = sw.scrape_page("https://boutique.example.com/products/wool-coat")
print(item.model_dump(exclude={"raw"}))

# Crawl mode — walk a WHOLE custom store from one listing/category URL.
# The frontier discovers product pages (deterministic, free); the first page
# pays the single synthesis cost, every other page replays the recipe.
for product in sw.crawl("https://boutique.example.com/collection", max_items=100):
    print(product.title, product.price)

CLI

scrapewright detect https://shop.example.com          # platform + strategy
scrapewright run    https://shop.example.com --max 50 # scrape a catalog → JSONL
scrapewright crawl  https://boutique.example.com/collection -o products.xlsx
scrapewright run    https://shop.example.com -o products.csv   # Excel-ready CSV
scrapewright add    https://boutique.example.com/products/coat  # learn a site
scrapewright run    https://boutique.example.com/products/coat --no-llm
scrapewright list                                     # cached recipe domains

-o writes .csv (Excel-ready, UTF-8 BOM), .xlsx (pip install scrapewright[excel]), or .jsonl; without it, products stream to stdout as JSONL.

Know what you are dealing with

detect answers the routing question before a job starts:

$ scrapewright detect https://some-store.com
https://some-store.com
  platform: bigcommerce
  catalog:  -
  strategy: crawl
  note:     BigCommerce (Stencil) markup

Twelve platforms are recognized: Shopify and WooCommerce publish a free JSON catalog, so those route to catalog — deterministic, no LLM, no browser. Magento, BigCommerce, Salesforce Commerce Cloud, Squarespace, Wix, Webflow, PrestaShop, Shopware, Ecwid and OpenCart are recognized by fingerprint and route to crawl, where the recipe path handles them like any custom site — the point of naming them is knowing what you face, not writing twelve parsers. Wix and Ecwid render client-side, so detection says crawl+js up front.

A site behind an anti-bot wall reports strategy: blocked with the HTTP status, rather than pretending it found nothing.

Bring your own schema

Products are just the built-in default. Declare the fields you want and the same compile-once/replay-free loop works on any structured page — job posts, listings, registry records:

scrapewright run https://jobs.example.com/p/123 -f title -f company -f salary:number -f tags:list --schema-name job
from scrapewright import Scrapewright, Schema

job = Schema.from_names(["title", "company", "salary:number", "tags:list"], name="job")
record = Scrapewright().extract("https://jobs.example.com/p/123", job)
print(record.data)   # {'title': ..., 'company': ..., 'salary': ..., 'tags': [...]}

Field kinds are text (default), number, url, and list. Recipes are cached per site and per schema, so one domain can be compiled against several field sets without them overwriting each other.

Use it from an AI agent (MCP)

scrapewright ships an MCP server, so an agent can call it as a tool instead of reading raw HTML itself:

pip install "scrapewright[mcp,llm]"
scrapewright mcp

Point any MCP client at that command and the agent gains five tools: detect_site, scrape_catalog, extract_page, crawl_site, and list_learned_sites.

The economics are the point. An agent that reads pages itself pays model tokens per page, forever. These tools pay once per site — an agent crawling 500 pages spends one synthesis, not five hundred, and platform stores (Shopify, WooCommerce) cost nothing at all.

Run it as a service

The same core behind an HTTP API, with keys, quotas, metering and background jobs:

pip install "scrapewright[service,llm]"
scrapewright keys create --label alice --plan free
scrapewright serve --port 8000
curl -X POST localhost:8000/v1/extract   -H "X-API-Key: sw_..." -H "Content-Type: application/json"   -d '{"url": "https://shop.example.com/products/coat"}'
Endpoint Purpose
POST /v1/detect platform + strategy (cheap)
POST /v1/extract one page -> structured record
POST /v1/crawl a whole site -> job id (crawls outlive a request)
GET /v1/jobs/{id} poll a crawl
GET /v1/usage what this key has consumed, against its plan

Prepaid credits, no subscription

One action costs real money: compiling a new site, a single LLM pass over a page, measured at $0.02 on a small product page and $0.15 on a heavy rendered one. Everything after that is BeautifulSoup — the ten-thousandth record from a compiled site is free to serve. So credits are priced off that one action, and everything else is denominated relative to it:

Action Credits
1 record delivered 1
1 browser render 5
1 new site compiled 300
page fetches, detect free
$ scrapewright plans
pack         credits   price   $/credit   margin
starter       10,000     $10    0.00100    80.0%
growth        50,000     $40    0.00080    75.0%
scale        250,000    $150    0.00060    66.7%

Free: 1,000 credits a month, resetting.

Margin is measured on compiling a site, because that is the only step that costs anything; a test fails if a price edit drops any pack below 60%. A free account can cost us at most $0.20 a month, even if every free credit goes to the most expensive action there is.

Credits are a ledger, not a counter — every grant and every charge is a row, so a disputed bill can be reconstructed line by line, and a replayed payment webhook cannot double-credit (grants take an idempotency key). Running out returns 402 with the balance and what to do about it; a crawl is capped by the credits on hand, so a job stops at what the caller can pay for instead of overdrawing.

scrapewright credits grant <key_id> --pack starter --idempotency <payment_id>
scrapewright credits balance <key_id>

Taking payment

Stripe is wired in and turned on by environment, not by a code change:

pip install "scrapewright[service,stripe]"
export STRIPE_SECRET_KEY=sk_test_...      # absent -> nothing is for sale
export STRIPE_WEBHOOK_SECRET=whsec_...    # absent -> webhooks are refused
scrapewright serve
Endpoint Purpose
GET /v1/credits/packs the price list — public, no key needed
POST /v1/credits/checkout start a purchase, returns a Stripe Checkout URL
POST /v1/webhooks/stripe payment notifications from Stripe

The webhook endpoint takes no API key — Stripe is the caller, so the signature is the credential, and an unverified endpoint would be a free credit printer for anyone who guessed the URL. Three rules hold the integration up:

  • Verify every signature. No signing secret configured means webhooks are refused outright, rather than accepted unverified.
  • Never trust an amount off the wire. The event names a pack; how many credits that pack is worth is looked up from our own price list, so a tampered payload buys exactly what it paid for or nothing at all.
  • Grant idempotently, keyed on the Checkout session. Stripe retries deliveries, and one payment can produce several event types — the session id is what identifies the money that actually moved.

examples/stripe_smoke_test.py runs the whole path against Stripe's test mode with the 4242 card. Any other provider plugs into the same two-method BillingProvider protocol in scrapewright.service.billing; without one, the service simply runs free, which is the right default for a demo or a self-hosted instance.

Docker:

docker build -t scrapewright .                       # static paths
docker build -t scrapewright --build-arg WITH_JS=1 . # + headless Chromium
docker run -p 8000:8000 -v sw-data:/data scrapewright

Client-side-rendered stores

Add --js (or Scrapewright(js=True)) and pages that render their catalog in the browser become extractable:

scrapewright run https://spa-store.example.com/products/x --page --js
scrapewright crawl https://spa-store.example.com/shop --js -o products.xlsx

Rendering stays rare by construction: the static fetch runs first, and Chromium is only started when the static HTML is an empty client-side shell or extraction on it fails. A recipe learned from rendered HTML is tagged needs_js, so later runs on that site skip the wasted static hop. The browser starts at most once per run and is reused for every page.

The Product shape

url: str            # canonical product URL
title: str
brand: str | None
price: Decimal | None   # parsed from "1,250.00" / "1.250,00" / "€1290" alike
currency: str | None
available: bool | None
images: list[str]       # absolute URLs
sizes: list[str]
description: str | None
sku: str | None
source_platform: str    # shopify | woocommerce | json-ld | selector

A record is usable when it carries a title, a price, and a URL. The validator (scrapewright.coverage) reports the usable ratio across a batch — the number a recipe is trusted on before it's cached.

How the pieces fit

Module Role
detect Platform registry: free-catalog probes, then fingerprints for 12 platforms; returns the strategy to use
extract/shopify, extract/woocommerce Deterministic catalog extractors
extract/jsonld schema.org/Product from <script type="application/ld+json"> — free, ~common
extract/llm Synthesizes a SelectorRecipe from HTML — the one-time compile step
extract/selectors Replays a recipe with BeautifulSoup — the deterministic runtime
schema Schema/Field — declare what to extract; PRODUCT_SCHEMA is the built-in default
service/ FastAPI app: API keys (stored hashed), record-based quotas, cost metering, background crawl jobs, pluggable billing
service/credits Credit prices, packs, and the free allowance
service/stripe_billing Stripe Checkout + signature-verified webhook
service/pricing Measured unit costs and the margin each pack clears
mcp_server Five MCP tools so AI agents can call scrapewright directly
fetch StaticFetcher (plain HTTP) and BrowserFetcher (headless Chromium), plus the shell heuristic that decides when a render is worth paying for
crawl Frontier: turns one listing URL into product URLs (pattern match + card-template fallback + pagination) — deterministic, no LLM
cache Persists recipes keyed by domain, so the compile happens once
validate Field-coverage scoring
export Batch → .csv / .xlsx / .jsonl
pipeline Orchestrates detect → extract → validate → cache → heal

Design notes

  • Deterministic paths run first. Shopify JSON, the WooCommerce Store API, and JSON-LD cover a large share of real stores for free. The LLM is only ever reached for genuinely custom HTML.
  • Self-healing. When a cached recipe stops producing usable products — the site changed its DOM — the page falls through to the free JSON-LD path and, failing that, a fresh synthesis replaces the stale recipe. A broken site heals on the next run instead of silently returning empty fields.
  • Bounded model spend. Batch and crawl runs cap LLM calls at max_synth_per_run (default 3) — a site that resists synthesis cannot burn one model call per page. The bill is bounded no matter how large the crawl.
  • Provider-configurable. The LLM extractor takes a model and works with any injected client; the default targets Anthropic's Claude via the official SDK.

Testing

The deterministic paths are fully covered by offline fixtures — no network, no model calls — so CI is green without an API key:

pip install "scrapewright[dev]"
pytest

Status

v0.9 (alpha). Implemented and tested: an HTTP service with API keys, prepaid credits (priced off the one action that costs money, on an auditable ledger) and Stripe checkout with a signature-verified webhook, cost metering and background jobs; platform detection across 12 storefronts with a recommended strategy per site, catalog extraction (Shopify, WooCommerce), page extraction (JSON-LD, LLM-synthesized selectors), recipe caching, self-healing re-synthesis with a bounded per-run model budget, a crawl frontier (one listing URL → the whole site), JS rendering via an optional Playwright fetcher with automatic escalation, schema-agnostic extraction (bring your own fields), an MCP server for AI agents, coverage validation, and CSV / XLSX / JSONL export. 141 offline tests.

Known limit, stated plainly: it does not defeat anti-bot walls — deliberately out of scope. Sites behind Akamai/Fastly-style challenges return an honest miss.

Roadmap: pagination strategies for infinite-scroll listings, and a deployed instance of the service.

License

MIT — see LICENSE.

Metadata

Release files for scrapewright 0.9.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrapewright 0.9.1
File Size Uploaded
scrapewright-0.9.1.tar.gz 116.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrapewright 0.9.1
File Interpreter ABI Platform
scrapewright-0.9.1-py3-none-any.whl Python 3 none any Details

Total release size: 184.8 kB

Release files / scrapewright-0.9.1.tar.gz

Download URL scrapewright-0.9.1.tar.gz
Size 116.4 kB
Tags Source
SHA-256 checksum
How to use checksums
e535b21893286ff9199305673a76feafcfe226a0650b71f6efe74cb8cc6413e2
BLAKE2b-256 checksum
How to use checksums
1a17fec64b3f1e98185218e62584068cb606211dbb919c48729f941647f0fad0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / scrapewright-0.9.1-py3-none-any.whl

Download URL scrapewright-0.9.1-py3-none-any.whl
Size 68.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b8eb27c1a0aeb194ef13afba1e84ba6f144b9fdd14c469f3df677ac91b9a4202
BLAKE2b-256 checksum
How to use checksums
c3cb06dabbf84a41a023d3a181f942c11902003d388f875e7f52e108a030e4c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

This release

0.9.1 This release

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page