Skip to main content

scrapewright

PyPI Python License: MIT

Give it a store URL. It writes the scraper.

Most e-commerce catalog scraping splits into two worlds: sites on a known platform (Shopify, WooCommerce) that expose a clean JSON feed, and everything else — bespoke HTML where you hand-write a parser per site and re-write it every time the markup shifts. scrapewright collapses both into one call:

  1. Detect the platform behind a URL.
  2. For known platforms, extract deterministically from their public catalog API — free, stable, no LLM.
  3. For custom HTML, synthesize a reusable extractor once with an LLM, cache it, and replay it deterministically forever after.

The LLM is a compiler, not a runtime. It runs once per site to produce a recipe of CSS selectors; every page after that is parsed by plain BeautifulSoup at zero marginal cost. That is the whole cost-control story — no per-page model calls, no token bill that scales with your crawl.

                    ┌─────────────┐
   store URL  ───▶  │   detect    │
                    └──────┬──────┘
        ┌──────────────────┼──────────────────┐
        ▼                  ▼                   ▼
    shopify            woocommerce         generic HTML
   products.json      wc/store/products    (page mode)
        │                  │                   │
        │  deterministic   │                   ▼
        │  (free)          │            cached recipe? ──yes──▶ replay (free)
        └────────┬─────────┘                   │ no
                 ▼                              ▼
             Product{}  ◀───── selectors ── JSON-LD? ──yes──▶ Product{} (free)
                 ▲                              │ no
                 │                              ▼
                 └──────── replay ◀── LLM synthesizes recipe ONCE ──▶ cache

Everything normalizes to one Product shape, so downstream code never knows or cares which path a record came from.

Install

pip install scrapewright               # deterministic paths (Shopify, Woo, JSON-LD)
pip install "scrapewright[llm]"        # + LLM recipe synthesis for custom HTML
pip install "scrapewright[llm,js,excel]" && playwright install chromium   # + JS rendering, XLSX

Use it

from scrapewright import Scrapewright

sw = Scrapewright()

# Catalog mode — a whole Shopify/WooCommerce store, deterministically
for product in sw.scrape_catalog("https://shop.example.com", max_items=200):
    print(product.brand, product.title, product.price, product.currency)

# Page mode — one custom-HTML product page.
# First call: tries JSON-LD (free); if absent, the LLM writes a recipe once.
# Every later call on that domain: replayed from the cached recipe, no LLM.
item = sw.scrape_page("https://boutique.example.com/products/wool-coat")
print(item.model_dump(exclude={"raw"}))

# Crawl mode — walk a WHOLE custom store from one listing/category URL.
# The frontier discovers product pages (deterministic, free); the first page
# pays the single synthesis cost, every other page replays the recipe.
for product in sw.crawl("https://boutique.example.com/collection", max_items=100):
    print(product.title, product.price)

CLI

scrapewright detect https://shop.example.com          # what platform is this?
scrapewright run    https://shop.example.com --max 50 # scrape a catalog → JSONL
scrapewright crawl  https://boutique.example.com/collection -o products.xlsx
scrapewright run    https://shop.example.com -o products.csv   # Excel-ready CSV
scrapewright add    https://boutique.example.com/products/coat  # learn a site
scrapewright run    https://boutique.example.com/products/coat --no-llm
scrapewright list                                     # cached recipe domains

-o writes .csv (Excel-ready, UTF-8 BOM), .xlsx (pip install scrapewright[excel]), or .jsonl; without it, products stream to stdout as JSONL.

Client-side-rendered stores

Add --js (or Scrapewright(js=True)) and pages that render their catalog in the browser become extractable:

scrapewright run https://spa-store.example.com/products/x --page --js
scrapewright crawl https://spa-store.example.com/shop --js -o products.xlsx

Rendering stays rare by construction: the static fetch runs first, and Chromium is only started when the static HTML is an empty client-side shell or extraction on it fails. A recipe learned from rendered HTML is tagged needs_js, so later runs on that site skip the wasted static hop. The browser starts at most once per run and is reused for every page.

The Product shape

url: str            # canonical product URL
title: str
brand: str | None
price: Decimal | None   # parsed from "1,250.00" / "1.250,00" / "€1290" alike
currency: str | None
available: bool | None
images: list[str]       # absolute URLs
sizes: list[str]
description: str | None
sku: str | None
source_platform: str    # shopify | woocommerce | json-ld | selector

A record is usable when it carries a title, a price, and a URL. The validator (scrapewright.coverage) reports the usable ratio across a batch — the number a recipe is trusted on before it's cached.

How the pieces fit

Module Role
detect Platform probe: Shopify → WooCommerce → generic
extract/shopify, extract/woocommerce Deterministic catalog extractors
extract/jsonld schema.org/Product from <script type="application/ld+json"> — free, ~common
extract/llm Synthesizes a SelectorRecipe from HTML — the one-time compile step
extract/selectors Replays a recipe with BeautifulSoup — the deterministic runtime
fetch StaticFetcher (plain HTTP) and BrowserFetcher (headless Chromium), plus the shell heuristic that decides when a render is worth paying for
crawl Frontier: turns one listing URL into product URLs (pattern match + card-template fallback + pagination) — deterministic, no LLM
cache Persists recipes keyed by domain, so the compile happens once
validate Field-coverage scoring
export Batch → .csv / .xlsx / .jsonl
pipeline Orchestrates detect → extract → validate → cache → heal

Design notes

  • Deterministic paths run first. Shopify JSON, the WooCommerce Store API, and JSON-LD cover a large share of real stores for free. The LLM is only ever reached for genuinely custom HTML.
  • Self-healing. When a cached recipe stops producing usable products — the site changed its DOM — the page falls through to the free JSON-LD path and, failing that, a fresh synthesis replaces the stale recipe. A broken site heals on the next run instead of silently returning empty fields.
  • Bounded model spend. Batch and crawl runs cap LLM calls at max_synth_per_run (default 3) — a site that resists synthesis cannot burn one model call per page. The bill is bounded no matter how large the crawl.
  • Provider-configurable. The LLM extractor takes a model and works with any injected client; the default targets Anthropic's Claude via the official SDK.

Testing

The deterministic paths are fully covered by offline fixtures — no network, no model calls — so CI is green without an API key:

pip install "scrapewright[dev]"
pytest

Status

v0.3 (alpha). Implemented and tested: catalog extraction (Shopify, WooCommerce), page extraction (JSON-LD, LLM-synthesized selectors), recipe caching, self-healing re-synthesis with a bounded per-run model budget, a crawl frontier (one listing URL → the whole store), JS rendering via an optional Playwright fetcher with automatic escalation, coverage validation, and CSV / XLSX / JSONL export. 42 offline tests.

Known limits, stated plainly: it does not defeat anti-bot walls (deliberately out of scope), and the schema is products-only for now.

Roadmap: schema-agnostic extraction (bring your own field schema — the same compile-once/replay-free loop for any structured site, not just product pages), and BigCommerce / Salesforce Commerce detectors.

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapewright-0.3.0.tar.gz (30.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapewright-0.3.0-py3-none-any.whl (31.6 kB view details)

Uploaded Python 3

File details

Details for the file scrapewright-0.3.0.tar.gz.

File metadata

  • Download URL: scrapewright-0.3.0.tar.gz
  • Upload date:
  • Size: 30.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scrapewright-0.3.0.tar.gz
Algorithm Hash digest
SHA256 256b5494eeeb32b2c61183ffaeccd843e3e686d37aae6e4eeb1da545b2061415
MD5 7b4135e4af63a7a7ede081c6ee6537a7
BLAKE2b-256 2f8ffa851d235ce197a89e428895cdf8aac8ddf6bb55f24aea1e22265a63c0b6

See more details on using hashes here.

File details

Details for the file scrapewright-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: scrapewright-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 31.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scrapewright-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 42b5f84119ff396060c0186e99f9af8632401242e258f3845cbdb61aa484df24
MD5 c29c44e756af4d880a5c3dc14a81a406
BLAKE2b-256 8b1436a0a4ea09426243ff56ca6c21747873a858f866caa942f5326398029882

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.0

2 files

This release

0.3.0 This release

2 files

0.2.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page