Skip to main content

wintergrab

Friendly web scraping that scales from one page to big crawls.

import wintergrab as wg

page = wg.get("https://quotes.toscrape.com/")
for quote in page.css(".quote"):
    print(quote.css(".text::text").get(), "-", quote.css(".author::text").get())

wintergrab is a Python toolkit for grabbing data from websites. Simple things are one line. When a site is harder (JavaScript, bot checks, rate limits, thousands of pages), the same API scales up.

  • Fetch like a real browser. HTTP requests carry Chrome/Firefox/Safari TLS and HTTP/2 fingerprints (via curl_cffi). A headless Chromium (via Playwright) is one flag away for JavaScript pages. It hides common automation tells and waits out "checking your browser" interstitials.
  • Parse with CSS or XPath. Scrapy-style ::text / ::attr(href), extraction schemas, search by text, "find similar elements", and HTML → Markdown/text conversion.
  • Selectors that adapt. With adaptive=True, wintergrab remembers what a selector matched. After a redesign that breaks it, it finds the most similar elements on the new page.
  • Crawl at scale. Async spiders with concurrency limits, multiple sessions (HTTP + browser, several accounts…), proxy rotation with health checks, AutoThrottle that backs off when a site pushes back, robots.txt support, and pause/resume (Ctrl+C, then run again).
  • Scrape without selectors. Pull JSON-LD/microdata/OpenGraph, the JSON state that React/Next/Vue apps embed in their HTML, and every table. auto_extract() finds a page's product grid or result list and names the fields. learn({"title": "…", "price": "…"}) writes the selectors for you from values you can see on the page.
  • Built for big, long crawls. An HTTP cache that revalidates with 304s and replays whole crawls offline. A disk-backed queue with a Bloom filter that keeps memory flat at millions of URLs and survives kill -9. Sitemap crawling, SQLite output with upserts, and a live progress line.
  • Fast. In a reproducible benchmark against a local test shop, a wintergrab spider crawled about 1,000 pages/s on one core. That is 1.8× Crawlee and 3.8× Scrapy at the same concurrency, with under half their memory. On real sites, the site and your politeness settings usually set the pace, not the crawler.
  • Browser superpowers. Capture the JSON API calls a page makes while it renders. Clear a login or JS check once in the browser, then continue over fast HTTP with the same cookies.
  • A small CLI. wintergrab get and wintergrab crawl cover the common jobs with no code at all, including --auto, --learn and --offline.

Install

pip install wintergrab                 # HTTP fetching, parsing, spiders, CLI
pip install "wintergrab[browser]"      # + headless browser support
pip install "wintergrab[speed]"        # + uvloop and orjson
playwright install chromium            # one-time browser download (browser extra only)
wintergrab doctor                      # check what is installed

Python 3.10+ on Linux, macOS and Windows. On a fresh Linux machine, use playwright install --with-deps chromium to get the browser's system libraries too. The development version installs straight from GitHub: pip install "wintergrab @ git+https://github.com/opensourcewinter/wintergrab".

A quick tour

Fetch and parse

import wintergrab as wg

page = wg.get("https://books.toscrape.com/")      # looks like Chrome, retries hiccups
page.status, page.title                           # (200, 'All products | Books to Scrape')

page.css("h3 a::attr(title)").getall()            # every title
page.css(".price_color::text").get()              # first price: '£51.77'
page.xpath("//p[contains(@class, 'star-rating')]/@class").get()

for book in page.css("article.product_pod"):      # loop and query inside
    print(book.css("h3 a").attr("title"), book.css(".price_color").text)

page.links(".pager")                              # absolute URLs of links in the pager
page.find_by_text("Tipping the Velvet")           # search by visible text
page.markdown(main_content=True)                  # the page as Markdown

Pull out structured data

from wintergrab import Field

books = page.extract_all("article.product_pod", {
    "title": "h3 a::attr(title)",
    "price": Field(".price_color::text", transform=lambda p: float(p.lstrip("£"))),
    "rating": Field("p.star-rating", attr="class", regex=r"star-rating (\w+)"),
})
page.extract({"titles": ["h3 a::attr(title)"]})       # a one-item list = all matches

Survive layout changes

products = page.css(".product-card", adaptive=True)

The first time, wintergrab saves a fingerprint of what matched: tag, attributes, text, position, parent and neighbours. If the site later renames .product-card or wraps it in new containers, the same call scores every element on the new page and returns the closest matches. It logs a warning so you know to update the selector. See docs/adaptive-selectors.md.

Scrape without writing selectors

page.auto_extract()          # [{"title", "url", "image", "price", "rating"...}, ...] from the main record list
schema = page.learn({"title": "A Light in the Attic", "price": "£51.77"})
schema.extract(other_page)   # the learned selectors work on every page of that template
page.structured_data()       # JSON-LD, microdata, OpenGraph, meta tags
page.embedded_json()         # __NEXT_DATA__, window.__INITIAL_STATE__, ... (SPAs without a browser)
page.tables()                # every table as records
page.next_page()             # pagination, auto-detected

JavaScript pages

page = wg.render("https://quotes.toscrape.com/js/", wait_for=".quote")

with wg.BrowserFetcher(headless=True) as browser:     # reuse one browser
    page = browser.get(url, scroll=True, screenshot="page.png")

Many pages at once

async with wg.AsyncFetcher() as fetcher:
    pages = await fetcher.get_many(urls, concurrency=10)

Crawl a site

from wintergrab import Spider

class BooksSpider(Spider):
    start_urls = ["https://books.toscrape.com/"]
    allowed_domains = ["books.toscrape.com"]
    concurrency = 16                 # AutoThrottle adapts the real speed per domain
    crawl_dir = ".crawl/books"       # makes it resumable: Ctrl+C pauses, re-run resumes
    output = "books.jsonl"           # items stream here (.jsonl / .json / .csv)

    def parse(self, response):
        for link in response.css("article.product_pod h3 a"):
            yield response.follow(link, callback=self.parse_book)
        yield from response.follow_all("li.next a")

    def parse_book(self, response):
        yield {
            "title": response.css("h1::text").get(),
            "price": response.css(".product_main .price_color::text").get(),
        }

result = BooksSpider().run()
print(result.status, result.stats["pages"], result.stats["items"])

Scaling up is a few attributes away:

class BigCrawl(Spider):
    sitemap_urls = ["https://shop.example/robots.txt"]   # discover pages from sitemaps
    frontier = "disk"            # flat memory for millions of URLs, crash-safe queue
    crawl_dir = ".crawl/big"
    cache = ".cache/big"         # revalidating HTTP cache; cache_mode="offline" replays the crawl
    output = "catalog.db"        # SQLite...
    unique_key = "url"           # ...with upserts: re-crawls update rows in place
    fallback_session = "browser" # blocked page? retry it in a headless browser, share its cookies

Spiders also give you:

  • Sessions. Route requests through different fetchers with Request(url, session="browser"). Set fallback_session="browser" to retry blocked pages in a headless browser automatically.
  • Proxy rotation. proxies = [...] (or a ProxyRotator). Proxies that keep failing are benched for a while.
  • Speed control. Per-domain concurrency and delays that back off on 429/503/block pages, honour Retry-After and robots.txt Crawl-delay, and recover gradually.
  • Limits and hooks. max_pages, max_items, max_depth, process_item(), on_error(), on_start() / on_close(), and async for item in spider.stream().

Command line

wintergrab get https://quotes.toscrape.com                          # page as Markdown
wintergrab get https://quotes.toscrape.com --css ".quote .text::text"
wintergrab get https://books.toscrape.com --each article.product_pod \
    --field title="h3 a::attr(title)" --field price=.price_color::text -o books.csv
wintergrab get https://quotes.toscrape.com/js/ --browser --wait-for .quote

wintergrab get https://books.toscrape.com --auto                     # records, no selectors
wintergrab get https://books.toscrape.com --learn "title=A Light in the Attic" --save-schema books.json
wintergrab crawl https://books.toscrape.com --schema books.json --paginate -o books.jsonl
wintergrab get https://shop.example/p/1 --structured                # JSON-LD, OpenGraph...

wintergrab crawl my_spider.py -o items.jsonl --crawl-dir .crawl/mine   # run a spider file
wintergrab crawl https://books.toscrape.com --follow "li.next a" --follow "h3 a" \
    --each ".product_main" --field title=h1::text --max-pages 50 -o books.jsonl

wintergrab shell https://quotes.toscrape.com                        # explore interactively

Documentation

Guide What's inside
Getting started Install, first scrape, first spider, in 10 minutes
Fetching get/Fetcher/AsyncFetcher/BrowserFetcher, options, errors
Parsing Selectors, extraction schemas, text search, Markdown
Adaptive selectors How relocation works and how to tune it
Spiders Crawling, sessions, pause/resume, output, every setting
Power features Zero-selector extraction, cache & offline replay, API capture, cookie handoff, sitemaps, disk frontier, SQLite
Tough sites Impersonation, browsers, proxies, AutoThrottle, etiquette
CLI get, crawl and shell reference
Examples Runnable scripts for every feature

Scrape responsibly

wintergrab makes polite crawling the default. Spiders obey robots.txt, adapt their speed to each site, and back off when asked. Stealth features exist so legitimate automation isn't misclassified. They don't make it OK to ignore a site's terms, hammer servers, or collect personal data you have no right to. Check the rules of each site you scrape.

Development

python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"
playwright install chromium        # for the browser tests (skipped otherwise)
pytest                             # runs against a local test site; no internet needed
ruff check . && ruff format --check .

See CONTRIBUTING.md for the live tests and the release process, and SECURITY.md to report a vulnerability.

License

MIT

Release files for wintergrab 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wintergrab 0.2.0
File Size Uploaded
wintergrab-0.2.0.tar.gz 228.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for wintergrab 0.2.0
File Interpreter ABI Platform
wintergrab-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 382.7 kB

Release files / wintergrab-0.2.0.tar.gz

Download URL wintergrab-0.2.0.tar.gz
Size 228.2 kB
Tags Source
SHA-256 checksum
How to use checksums
9511a5ab8a0031141ea25d8f6314fb7b95d028dca93989d2c36b76547ca31552
BLAKE2b-256 checksum
How to use checksums
2c9f821b4c22d95252103fe6aca99b8d2d1ec24764402e8b1c9aee3ddecb2573
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log

Release files / wintergrab-0.2.0-py3-none-any.whl

Download URL wintergrab-0.2.0-py3-none-any.whl
Size 154.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
35b50036de0e270e4726d71a1826fca6cc5fc6b6cec1203aa4c60f81ff6f19a5
BLAKE2b-256 checksum
How to use checksums
cf1b53a501bd24233f8bb21ed3a9417e1615d0c07ce6acb924d9e6f9f4b537e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page