Skip to main content

scrapy-stealth logo

scrapy-stealth

Stealthy Crawling. Maximum Results.

A pluggable anti-bot and stealth framework for Scrapy.

PyPI version Python versions Downloads GitHub release License: MIT Changelog

scrapy-stealth extends Scrapy with browser impersonation, proxy rotation, fingerprint cycling, and intelligent retry strategies โ€” designed for large-scale, production-grade crawling.


๐Ÿ’œ Sponsor

NodeMaven NodeMaven โ€” the most reliable proxy provider with the highest-quality IP on the market. Best solution for automation, web scraping, SEO research, and social media management.

Why NodeMaven?
  • 99.9% uptime
  • Sticky sessions up to 7 days
  • IP filtering: all proxies have fraud score <97%
  • No KYC required
  • Cashback on traffic โ€” burn GB and earn up to 10% back
Special codes for scrapy-stealth users: SCRAPYSTEALTH35 โ€” 35% off Mobile and Residential Proxies; SCRAPYSTEALTH40 โ€” 40% off ISP (Static) Proxies.

๐Ÿง  Why scrapy-stealth?

Scrapy is fast and powerful, but modern websites use advanced anti-bot protections such as:

  • TLS fingerprinting
  • Browser behavior detection
  • Rate limiting and IP blocking

scrapy-stealth helps by adding:

  • ๐Ÿงฌ Browser-level impersonation (TLS + HTTP/2 fingerprints)
  • ๐Ÿ” Smarter retry strategies
  • ๐ŸŒ Proxy and fingerprint rotation
  • ๐Ÿ›ก๏ธ Anti-bot detection

Result

  • Higher success rate
  • Lower proxy cost
  • More stable crawls

๐Ÿ“Š Comparison

Feature scrapy-stealth scrapy-impersonate scrapy-playwright scrapy-splash Scrapy (default)
TLS fingerprint spoofing โœ… โœ… โŒ โŒ โŒ
HTTP/2 support โœ… โœ… โœ… โŒ โŒ
Browser impersonation โœ… โœ… โš ๏ธ partial โŒ โŒ
Proxy rotation (built-in) โœ… โŒ โŒ โŒ โŒ
Fingerprint rotation โœ… โŒ โŒ โŒ โŒ
Anti-bot detection โœ… โŒ โŒ โŒ โŒ
Smart browser selection โœ… โŒ โŒ โŒ โŒ
Smart retry logic โœ… โŒ โŒ โŒ โŒ
Per-request engine switching โœ… โŒ โŒ โŒ โŒ
Headless browser required โœ… โŒ โœ… โœ… โŒ
JavaScript rendering ๏ธโœ… โŒ โœ… โœ… โŒ
Screenshot / snapshot โœ… โŒ โœ… โœ… โŒ
Native Scrapy integration โœ… โœ… โœ… โœ… โœ…
Memory footprint ๐ŸŸข Low ๐ŸŸข Low ๐Ÿ”ด High ๐Ÿ”ด High ๐ŸŸข Low

โš ๏ธ scrapy-playwright passes real browser TLS but does not spoof fingerprint profiles like scrapy-stealth does. scrapy-impersonate provides TLS/HTTP2 impersonation via curl_cffi but lacks built-in rotation, detection, or per-request engine switching. JavaScript rendering is available via the optional browser driver โ€” use it selectively for pages that require a full browser.


โœจ Features

  • ๐Ÿ”Œ Pluggable engine system (scrapy, stealth)
  • ๐Ÿง  Per-request engine selection via request.meta
  • ๐ŸŒ Proxy support and rotation
  • ๐Ÿงฌ Browser fingerprint rotation
  • ๐Ÿ” Smart retry logic
  • ๐Ÿ›ก๏ธ Anti-bot detection (status + content-based, Cloudflare, Akamai)
  • ๐Ÿง  Smart browser selection โ€” start with fast basic / turbo, auto-retry once with visible Chrome when a JS challenge or ban is detected
  • โšก Thread-safe async integration
  • ๐Ÿ–ฅ๏ธ Real-browser engine (CDP) for JS-heavy pages
  • ๐Ÿ”„ Intelligent session recycle โ€” after consecutive bans, browser restarts Chrome; basic/turbo clear HTTP sessions
  • ๐Ÿšซ Static asset blocking โ€” skip images, fonts, CSS, and media for faster, lighter browser fetches
  • ๐ŸŽฏ Proxy bypass list โ€” send chosen domains straight to the origin instead of through the proxy (--proxy-bypass-list)
  • ๐Ÿงญ Custom DNS overrides โ€” pin hosts to fixed IPs (connect via IP, keep hostname for TLS/SNI/Host) to dodge poisoned or geo-shifted public DNS
  • ๐Ÿ“ธ Built-in snapshot decorator (scrapy_stealth.decorators.snapshot)

๐Ÿ“ฆ Installation

pip install scrapy-stealth

Requires Python 3.11+ and Scrapy 2.12โ€“2.x


โš™๏ธ Setup

Option 1 โ€” Global (settings.py)

# 1. Enable the middleware
DOWNLOADER_MIDDLEWARES = {
    "scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
}

# 2. (Optional) Route ALL requests through stealth automatically โ€” no meta needed per request
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo"  # "basic" (default), "turbo", "browser", or "auto"
STEALTH_AUTO_FALLBACK = True  # retry basic/turbo JS challenges once with browser (headless=False)

# 3. (Optional) Proxy list โ€” seeded as engine default; rotated on ban-streak session recycle
#    Supported schemes: http, https, socks4, socks5
STEALTH_PROXIES = [
    "http://proxy1:8080",
    "http://proxy2:8080",
    "http://user:pass@proxy3:8080",  # with authentication
    "socks5://proxy4:1080",
]

# 4. (Optional) Pin hosts to fixed origin IPs (bypass public DNS)
#    Connects to the IP while keeping the hostname for TLS SNI / Host / certs
STEALTH_DNS_OVERRIDES = {
    "example.com": "203.0.113.10",
    "www.example.com": "203.0.113.10",
}

Option 2 โ€” Per-spider (custom_settings)

Configure the middleware and all stealth settings directly on the spider โ€” no changes to settings.py required.

class MySpider(scrapy.Spider):
    name = "example"

    custom_settings = {
        "DOWNLOADER_MIDDLEWARES": {
            "scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
        },
        "STEALTH_ENABLED": True,
        "STEALTH_DRIVER": "turbo",
        "STEALTH_AUTO_FALLBACK": True,
        "STEALTH_PROXIES": [
            "http://proxy1:8080",
            "http://user:pass@proxy2:8080",
            "socks5://proxy3:1080",
        ],
        "STEALTH_DNS_OVERRIDES": {
            "example.com": "203.0.113.10",
        },
    }

Proxies are validated at startup โ€” invalid format or unsupported scheme raises ValueError immediately. DNS overrides are validated the same way โ€” invalid IPs raise ValueError immediately.


๐Ÿš€ Quick Start

Option A โ€” Per-request (stealth only on specific requests):

yield scrapy.Request(
    url="https://example.com",
    meta={"stealth": {}},
)

Option B โ€” Global mode (stealth on every request automatically):

# settings.py or custom_settings
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo"
STEALTH_AUTO_FALLBACK = True  # optional: basic/turbo -> browser on JS challenge
# No meta needed โ€” all requests go through stealth
yield scrapy.Request(url="https://example.com")

# Opt out for a specific request
yield scrapy.Request(url="https://api.internal/health", meta={"stealth": False})

Option C โ€” Smart browser selection (fast HTTP first, Chrome only when needed):

# settings.py or custom_settings โ€” enable fallback for all stealth requests
STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo"  # primary: fast HTTP driver
STEALTH_AUTO_FALLBACK = True  # retry once with browser (headless=False) on JS challenge / ban
# Or per-request โ€” same behaviour without a global setting
yield scrapy.Request(
    url="https://example.com",
    meta={"stealth": {"driver": "auto"}},
)

See Smart browser selection for full details.


๐Ÿ”ง Global Configuration

Customise package-wide defaults via the shared config instance. All settings must be applied at module level, before the spider class โ€” the engine client is created at middleware initialisation, so changes inside start_requests or parse will have no effect.

# myspider.py
import scrapy
from scrapy_stealth.config import config

config.DEFAULT_ENGINE = "stealth"  # "scrapy" (native) or "stealth" (browser impersonation)
config.DEFAULT_PROFILE = "chrome_147"  # browser profile when meta["stealth"]["profile"] is not set
config.DEFAULT_TIMEOUT = 30  # stealth request timeout in seconds
config.STEALTH_DRIVER = "turbo"  # "basic" (default), "turbo", "browser", or "auto"
config.STEALTH_AUTO_FALLBACK = True  # basic/turbo -> browser on JS challenge (headless=False)
config.HTTP2 = True  # False for servers that only support HTTP/1.1
config.BLOCK_CODES |= {407}  # extend blocked status codes (|= keeps defaults)
config.BLOCK_KEYWORDS.append("banned")  # extend blocked body-text patterns
config.BROWSER_HEADLESS = True  # browser driver: headless mode (False = visible window, more stealthy)
config.BROWSER_SETTLE_S = 4.0  # browser driver: seconds to wait after navigation for JS to finish
config.BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser"  # custom browser binary (default: auto-detect Chrome)
config.STEALTH_RECYCLE_AFTER_BANS = 5  # recycle Chrome / HTTP sessions after 5 consecutive bans
config.BROWSER_STATIC_ASSETS_BLOCK = True  # block images/fonts/CSS/media (skipped when snapshot=True)
config.BROWSER_PROXY_BYPASS_LIST = ["example.com", "*.internal"]  # these bypass the proxy
config.STEALTH_DNS_OVERRIDES = {"example.com": "203.0.113.10"}  # pin host โ†’ origin IP


class MySpider(scrapy.Spider):
    name = "example"
    ...
# โŒ wrong โ€” too late, the engine client is already created
class MySpider(scrapy.Spider):
    def start_requests(self):
        config.HTTP2 = False  # has no effect
        ...

You can also read any value programmatically:

config.get("DEFAULT_ENGINE")  # "scrapy"
config.get("MISSING_KEY", "default")  # "default"
Attribute Type Default Description
DEFAULT_ENGINE str "scrapy" Engine used when request.meta["stealth"] key is absent
DEFAULT_PROFILE str "chrome_147" Browser profile used when none is specified
DEFAULT_TIMEOUT int 30 Request timeout in seconds
STEALTH_DRIVER str "basic" Default driver: "basic", "turbo", "browser", or "auto". Also readable from Scrapy settings as STEALTH_DRIVER
STEALTH_AUTO_FALLBACK bool False Smart browser selection: when True, basic / turbo responses that look like a JS challenge or session ban are retried once with browser in visible mode (headless=False). Per-request equivalent: meta["stealth"]["driver"] = "auto"
HTTP2 bool True HTTP/2 mode; overridable per-request via meta["stealth"]["http2"]
BLOCK_CODES frozenset[int] {403, 429, 503} HTTP status codes considered blocked
BLOCK_KEYWORDS list[str] ["captcha", "access denied", โ€ฆ] Body-text patterns considered blocked
BROWSER_HEADLESS bool True Browser driver: headless mode (False = visible window, more stealthy)
BROWSER_SETTLE_S float 4.0 Browser driver: seconds to wait after navigation for JS to finish rendering
BROWSER_NO_SANDBOX bool | None None Browser driver: disable Chrome sandbox. None = auto-detect (enabled when running as root, e.g. Docker)
BROWSER_EXECUTABLE_PATH str | None None Browser driver: path to the browser binary. None = auto-detect Chrome/Chromium. Set to use Brave or a custom install (e.g. "/usr/bin/brave-browser")
BROWSER_MAX_TABS int 10 Browser driver: max concurrent Chrome tabs across in-flight requests
STEALTH_RECYCLE_AFTER_BANS int 5 After this many consecutive bans: browser restarts Chrome; basic / turbo clear cached HTTP sessions/clients. Any clean response resets the count
BROWSER_STATIC_ASSETS_BLOCK bool False Browser driver: block images, fonts, CSS, and media via CDP. Overridable per-request via meta["stealth"]["static_assets_block"]; always off when snapshot=True
BROWSER_PROXY_BYPASS_LIST list[str] [] Browser driver: domains/patterns that bypass the proxy and connect to the origin directly, via Chrome's --proxy-bypass-list. Supports wildcards (*.example.com), IP/CIDR, ports, and <local>. Only applies when a proxy is in use; set at browser launch (config/settings, not per-request)
STEALTH_DNS_OVERRIDES dict[str, str] {} Hostโ†’IP map used by basic / turbo (and Chrome --host-resolver-rules for browser). Connects to the IP while keeping the hostname for TLS SNI, Host header, and cert verification. Also readable from Scrapy settings as STEALTH_DNS_OVERRIDES. Per-request override via meta["stealth"]["dns"]

For one-off overrides on a single request, set meta["stealth"]["driver"] or meta["stealth"]["http2"] (see Per-Request Configuration below).


โš™๏ธ Per-Request Configuration

All options are passed via request.meta["stealth"].

The presence of meta["stealth"] (a dict) activates the stealth engine. Omit the key to use the default Scrapy engine. When STEALTH_ENABLED = True, all requests are stealth by default โ€” pass meta={"stealth": False} to opt out for a specific request.

yield scrapy.Request(
    url,
    meta={
        "stealth": {
            "driver": "turbo",
            # optional overrides โ€” otherwise profile/proxy come from defaults and
            # rotate automatically when the session recycles after consecutive bans
            "profile": "chrome_147",
            "proxy": "http://user:pass@proxy:8080",
            "stealth_timeout": 60,
            "http2": True,
            "dns": "203.0.113.10",  # or {"example.com": "203.0.113.10"}
        }
    },
)
Key Type Description
driver str "basic", "turbo", "browser", or "auto" โ€” "auto" enables smart browser selection for this request (HTTP first, then browser on challenge/ban). Requires STEALTH_AUTO_FALLBACK = True globally, or use "auto" alone to opt in per-request
fallback bool Set to False to opt out of auto-fallback for this request (when STEALTH_AUTO_FALLBACK or driver="auto" is active)
profile str Browser profile (e.g. "chrome_147", "safari_ios_18_1_1"). Omit to use engine default; default rotates on ban-streak session recycle
proxy str Explicit proxy URL. Omit to use STEALTH_PROXIES default; default rotates on ban-streak session recycle
dns str or dict Pin DNS: bare IP for this request's hostname, or {host: ip} mapping. Merges over STEALTH_DNS_OVERRIDES. Works with basic/turbo per-request; browser uses global overrides at Chrome launch only
stealth_timeout int Per-request timeout in seconds (overrides default 30s)
http2 bool True = HTTP/2, False = HTTP/1.1 (overrides config.HTTP2 for this request)
headless bool Browser driver only: True = headless, False = visible window (more stealthy)
settle float Browser driver only: seconds to wait for JS after navigation (default 4.0)
snapshot bool Browser driver only: capture a PNG snapshot โ€” result available as response.meta["snapshot_content"] (bytes)
static_assets_block bool Browser driver only: block images, fonts, CSS, and media for this request (overrides config.BROWSER_STATIC_ASSETS_BLOCK). Ignored โ€” always unblocked โ€” when snapshot is True

๐Ÿงญ Custom DNS Overrides

Pin a hostname to a fixed origin IP so the package dials that address directly instead of trusting public DNS. The request URL stays as https://example.com/... โ€” TLS SNI, the Host header, and certificate verification still use the hostname.

Global (settings.py / custom_settings / config):

STEALTH_DNS_OVERRIDES = {
    "shop.example.com": "203.0.113.10",
    "cdn.example.com": "203.0.113.11",
}

Per-request (overrides or extends the global map):

yield scrapy.Request(
    "https://shop.example.com/item/1",
    meta={"stealth": {"driver": "turbo", "dns": "203.0.113.10"}},
)

# Or a full mapping:
meta = {"stealth": {"dns": {"shop.example.com": "203.0.113.10"}}}

Supported on basic and turbo per-request. The browser driver applies the effective map (config + meta["stealth"]["dns"]) via a local CONNECT relay that dials the pinned IP (Chrome's --host-resolver-rules is not used โ€” it is unreliable). Chrome is pointed at the relay with --proxy-server; when the DNS map changes, the browser restarts so the relay is rebuilt. Do not put DNS-pinned hosts on BROWSER_PROXY_BYPASS_LIST or they will skip the relay. With an HTTP proxy on basic/turbo, DNS is often resolved by the proxy โ€” prefer direct connections or SOCKS when using overrides.

๐Ÿง  Smart browser selection

Pick the right driver automatically: stay on fast HTTP impersonation (basic / turbo) for normal pages, and escalate to real Chrome only when the response looks like a JS challenge or session ban (403/429/503, Cloudflare โ€œJust a momentโ€, Akamai, DataDome, and similar signals).

Phase Driver When
1 basic or turbo Default โ€” low memory, high throughput
2 browser (headless=False) One retry when phase 1 is blocked or challenged

The fallback always opens a visible Chrome window (headless=False) for better evasion โ€” regardless of BROWSER_HEADLESS or any prior meta["stealth"]["headless"] value.

Global (settings.py / custom_settings):

STEALTH_ENABLED = True
STEALTH_DRIVER = "turbo"  # or "basic" โ€” primary HTTP driver
STEALTH_AUTO_FALLBACK = True  # off by default; set True to enable smart selection globally

Per-request (no global setting โ€” equivalent to enabling fallback for that URL only):

yield scrapy.Request(
    url,
    meta={"stealth": {"driver": "auto"}},  # primary = STEALTH_DRIVER, else basic
)

Always use browser (skip phase 1):

meta={"stealth": {"driver": "browser"}}

Opt out for one request:

meta={"stealth": {"driver": "auto", "fallback": False}}
# or globally: STEALTH_AUTO_FALLBACK = False

Each request is retried at most once. If the browser fetch fails, the original basic / turbo response is returned. Console output and stats (stealth/fallbacks, stealth/requests/browser) show when escalation happened.


๐Ÿ–ฅ๏ธ Browser Engine

For sites protected by Cloudflare JS challenges or heavy JavaScript rendering, use the browser driver. It runs a real Chrome instance via the DevTools Protocol (no WebDriver), keeping one persistent browser and opening a new tab per request.

Per-request (most common):

yield scrapy.Request(
    url,
    meta={
        "stealth": {
            "driver": "browser",
            "headless": False,  # visible window โ€” harder to detect (default: True)
            "settle": 4.0,  # seconds to wait for JS after page load
        }
    },
)

Heavy Cloudflare sites โ€” increase settle time:

meta = {"stealth": {"driver": "browser", "headless": False, "settle": 12}}

Global default (all stealth requests use browser engine):

from scrapy_stealth.config import config

config.STEALTH_DRIVER = "browser"
config.BROWSER_HEADLESS = False  # more stealthy
config.BROWSER_SETTLE_S = 6.0  # longer wait for JS

Custom browser binary (Brave, Chromium, or a non-default Chrome install):

from scrapy_stealth.config import config

config.BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser"  # Linux
# config.BROWSER_EXECUTABLE_PATH = r"C:\Program Files\BraveSoftware\Brave-Browser\Application\brave.exe"  # Windows

Or via settings.py / custom_settings:

BROWSER_EXECUTABLE_PATH = "/usr/bin/brave-browser"

When BROWSER_EXECUTABLE_PATH is None (the default), scrapy-stealth auto-detects Google Chrome or Chromium from standard system paths. Set it explicitly when using Brave or a non-standard Chrome installation โ€” a clear error is raised if the path does not exist.

Intelligent restart / session recycle:

After STEALTH_RECYCLE_AFTER_BANS consecutive banned/challenged responses (as classified by Anti-Bot Detection), scrapy-stealth recycles the driver session:

  • browser โ€” restarts Chrome (fresh fingerprint, cookies, CDP session)
  • basic / turbo โ€” clears cached HTTP clients/sessions and rotates default fingerprint profile + proxy from STEALTH_PROXIES (no per-request rotate_* needed)

A single clean response resets the streak, so a healthy crawl is never recycled just because it has served a lot of requests.

from scrapy_stealth.config import config

config.STEALTH_RECYCLE_AFTER_BANS = 5  # recycle after 5 consecutive bans (default)

Static asset blocking:

scrapy-stealth can block static assets (images, fonts, CSS, and media) in the browser to speed up page loads and cut bandwidth, via the CDP Fetch domain. It's off by default โ€” enable it globally with BROWSER_STATIC_ASSETS_BLOCK = True in settings.py, or per-request via meta["stealth"]["static_assets_block"]. Blocking is always skipped when snapshot=True, since a snapshot needs the fully rendered page.

# settings.py
BROWSER_STATIC_ASSETS_BLOCK = True
# per-request
meta = {"stealth": {"driver": "browser", "static_assets_block": True}}
# snapshot always wins โ€” assets are never blocked here, even with the global default on
meta = {"stealth": {"driver": "browser", "snapshot": True}}

Proxy bypass list:

When a proxy is configured, you can send specific domains straight to the origin instead of through the proxy. The list is passed to Chrome's --proxy-bypass-list launch flag, so it supports the full Chrome bypass syntax โ€” bare hostnames, wildcards (*.example.com), IP/CIDR ranges, ports, and the special <local> token.

# settings.py
STEALTH_DRIVER = "browser"
STEALTH_PROXIES = ["http://user:pass@proxy:8080"]
BROWSER_PROXY_BYPASS_LIST = [
    "example.com",  # exact host
    "*.internal.net",  # wildcard subdomains
    "127.0.0.1",  # IP
    "<local>",  # any plain hostname without dots
]
# or via config
from scrapy_stealth.config import config

config.BROWSER_PROXY_BYPASS_LIST = ["example.com", "*.internal.net"]

The bypass list is a Chrome launch flag, so it's read once when the browser starts and applies to the whole browser lifetime โ€” it's configured globally (config/settings), not per-request. It has no effect unless a proxy is in use.

Docker (running as root):

Chrome requires --no-sandbox when the process runs as root. scrapy-stealth detects this automatically, but you can also set it explicitly in settings.py:

BROWSER_NO_SANDBOX = True  # force no-sandbox (Docker, any root environment)
BROWSER_EXECUTABLE_PATH = "/usr/bin/chromium"  # use Chromium instead of Chrome in Docker

Or via config:

config.BROWSER_NO_SANDBOX = True
config.BROWSER_EXECUTABLE_PATH = "/usr/bin/chromium"

Performance note: the browser engine is slower than basic/turbo (~5-15s per page vs <2s). Use it selectively โ€” route only JS-protected URLs to "browser" and keep everything else on "turbo".


๐Ÿ“ธ Screenshots

Capture a PNG screenshot of any page rendered by the browser driver and save it to disk.

Enable on the request

yield scrapy.Request(
    url,
    meta={
        "stealth": {
            "driver": "browser",
            "snapshot": True,
        }
    },
    callback=self.parse,
)

The raw PNG bytes are available at response.meta["snapshot_content"] inside your callback.

Auto-save with snapshot decorator

from scrapy_stealth.decorators import snapshot


class MySpider(scrapy.Spider):

    @snapshot
    def parse(self, response): ...

    @snapshot(path="stealth_shots/page.png")
    def parse(self, response): ...

    @snapshot(path=lambda r: r.url.split("/")[-1] + ".png")
    def parse(self, response): ...

Note: Requires driver="browser" and snapshot=True in the request meta. Logs an error if no snapshot data is found in the response.

Custom handling (without the built-in helper)

The screenshot is just bytes in response.meta["snapshot_content"] โ€” do anything you like with it:

def parse(self, response):
    shot: bytes | None = response.meta.get("snapshot_content")
    if shot is None:
        return  # screenshot was not requested or capture failed

    # Save manually
    with open("page.png", "wb") as f:
        f.write(shot)

    # Pass to a pipeline via item
    yield {"url": response.url, "screenshot": shot}

๐Ÿ” Automatic Rotation

Profile + proxy stay stable for speed (session reuse). After STEALTH_RECYCLE_AFTER_BANS consecutive bans, the session recycles and a new default profile + proxy (from STEALTH_PROXIES) are chosen automatically.

# settings.py
STEALTH_PROXIES = ["http://proxy1:8080", "http://proxy2:8080"]
STEALTH_RECYCLE_AFTER_BANS = 5  # default

# spider
yield scrapy.Request(url, meta={"stealth": {}})

Scrapy stats: after the crawl (or mid-run via crawler.stats), inspect:

Key Meaning
stealth/requests / stealth/requests/{driver} Stealth fetches
stealth/responses / stealth/responses/{driver} Completed responses
stealth/successes / stealth/successes/{driver} Non-banned responses below HTTP 400
stealth/failures / stealth/failures/{driver} Banned responses or HTTP 400+
stealth/status/{code} Response count by HTTP status
stealth/bans / stealth/bans/{driver} Session-ban responses
stealth/recycles / stealth/recycles/{driver} Session / Chrome recycles
stealth/ban_streak Current consecutive ban streak
stealth/driver Last stealth driver used
stealth/profile Last fingerprint profile used
stealth/proxy Last proxy as host:port (no credentials)
stealth/proxy/requests/{driver} Requests sent through a proxy
stealth/dns/requests/{driver} Requests using DNS overrides
stealth/dns/hosts Total pinned hosts applied
stealth/dns/active_hosts Pinned hosts on latest request
# e.g. in spider_closed
stats = spider.crawler.stats.get_stats()
print(stats.get("stealth/bans"), stats.get("stealth/recycles"))

๐Ÿงฉ Strategies

Proxy Rotation

from scrapy_stealth.strategies.proxy import ProxyRotator

proxy_rotator = ProxyRotator([
    "http://proxy1:8080",
    "http://proxy2:8080",
])

yield scrapy.Request(
    url,
    meta={
        "stealth": {
            "proxy": proxy_rotator.get(),
        }
    },
)

Fingerprint Rotation

from scrapy_stealth.strategies.fingerprint import ProfileRotator

fp = ProfileRotator()

yield scrapy.Request(
    url,
    meta={
        "stealth": {
            "profile": fp.get(),
        }
    },
)

Intelligent Retry

from scrapy_stealth.strategies.retry import RetryHandler

retry = RetryHandler()


def parse(self, response):
    if retry.should_retry(response):
        yield retry.build(response.request)
        return

๐Ÿ›ก๏ธ Anti-Bot Detection

from scrapy_stealth.detectors.antibot import AntiBotDetector

detector = AntiBotDetector()

if detector.is_blocked(response):
    print("Blocked!")

๐Ÿ“Š Full spider example

Keep the README short โ€” the complete working spider lives in examples/full_spider.py.

It shows:

  • middleware + STEALTH_ENABLED via custom_settings
  • default turbo driver
  • per-request basic / browser overrides
  • optional snapshot with @snapshot
  • ban detection and stealth stats on close
# from a Scrapy project
scrapy crawl stealth_demo

# or one-off
scrapy runspider examples/full_spider.py

Minimal version:

import scrapy


class ExampleSpider(scrapy.Spider):
    name = "example"
    custom_settings = {
        "DOWNLOADER_MIDDLEWARES": {
            "scrapy_stealth.middlewares.StealthDownloaderMiddleware": 950,
        },
        "STEALTH_ENABLED": True,
        "STEALTH_DRIVER": "turbo",
    }

    def start_requests(self):
        yield scrapy.Request("https://example.com")

    def parse(self, response):
        yield {"title": response.css("title::text").get(), "url": response.url}

โšก Performance Insight

Using stealth selectively:

  • โšก Faster crawling (Scrapy for simple pages)
  • ๐Ÿ’ฐ Lower proxy cost
  • ๐Ÿ›ก๏ธ Better success rate on protected pages

๐Ÿ“œ Changelog

See CHANGELOG.md for a full history of changes, or browse GitHub Releases.


๐Ÿค Contributing

See CONTRIBUTING.md for guidelines on how to contribute.


๐Ÿ“„ License

This project is licensed under the MIT License โ€” free to use, modify, and distribute. See LICENSE for the full text.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapy_stealth-0.6.12.tar.gz (225.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapy_stealth-0.6.12-py3-none-any.whl (216.4 kB view details)

Uploaded Python 3

File details

Details for the file scrapy_stealth-0.6.12.tar.gz.

File metadata

  • Download URL: scrapy_stealth-0.6.12.tar.gz
  • Upload date:
  • Size: 225.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scrapy_stealth-0.6.12.tar.gz
Algorithm Hash digest
SHA256 a95b8d483050c8916b7a49844b2a938681df3067b1125e0c97ac4c0a5bd1776f
MD5 027f3750ea15e247972e3c099356473a
BLAKE2b-256 a1ac736f35394dd8b58c11b905838b6284dbe0a6fd8d4d01bd67f2a307f5934c

See more details on using hashes here.

Provenance

The following attestation bundles were made for scrapy_stealth-0.6.12.tar.gz:

Publisher: publish.yml on fawadss1/scrapy-stealth

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file scrapy_stealth-0.6.12-py3-none-any.whl.

File metadata

  • Download URL: scrapy_stealth-0.6.12-py3-none-any.whl
  • Upload date:
  • Size: 216.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scrapy_stealth-0.6.12-py3-none-any.whl
Algorithm Hash digest
SHA256 fa6935b5e7a63c6d8d85d6cb9a4162e859d5affc8648f26b94a950a42486e90f
MD5 d4881fe6cdd639fd0c4da72db0665428
BLAKE2b-256 12467bf5b1df4cb91d268a5caa39ebf7333541a417112aa7d0e5d2d48806d4b8

See more details on using hashes here.

Provenance

The following attestation bundles were made for scrapy_stealth-0.6.12-py3-none-any.whl:

Publisher: publish.yml on fawadss1/scrapy-stealth

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page