Skip to main content

scrapy-cloud-browser

PyPI - Version PyPI - Python Version

A Scrapy extension that integrates with surfsky.io for web scraping. This extension allows you to use Chrome anti-detection browsers through Surfsky.io's cloud browser service, helping you avoid detection while scraping challenging websites.


Installation

pip install scrapy-cloud-browser

Usage

Setup environment variables in settings.py in CLOUD_BROWSER namespace:

CLOUD_BROWSER = {
    "API_HOST": <HOST>,
    "API_TOKEN": <API_TOKEN>,
    "NUM_BROWSERS": <NUM_BROWSERS>,
    "PROXIES": [<proxy>],
    "PAGES_PER_BROWSER": <PAGES_PER_BROWSER>,
    "START_SEMAPHORES": <START_SEMAPHORES>,
    "PROXY_ORDERING": <PROXY_ORDERING>,
    "BROWSER_SETTINGS: <BROWSER_SETTINGS>,
    "FINGERPRINT": <FINGERPRINT>
}

Configuration Parameters

  • API_HOST: The URL of the Surfsky.io API host.
  • API_TOKEN: Your authentication token for the Surfsky.io service.
  • NUM_BROWSERS (default: 1): Number of browser instances to run in parallel. Increase this value to improve throughput for large-scale scraping.
  • PROXIES: A list of proxy URLs to use with the browsers. Each browser will be assigned a proxy from this list according to the PROXY_ORDERING strategy.
  • PAGES_PER_BROWSER (default: 100): Maximum number of pages a browser instance will process before being recycled.
  • START_SEMAPHORES (default: 10): Controls how many browsers can be started simultaneously. This prevents overwhelming the system with too many concurrent browser startups.
  • PROXY_ORDERING (default: 'random'): Strategy for assigning proxies to browsers. Options are:
    • 'random': Randomly select a proxy from the list for each browser
    • 'round-robin': Cycle through the proxy list in order

For detailed browser settings (BROWSER_SETTINGS) and fingerprint configuration options (FINGERPRINT), please refer to the Surfsky API Reference.

Add cloud browser handlers and change reactor in settings.py:

TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"

EXTENSIONS = {
    'scrapy_cloud_browser.CloudBrowserExtension': 100,
}

License

scrapy-cloud-browser is distributed under the terms of the MIT license.

Release files for scrapy-cloud-browser 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrapy-cloud-browser 0.1.1
File Size Uploaded
scrapy_cloud_browser-0.1.1.tar.gz 12.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrapy-cloud-browser 0.1.1
File Interpreter ABI Platform
scrapy_cloud_browser-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 25.4 kB

Release files / scrapy_cloud_browser-0.1.1.tar.gz

Download URL scrapy_cloud_browser-0.1.1.tar.gz
Size 12.8 kB
Tags Source
SHA-256 checksum
How to use checksums
39398415a03015694ad1b432f6cce2f54500ea37e5809d210203d0b7d360be36
BLAKE2b-256 checksum
How to use checksums
d2aa2592e8adf8d4b211cf9ae599c0235d1df8fa0b1a22cebb46b5c9817370c4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.3

Release files / scrapy_cloud_browser-0.1.1-py3-none-any.whl

Download URL scrapy_cloud_browser-0.1.1-py3-none-any.whl
Size 12.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9d1920b9ded4954afc805882201c95d7c3d505e786562e038d3a97027f86ffa1
BLAKE2b-256 checksum
How to use checksums
cfc25a68bc198baede5863b6a512c18907ab36d892fd89a39e889fce17e9fd9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page