scrapy-cloud-browser
A Scrapy extension that integrates with surfsky.io for web scraping. This extension allows you to use Chrome anti-detection browsers through Surfsky.io's cloud browser service, helping you avoid detection while scraping challenging websites.
Installation
pip install scrapy-cloud-browser
Usage
Setup environment variables in settings.py in CLOUD_BROWSER namespace:
CLOUD_BROWSER = {
"API_HOST": <HOST>,
"API_TOKEN": <API_TOKEN>,
"NUM_BROWSERS": <NUM_BROWSERS>,
"PROXIES": [<proxy>],
"PAGES_PER_BROWSER": <PAGES_PER_BROWSER>,
"START_SEMAPHORES": <START_SEMAPHORES>,
"PROXY_ORDERING": <PROXY_ORDERING>,
"BROWSER_SETTINGS: <BROWSER_SETTINGS>,
"FINGERPRINT": <FINGERPRINT>
}
Configuration Parameters
- API_HOST: The URL of the Surfsky.io API host.
- API_TOKEN: Your authentication token for the Surfsky.io service.
- NUM_BROWSERS (default: 1): Number of browser instances to run in parallel. Increase this value to improve throughput for large-scale scraping.
- PROXIES: A list of proxy URLs to use with the browsers. Each browser will be assigned a proxy from this list according to the PROXY_ORDERING strategy.
- PAGES_PER_BROWSER (default: 100): Maximum number of pages a browser instance will process before being recycled.
- START_SEMAPHORES (default: 10): Controls how many browsers can be started simultaneously. This prevents overwhelming the system with too many concurrent browser startups.
- PROXY_ORDERING (default: 'random'): Strategy for assigning proxies to browsers. Options are:
- 'random': Randomly select a proxy from the list for each browser
- 'round-robin': Cycle through the proxy list in order
For detailed browser settings (BROWSER_SETTINGS) and fingerprint configuration options (FINGERPRINT), please refer to the Surfsky API Reference.
Add cloud browser handlers and change reactor in settings.py:
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
EXTENSIONS = {
'scrapy_cloud_browser.CloudBrowserExtension': 100,
}
License
scrapy-cloud-browser is distributed under the terms of the MIT license.
Release files for scrapy-cloud-browser 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scrapy_cloud_browser-0.1.1.tar.gz | 12.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scrapy_cloud_browser-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.4 kB
Release files / scrapy_cloud_browser-0.1.1.tar.gz
| Download URL | scrapy_cloud_browser-0.1.1.tar.gz |
|---|---|
| Size | 12.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
39398415a03015694ad1b432f6cce2f54500ea37e5809d210203d0b7d360be36
|
|
BLAKE2b-256 checksum How to use checksums |
d2aa2592e8adf8d4b211cf9ae599c0235d1df8fa0b1a22cebb46b5c9817370c4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.3
|
Release files / scrapy_cloud_browser-0.1.1-py3-none-any.whl
| Download URL | scrapy_cloud_browser-0.1.1-py3-none-any.whl |
|---|---|
| Size | 12.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9d1920b9ded4954afc805882201c95d7c3d505e786562e038d3a97027f86ffa1
|
|
BLAKE2b-256 checksum How to use checksums |
cfc25a68bc198baede5863b6a512c18907ab36d892fd89a39e889fce17e9fd9a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.3
|