Scrapy downloader middleware for FlareSolverr with multiple backends and concurrency control
Project description
scrapy-flaresolverr
A Scrapy downloader middleware for routing opt-in requests through one or more FlareSolverr backends.
scrapy-flaresolverr provides a small integration layer between Scrapy and the
FlareSolverr v1 API, with support for multiple backends, round-robin selection,
request-level options, bounded concurrency, response reconstruction, and Scrapy
stats.
Features
- Stateless FlareSolverr
request.getintegration - Multiple FlareSolverr backends with thread-safe round-robin selection
- Request-level backend override
- Optional Bearer authentication
- Configurable request and FlareSolverr timeouts
- Configurable concurrency limits
- Request-level media disabling and wait time
- Optional screenshot return
- Reconstruction of complete Scrapy
HtmlResponseobjects - FlareSolverr response metadata exposed through
response.meta - Scrapy stats for requests, responses, errors, and concurrency timeouts
- Explicit package-specific exceptions for configuration, transport, response, concurrency, and unsupported-request failures
Requirements
- Python 3.10+
- Scrapy 2.12+
- requests 2.31+
The package supports Scrapy 2.x and currently declares:
Python >=3.10
Scrapy >=2.12,<3
requests >=2.31,<3
CI compatibility matrix
The build and functional test workflows cover the following compatibility baseline:
| Python | Scrapy 2.12.x | Latest supported Scrapy 2.x |
|---|---|---|
| 3.10 | ✓ | ✓ |
| 3.11 | ✓ | ✓ |
| 3.12 | ✓ | ✓ |
| 3.13 | — | ✓ |
| 3.14 | — | ✓ |
The minimum Scrapy line is tested on Python versions supported by that combination, while current Scrapy 2.x is tested across all supported Python versions.
Installation
Install from PyPI:
pip install scrapy-flaresolverr
For local development:
git clone https://github.com/geeone/scrapy-flaresolverr.git
cd scrapy-flaresolverr
pip install -r requirements-dev.txt
Quick start
Enable the downloader middleware in your Scrapy settings:
DOWNLOADER_MIDDLEWARES = {
"scrapy_flaresolverr.middleware.FlareSolverrMiddleware": 100,
}
FLARESOLVERR_URLS = [
"http://127.0.0.1:8191",
]
FLARESOLVERR_MAX_CONCURRENT = 3
Then opt individual requests into FlareSolverr:
yield scrapy.Request(
"https://example.com",
meta={"flaresolverr": True},
)
Request-specific options can be passed as a dictionary:
yield scrapy.Request(
"https://example.com",
meta={
"flaresolverr": {
"disable_media": True,
"wait_in_seconds": 2,
"return_screenshot": True,
},
},
)
Requests without the flaresolverr opt-in continue through Scrapy's normal
downloader middleware chain.
Configuration
Backend URLs
Configure one or more FlareSolverr backends:
FLARESOLVERR_URLS = [
"http://127.0.0.1:8191",
"http://127.0.0.1:8192",
]
Configured backend URLs are normalized to the FlareSolverr /v1 endpoint.
For example:
http://127.0.0.1:8191 -> http://127.0.0.1:8191/v1
http://127.0.0.1:8191/fs1/ -> http://127.0.0.1:8191/fs1/v1
A single backend can alternatively be configured with:
FLARESOLVERR_URL = "http://127.0.0.1:8191"
When multiple backends are configured, requests are distributed using round-robin selection.
Authentication
If the FlareSolverr endpoint or an upstream reverse proxy requires Bearer authentication:
FLARESOLVERR_AUTH_TOKEN = "your-token"
The token is sent as:
Authorization: Bearer <token>
Settings reference
| Setting | Default | Description |
|---|---|---|
FLARESOLVERR_URLS |
— | List of FlareSolverr backend URLs |
FLARESOLVERR_URL |
— | Single-backend fallback when FLARESOLVERR_URLS is not set |
FLARESOLVERR_AUTH_TOKEN |
None |
Optional Bearer token |
FLARESOLVERR_MAX_CONCURRENT |
3 |
Maximum number of concurrent FlareSolverr calls |
FLARESOLVERR_CONCURRENCY_WAIT_TIMEOUT |
120.0 |
Maximum seconds to wait for a concurrency slot |
FLARESOLVERR_MAX_TIMEOUT |
60000 |
FlareSolverr maxTimeout in milliseconds |
FLARESOLVERR_REQUEST_TIMEOUT |
75.0 |
HTTP timeout in seconds for the call to FlareSolverr |
FLARESOLVERR_DISABLE_MEDIA |
False |
Disable media by default for FlareSolverr requests |
FLARESOLVERR_REQUEST_TIMEOUT must be greater than
FLARESOLVERR_MAX_TIMEOUT converted to seconds.
Request options
Per-request behavior is configured through request.meta["flaresolverr"].
yield scrapy.Request(
"https://example.com",
meta={
"flaresolverr": {
"backend": "http://127.0.0.1:8191",
"max_timeout": 90_000,
"request_timeout": 110,
"wait_in_seconds": 2,
"disable_media": True,
"return_screenshot": False,
},
},
)
| Option | Type | Description |
|---|---|---|
backend |
str |
Use a specific configured backend for this request |
max_timeout |
int |
Override FlareSolverr maxTimeout in milliseconds |
request_timeout |
float |
Override the HTTP timeout for the FlareSolverr call |
wait_in_seconds |
float |
Additional FlareSolverr wait time; may be zero |
disable_media |
bool |
Override the global media-loading setting |
return_screenshot |
bool |
Request a screenshot from FlareSolverr |
A backend override must resolve to one of the configured backends.
Multiple backends
Multiple backends are selected in round-robin order:
FLARESOLVERR_URLS = [
"http://127.0.0.1:8191/fs1/",
"http://127.0.0.1:8191/fs2/",
]
The backend selected for a request is exposed as:
response.meta["flaresolverr_backend"]
A request can explicitly target one configured backend:
yield scrapy.Request(
"https://example.com",
meta={
"flaresolverr": {
"backend": "http://127.0.0.1:8191/fs2/",
},
},
)
Response metadata
Successful FlareSolverr requests are returned as Scrapy HtmlResponse
instances.
The middleware marks the request with:
response.meta["flaresolverr_used"]
response.meta["flaresolverr_backend"]
Additional FlareSolverr response data is available under:
response.meta["flaresolverr_response"]
The dictionary contains:
{
"cookies": ...,
"user_agent": ...,
"screenshot": ...,
"start_timestamp": ...,
"end_timestamp": ...,
"version": ...,
}
The reconstructed response also includes the flaresolverr response flag.
Because FlareSolverr returns decoded HTML, transport-specific headers such as
Content-Encoding, Content-Length, and Transfer-Encoding are not copied to
the reconstructed Scrapy response. The response body is encoded as UTF-8 and
the Content-Type charset is normalized accordingly.
Stats
The middleware records counters through Scrapy's stats collector:
flaresolverr/request_count
flaresolverr/response_count
flaresolverr/request_error
flaresolverr/concurrency_timeout
flaresolverr/unsupported_request
They can be inspected together with the rest of the spider's Scrapy stats.
Error handling
scrapy-flaresolverr does not silently fall back to Scrapy's normal downloader
after a FlareSolverr failure. Errors are propagated through package-specific
exceptions:
FlareSolverrConfigurationErrorFlareSolverrConcurrencyErrorFlareSolverrRequestErrorFlareSolverrResponseErrorFlareSolverrUnsupportedRequestError
This keeps FlareSolverr failures explicit and observable by the calling spider or Scrapy error handling.
Current limitations
Version 0.1.0 supports stateless GET requests only.
Current limitations include:
- no FlareSolverr sessions or session affinity
- no proxy configuration or proxy affinity
- no automatic backend failover or health-based routing
See the roadmap for planned development.
Examples
The repository includes runnable examples:
- Basic spider — middleware configuration, request options, and selected-backend metadata
- Docker Compose — two FlareSolverr workers behind an authenticated Nginx reverse proxy
The Docker Compose example can be used together with the basic spider to exercise multiple-backend routing locally.
Development
Install the development dependencies:
pip install -r requirements-dev.txt
Run the test suite:
pytest --cov=scrapy_flaresolverr --cov-report=term-missing tests/
Run the configured pre-commit checks:
pre-commit run --all-files
Or install the hooks locally:
pre-commit install
The project uses Ruff for linting and formatting, mypy for static type checking, pytest for testing, and pytest-cov for coverage.
For an overview of the internal design, see the architecture documentation.
Contributing
Bug reports, feature requests, and pull requests are welcome.
Before contributing, please review:
Security-related issues should be reported according to the Security Policy, rather than through a public issue.
Release history is documented in the Changelog.
Support the project
If scrapy-flaresolverr is useful to you, consider giving the project a ⭐ on
GitHub.
License
MIT. See LICENSE for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file scrapy_flaresolverr-0.1.0.tar.gz.
File metadata
- Download URL: scrapy_flaresolverr-0.1.0.tar.gz
- Upload date:
- Size: 29.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
17d21dd5b5e378497f0fb7fcec7f835e717f2682adba06b99cd76565bb0510d2
|
|
| MD5 |
89da309f14ff7048a015b7ae9c44fd03
|
|
| BLAKE2b-256 |
28abbc58ffa889a363b1b50cb3ea416fb5c2e0a93ea82dd4933a1adee8f87dc3
|
Provenance
The following attestation bundles were made for scrapy_flaresolverr-0.1.0.tar.gz:
Publisher:
release.yml on geeone/scrapy-flaresolverr
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
scrapy_flaresolverr-0.1.0.tar.gz -
Subject digest:
17d21dd5b5e378497f0fb7fcec7f835e717f2682adba06b99cd76565bb0510d2 - Sigstore transparency entry: 2256580090
- Sigstore integration time:
-
Permalink:
geeone/scrapy-flaresolverr@9cfc008608be8ccf44eceab797cdc979232d6e74 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/geeone
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@9cfc008608be8ccf44eceab797cdc979232d6e74 -
Trigger Event:
release
-
Statement type:
File details
Details for the file scrapy_flaresolverr-0.1.0-py3-none-any.whl.
File metadata
- Download URL: scrapy_flaresolverr-0.1.0-py3-none-any.whl
- Upload date:
- Size: 16.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
395dd761967317f5934735e086df901bb35e776c1a4c202cb590c45f8fadb969
|
|
| MD5 |
b739462c7a4649a9fce39d504a6513e8
|
|
| BLAKE2b-256 |
707d16a2b721b13f5cf86dfe05066411c97339067d576098fef626da19f99072
|
Provenance
The following attestation bundles were made for scrapy_flaresolverr-0.1.0-py3-none-any.whl:
Publisher:
release.yml on geeone/scrapy-flaresolverr
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
scrapy_flaresolverr-0.1.0-py3-none-any.whl -
Subject digest:
395dd761967317f5934735e086df901bb35e776c1a4c202cb590c45f8fadb969 - Sigstore transparency entry: 2256580108
- Sigstore integration time:
-
Permalink:
geeone/scrapy-flaresolverr@9cfc008608be8ccf44eceab797cdc979232d6e74 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/geeone
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@9cfc008608be8ccf44eceab797cdc979232d6e74 -
Trigger Event:
release
-
Statement type: