Does your scrapy spider get identified and blocked by servers because you use the default user-agent or a generic one?
Use this random_useragent module and set a random user-agent for every request. You are limited only by the number of different user-agents you set in a text file.
Installing
Installing it is pretty simple.
pip install scrapy-random-useragent
Usage
In your settings.py file, update the DOWNLOADER_MIDDLEWARES variable like this.
DOWNLOADER_MIDDLEWARES = {
'scrapy.contrib.downloadermiddleware.useragent.UserAgentMiddleware': None,
'random_useragent.RandomUserAgentMiddleware': 400
}
This disables the default UserAgentMiddleware and enables the RandomUserAgentMiddleware.
Then, create a new variable USER_AGENT_LIST with the path to your text file which has the list of all user-agents (one user-agent per line).
USER_AGENT_LIST = "/path/to/useragents.txt"
Now all the requests from your crawler will have a random user-agent picked from the text file.
Release files for scrapy-random-useragent 0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scrapy-random-useragent-0.2.tar.gz | 2.6 kB | Details |
Release files / scrapy-random-useragent-0.2.tar.gz
| Download URL | scrapy-random-useragent-0.2.tar.gz |
|---|---|
| Size | 2.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
570f87e26438f7e1b69890219a6e052b8c510a745a21ae129da8f8fe8161e102
|
|
BLAKE2b-256 checksum How to use checksums |
232e3a3ae91faf1d5d31526379186285817bda8ff66a221ec7085a9e549c1465
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |