Skip to main content

Does your scrapy spider get identified and blocked by servers because you use the default user-agent or a generic one?

Use this random_useragent module and set a random user-agent for every request.

Installing

Installing it is pretty simple.

pip install git+https://github.com/cleocn/scrapy-random-useragent.git

Usage

In your settings.py file, update the DOWNLOADER_MIDDLEWARES variable like this.

DOWNLOADER_MIDDLEWARES = {
    'scrapy.contrib.downloadermiddleware.useragent.UserAgentMiddleware': None,
    'random_useragent.RandomUserAgentMiddleware': 400
}

This disables the default UserAgentMiddleware and enables the RandomUserAgentMiddleware.

Now all the requests from your crawler will have a random user-agent.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapy-random-ua-0.3.tar.gz (3.2 kB view details)

Uploaded Source

File details

Details for the file scrapy-random-ua-0.3.tar.gz.

File metadata

File hashes

Hashes for scrapy-random-ua-0.3.tar.gz
Algorithm Hash digest
SHA256 0b21843e0eb86eb7351c0762c1851343b865d4a0b750146250f0c86581e2b972
MD5 2ac2a82c06b11252e0ce1b724416aaab
BLAKE2b-256 0b6fd48d5c0f651e2ede25637ca91584550f793dc967f68f56e56541a154e216

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.3 This release

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page