Concurrent Flood Scraper
It’s probably exactly what you think it is, based off the name
GET a page. scrape for urls, filter those according to some regex. Put all those in a master queue. Scrape page for any data you want. Repeat…
There’s a small demo in the wikipedia_demo. There you can see how easy it is to set up to fit your web scraping needs!
Specifics
Create a child class of concurrentfloodscraper.Scraper and implement the scrape_page(self, text) method. text is the raw html. In this method you do the specific scraping required. Note that only urls that match the class url_filter_regex will be added to the master queue.
Annotate your Scraper subclass with concurrentfloodscraper.Route. The single parameter is a regex; URL’s that match the regex will be parsed with that scraper.
Repeat steps 1 and 2 for as many different types of pages you expect to be scraping from.
Create an instance of concurrentfloodscraper.ConcurrentFloodScraper, pass it the root URL, the number of threads to use, and a page limit. Page limit defaults to None, which means ‘go forever’.
Start the ConcurrentFloodScraper instance, and enjoy the magic!
Release files for concurrentfloodscraper 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| concurrentfloodscraper-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Release files / concurrentfloodscraper-1.0.1-py3-none-any.whl
| Download URL | concurrentfloodscraper-1.0.1-py3-none-any.whl |
|---|---|
| Size | 9.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
27a56763c000c81d987efc3bf82835772f0695899d088b61e434d29bf0fac8a8
|
|
BLAKE2b-256 checksum How to use checksums |
e000995311f710f0a7217b65cbf128a636d7d14e764a5e4883b9b5e6beb31a84
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |