Scrapy Do is a daemon that provides a convenient way to run Scrapy spiders. It can either do it once - immediately; or it can run them periodically, at specified time intervals. It’s been inspired by scrapyd but written from scratch. It comes with a REST API, a command line client, and an interactive web interface.
Homepage: https://jany.st/scrapy-do.html
Documentation: https://scrapy-do.readthedocs.io/en/latest/
Quick Start
Install scrapy-do using pip:
$ pip install scrapy-doStart the daemon in the foreground:
$ scrapy-do -n scrapy-doOpen another terminal window, download the Scrapy’s Quotesbot example, and push the code to the server:
$ git clone https://github.com/scrapy/quotesbot.git $ cd quotesbot $ scrapy-do-cl push-project +----------------+ | quotesbot | |----------------| | toscrape-css | | toscrape-xpath | +----------------+Schedule some jobs:
$ scrapy-do-cl schedule-job --project quotesbot \ --spider toscrape-css --when 'every 5 to 15 minutes' +--------------------------------------+ | identifier | |--------------------------------------| | 0a3db618-d8e1-48dc-a557-4e8d705d599c | +--------------------------------------+ $ scrapy-do-cl schedule-job --project quotesbot --spider toscrape-css +--------------------------------------+ | identifier | |--------------------------------------| | b3a61347-92ef-4095-bb68-0702270a52b8 | +--------------------------------------+See what’s going on:
The web interface is available at http://localhost:7654 by default.
Building from source
Both of the steps below require nodejs to be installed.
Check if things work fine:
$ pip install -rrequirements-dev.txt $ toxBuild the wheel:
$ python setup.py bdist_wheel
ChangeLog
Version 0.5.0
Rewrite the log handling functionality to resolve duplication issues
Bump the JavaScript dependencies to resolve browser caching issues
Make the error message on failed spider listing more descriptive (Bug #28)
Make sure that the spider descriptions and payloads get handled properly on restart (Bug #24)
Clarify the documentation on passing arguments to spiders (Bugs #23 and #27)
Version 0.4.0
Migration to the Bootstrap 4 UI
Make it possible to add a short description to jobs
Make it possible to specify user-defined payload in each job that is passed on as a parameter to the python crawler
UI updates to support the above
New log viewers in the web UI
Release files for scrapy-do 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scrapy_do-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Release files / scrapy_do-0.5.0-py3-none-any.whl
| Download URL | scrapy_do-0.5.0-py3-none-any.whl |
|---|---|
| Size | 2.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ccccc7c14c3e38b0c6bd1f97128e4fc90120b3965248887dfdda5fbcc02f0750
|
|
BLAKE2b-256 checksum How to use checksums |
65e843811a916923f42ca802d1bdeac056ae33191d895d67d33a63270b44dd6d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/3.3.0 pkginfo/1.6.1 requests/2.25.1 setuptools/44.0.0 requests-toolbelt/0.9.1 tqdm/4.54.1 CPython/3.9.1
|