Crawler integration with INSPIRE-HEP using scrapy project HEPCrawl.
This module allows scheduling of crawler jobs to a Scrapyd instance serving a Scrapy project. E.g. in this case the default scrapy project is HEPCrawl.
It integrates directly with invenio-workflows module to create workflows for every record harvested by the crawler.
This module is meant to use only with INSPIRE-HEP overlay. Use at own risk.
Full documentation is hosted here: http://pythonhosted.org/inspire-crawler/
See also documentation of HEPCrawl: http://pythonhosted.org/hepcrawl/
Release files for inspire-crawler 3.0.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| inspire-crawler-3.0.4.tar.gz | 35.4 kB | Details |
Release files / inspire-crawler-3.0.4.tar.gz
| Download URL | inspire-crawler-3.0.4.tar.gz |
|---|---|
| Size | 35.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6ad11cb23666b7fdb4e5f179adae3cf73d2ff86ee052cc1f82c9ebf8f985f1f8
|
|
BLAKE2b-256 checksum How to use checksums |
9bb495946ca4f776b164935e0fc54ed3e7c7781876cfb486bbdcd2fd56d5fd16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/1.13.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.1.0 requests-toolbelt/0.9.1 tqdm/4.33.0 CPython/2.7.15
|