Crawler integration with INSPIRE-HEP using scrapy project HEPCrawl.
This module allows scheduling of crawler jobs to a Scrapyd instance serving a Scrapy project. E.g. in this case the default scrapy project is HEPCrawl.
It integrates directly with invenio-workflows module to create workflows for every record harvested by the crawler.
This module is meant to use only with INSPIRE-HEP overlay. Use at own risk.
Full documentation is hosted here: http://pythonhosted.org/inspire-crawler/
See also documentation of HEPCrawl: http://pythonhosted.org/hepcrawl/
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
inspire-crawler-3.0.4.tar.gz
(35.4 kB
view details)
File details
Details for the file inspire-crawler-3.0.4.tar.gz.
File metadata
- Download URL: inspire-crawler-3.0.4.tar.gz
- Upload date:
- Size: 35.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/1.13.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.1.0 requests-toolbelt/0.9.1 tqdm/4.33.0 CPython/2.7.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6ad11cb23666b7fdb4e5f179adae3cf73d2ff86ee052cc1f82c9ebf8f985f1f8
|
|
| MD5 |
92b6dca56a296f68cf69a6e2fe04a2b2
|
|
| BLAKE2b-256 |
9bb495946ca4f776b164935e0fc54ed3e7c7781876cfb486bbdcd2fd56d5fd16
|