Skip to main content

Django application which crawls and downloads online content following instructions

Project description

django-scraper is a Django application which crawls and downloads online content following configurable instructions.

  • Extract content of given online websites/pages using XPath queries.
  • Automatically browse and download content in related pages, with given depth.
  • Support metadata extract along with other content
  • Have content refinement rules and black words filtering
  • Store and prevent duplication of downloaded content
  • Support HTTP, HTTPS proxies.

Documentation

The full documentation is not ready yet, please go here for notes about installation and usage: https://github.com/zniper/django-scraper

Support

If you have any questions or any ideas regarding this application, please email to me[at]zniper.net

Project details


Release history Release notifications

History Node

0.3.8

History Node

0.3.0

History Node

0.2.3

History Node

0.2.2

History Node

0.2.0

This version
History Node

0.1

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Filename, size & hash SHA256 hash help File type Python version Upload date
django-scraper-0.1.tar.gz (6.9 kB) Copy SHA256 hash SHA256 Source None Jul 4, 2014

Supported by

Elastic Elastic Search Pingdom Pingdom Monitoring Google Google BigQuery Sentry Sentry Error logging CloudAMQP CloudAMQP RabbitMQ AWS AWS Cloud computing Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page