jaws-scraper

Just Another Web Scraper.

Project description

# Introduction JAWS is a system for quickly designing web scrapers. It contains a framework for designing custom resources, parsers and outputs for entirely custom scrapers as well as a few implemenations for common use cases.

# Dependenices JAWS is written in Python, for Python2. The dependencies for the latest version are: * mechanize==0.2.5 * requests==2.2.1

JAWS can also be installed with easy_install or pip.

# Components The core components of the JAWS framework can be found in core.py.

## Scraper The Scraper class is a collection of all the core components into one object which can be easily instantiated and used to scrape all data into your specified output.

## Resource The JAWSResource class is the abstract class describing the interface by which pages are provided to the parser for scraping. A resource could be as simple as a file reader or as complex as a full Web crawler.

## Parser The JAWSParser class is the abstract class describing the way your scraper will turn input from the resource into a python dictionary of keys and values to be fed to the output.

## Output The JAWSOutput class is the abstract class describing what to actually do with that data you have scraped. It could describe a file output format (a csv is probably simplest), a database interface, or whatever else you can think of.

# Future Work * Automatic Schema Detection * JSON parser * Examples for README * Better documentation in code * Python3 compatibility

# License All code and content distributed with JAWS is released under the [GNU GPLv3](http://www.gnu.org/licenses/gpl-3.0.html) unless otherwise specified or prohibited.

Project details

Release history Release notifications | RSS feed

This version

0.1.0

Mar 14, 2014

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

jaws-scraper-0.1.0.tar.gz (16.3 kB view details)

Uploaded Mar 14, 2014 Source

File details

Details for the file jaws-scraper-0.1.0.tar.gz.

File metadata

Download URL: jaws-scraper-0.1.0.tar.gz
Upload date: Mar 14, 2014
Size: 16.3 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for jaws-scraper-0.1.0.tar.gz
Algorithm	Hash digest
SHA256	`6c3814f838c0586131f0c5748d7ea2b20b324183d47dfc8832ae986d7748db05`
MD5	`99eda24ccc889d2dad41aa47c4dd547c`
BLAKE2b-256	`e8120642fade06bdd35f6f31de21d95be6970477752b8684e359a70b712e7cc5`

See more details on using hashes here.

jaws-scraper 0.1.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta