Fast asynchronous web scraper with minimalist API.
Project description
Hypotonic
Fast asynchronous web scraper with minimalist API inspired by awesome node-osmosis.
Hypotonic provides SQLAlchemy-like command chaining DSL to define HTML scrapers. Everything is executed asynchronously via asyncio
and all dependencies are pure Python. Supports querying by CSS selectors with Scrapy's pseudo-attributes. XPath is not supported due to libxml
requirement.
Hypotonic does not natively execute JavaScript on websites and it is recommended to use prerender.
Installing
Hypotonic requires Python 3.6+.
pip install hypotonic
Example
from hypotonic import Hypotonic
data, errors = (
Hypotonic()
.get('http://books.toscrape.com/')
.paginate('.next a::attr(href)', 5)
.find('.product_pod h3')
.set('title')
.follow('a::attr(href)')
.set({'price': '.price_color',
'availability': 'p.availability'})
.data()
)
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
hypotonic-0.0.14.tar.gz
(5.6 kB
view hashes)
Built Distribution
Close
Hashes for hypotonic-0.0.14-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | da6b303fb4e8d920d26cb22bb81dcc4c4ee3b860edd9e3e0d0517632f5486269 |
|
MD5 | 53cb85bb9741fefc08e9e5a089bca877 |
|
BLAKE2b-256 | 5c2bf7082ebfb61bc23a4c361f71d2cf08877f7af6f131c6ced0aa1c3d862564 |