Web scraping framework based on py3 asyncio & aiohttp libraries.
Usage Example
import re
from itertools import islice
from iob import Crawler, Request
RE_TITLE = re.compile(r'<title>([^<]+)</title>', re.S | re.I)
class TestCrawler(Crawler):
def task_generator(self):
for host in islice(open('var/domains.txt'), 100):
host = host.strip()
if host:
yield Request('http://%s/' % host, tag='page')
def handler_page(self, req, res):
print('Result of request to {}'.format(req.url))
try:
title = RE_TITLE.search(res.body).group(1)
except AttributeError:
title = 'N/A'
print('Title: {}'.format(title))
bot = TestCrawler(concurrency=10)
bot.run()
Installation
pip install iob
Dependencies
Python>=3.4
aiohttp
Metadata
Release files for iob 0.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| iob-0.0.2.tar.gz | 6.7 kB | Details |
Release files / iob-0.0.2.tar.gz
| Download URL | iob-0.0.2.tar.gz |
|---|---|
| Size | 6.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a86440277b1d91996c1f6c1c33a1219790962a504350adfc1f14771936bc3a77
|
|
BLAKE2b-256 checksum How to use checksums |
d69b4d2dadb823bdd27fe2ca45f19e33dce4d3aac3d02efa3a6087bd561bc145
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |