Skip to main content

An asyncio + aiolibs crawler imitate scrapy framework

Project description

aioscpy

Aioscpy

An asyncio + aiolibs crawler imitate scrapy framework

English | 中文

Overview

Aioscpy framework is base on opensource project Scrapy & scrapy_redis.

Aioscpy is a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages.

Dynamic variable injection is implemented and asynchronous coroutine feature support.

Distributed crawling/scraping.

Requirements

  • Python 3.7+
  • Works on Linux, Windows, macOS, BSD

Install

The quick way:

pip install aioscpy

Usage

create project spider:

aioscpy startproject project_quotes
cd project_quotes
aioscpy genspider quotes 

tree

quotes.py:

from aioscpy.spider import Spider


class QuotesSpider(Spider):
    name = 'quotes'
    custom_settings = {
        "SPIDER_IDLE": False
    }
    start_urls = [
        'https://quotes.toscrape.com/tag/humor/',
    ]

    async def parse(self, response):
        for quote in response.css('div.quote'):
            yield {
                'author': quote.xpath('span/small/text()').get(),
                'text': quote.css('span.text::text').get(),
            }

        next_page = response.css('li.next a::attr("href")').get()
        if next_page is not None:
            yield response.follow(next_page, self.parse)

create single script spider:

aioscpy onespider single_quotes

single_quotes.py:

from aioscpy.spider import Spider
from anti_header import Header
from pprint import pprint, pformat


class SingleQuotesSpider(Spider):
    name = 'single_quotes'
    custom_settings = {
        "SPIDER_IDLE": False
    }
    start_urls = [
        'https://quotes.toscrape.com/',
    ]

    async def process_request(self, request):
        request.headers = Header(url=request.url, platform='windows', connection=True).random
        return request

    async def process_response(self, request, response):
        if response.status in [404, 503]:
            return request
        return response

    async def process_exception(self, request, exc):
        raise exc

    async def parse(self, response):
        for quote in response.css('div.quote'):
            yield {
                'author': quote.xpath('span/small/text()').get(),
                'text': quote.css('span.text::text').get(),
            }

        next_page = response.css('li.next a::attr("href")').get()
        if next_page is not None:
            yield response.follow(next_page, callback=self.parse)

    async def process_item(self, item):
        self.logger.info("{item}", **{'item': pformat(item)})


if __name__ == '__main__':
    quotes = SingleQuotesSpider()
    quotes.start()

run the spider:

aioscpy crawl quotes
aioscpy runspider quotes.py

run

start.py:

from aioscpy import call_grace_instance
from aioscpy.utils.tools import get_project_settings


def load_file_to_execute():
    process = call_grace_instance("crawler_process", get_project_settings())
    process.load_spider(path='./spiders')
    process.start()


def load_name_to_execute():
    process = call_grace_instance("crawler_process", get_project_settings())
    process.crawl('[spider_name]')
    process.start()

more commands:

aioscpy -h

Ready

please submit your sugguestion to owner by issue

Thanks

aiohttp

scrapy

loguru

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

aioscpy-0.2.12.tar.gz (57.9 kB view details)

Uploaded Source

Built Distribution

aioscpy-0.2.12-py3-none-any.whl (79.9 kB view details)

Uploaded Python 3

File details

Details for the file aioscpy-0.2.12.tar.gz.

File metadata

  • Download URL: aioscpy-0.2.12.tar.gz
  • Upload date:
  • Size: 57.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.2 importlib_metadata/4.8.1 pkginfo/1.7.1 requests/2.27.1 requests-toolbelt/0.9.1 tqdm/4.62.2 CPython/3.9.6

File hashes

Hashes for aioscpy-0.2.12.tar.gz
Algorithm Hash digest
SHA256 166f8bfebfd83ba95853a0bd840b4351d5179c139aec4418e89846ca66c16f2f
MD5 5c79bfe0b5f94e2869f5a3d7a2a37064
BLAKE2b-256 4471e628b5ea49ff7d4b537def16e919ea76baa6ff5410e8617b0d520016c750

See more details on using hashes here.

File details

Details for the file aioscpy-0.2.12-py3-none-any.whl.

File metadata

  • Download URL: aioscpy-0.2.12-py3-none-any.whl
  • Upload date:
  • Size: 79.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.4.2 importlib_metadata/4.8.1 pkginfo/1.7.1 requests/2.27.1 requests-toolbelt/0.9.1 tqdm/4.62.2 CPython/3.9.6

File hashes

Hashes for aioscpy-0.2.12-py3-none-any.whl
Algorithm Hash digest
SHA256 0f46c45735c80f8e4979ea681f2521015945fc023aad18b105fa0b2eb2a7753b
MD5 c29da4b97b6d64728703aaf34afbd3db
BLAKE2b-256 bc65e8afa2102269285c5602e37c152d795622c11b38e2e10d5342e5f8b9ab88

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page