Skip to main content

简体中文 English

Introduction

GPT-Web-Crawler is a web crawler based on python and puppeteer. It can crawl web pages and extract content (including WebPages' title,url,keywords,description,all text content,all images and screenshot) from web pages. It is very easy to use and can be used to crawl web pages and extract content from web pages in a few lines of code. It is very suitable for people who are not familiar with web crawling and want to use web crawling to extract content from web pages. Crawler Working

The output of the spider can be a json file, which can be easily converted to a csv file, imported into a database or building an AI agent. Assistant demo

Getting Started

Step1. Install the package.

pip install gpt-web-crawler

Step2. Copy config_template.py and rename it to config.py. Then, edit the config.py file to config the openai api key and other settings, if you need use ProSpider to help you extract content from web pages. If you don't need to use ai help you extract content from web pages, you can keep the config.py file unchanged.

Step3. Run the following code to start a spider.

from gpt_web_crawler import run_spider,NoobSpider
run_spider(NoobSpider, 
           max_page_count= 10 ,
           start_urls="https://www.jiecang.cn/", 
           output_file = "test_pakages.json",
           extract_rules= r'.*\.html' )

Spiders

Spider Type Description
NoobSpider Basic web page scraping
CatSpider Web page scraping with screenshots
ProSpider Web page scraping with AI-extracted content
LionSpider Web page scraping with all images extracted

Cat Spider

Cat spider is a spider that can take screenshots of web pages. It is based on the Noob spider and uses puppeteer to simulate browser operations to take screenshots of the entire web page and save it as an image. So when you use the Cat spider, you need to install puppeteer first.

npm install puppeteer

TODO

  • 支持无需配置config.py

Release files for gpt-web-crawler 0.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpt-web-crawler 0.0.2
File Size Uploaded
gpt-web-crawler-0.0.2.tar.gz 16.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpt-web-crawler 0.0.2
File Interpreter ABI Platform
gpt_web_crawler-0.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 37.8 kB

Release files / gpt-web-crawler-0.0.2.tar.gz

Download URL gpt-web-crawler-0.0.2.tar.gz
Size 16.1 kB
Tags Source
SHA-256 checksum
How to use checksums
5df8005ce68ee51a3b74aef8c6e39b08e1fdb329cee3d3b625d301c13767d41f
BLAKE2b-256 checksum
How to use checksums
08f25840ca1241368a1075e19e1eb4bb14d55d98815a1ebfeab079fcf3fac9dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release files / gpt_web_crawler-0.0.2-py3-none-any.whl

Download URL gpt_web_crawler-0.0.2-py3-none-any.whl
Size 21.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
668bb06722fe917ad74daec9cf21e394b8af7cc428d333d220d941a26fb5cc09
BLAKE2b-256 checksum
How to use checksums
65aa41830345bd326154fb994b8bce40d20dfa50bc83ece0f612fc259d7ac047
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release history Release notifications | RSS feed

This release

0.0.2 This release

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page