Skip to main content

中文

reycrawler

reycrawler is a Python method integration package for web crawling.

It provides crawling methods for websites such as Baidu, Douban, Sina, Toutiao, and Weibo, and integrates crawling capabilities for different websites through a modular design.

It also provides browser-based human-like page crawling methods, which can simulate browser behavior to retrieve page HTML content and bypass anti-crawling mechanisms on some websites.

It is suitable for website data collection, information retrieval, and web crawling scenarios.

Features

  • Provides web crawling methods
  • Integrates crawling methods for multiple well-known websites
  • Supports crawling Baidu-related websites
  • Supports crawling Douban
  • Supports crawling Sina
  • Supports crawling Toutiao
  • Supports crawling Weibo
  • Provides crawling methods for other websites
  • Supports human-like page crawling through browsers
  • Supports retrieving page HTML content through browsers
  • Supports browser crawling tasks based on database polling
  • Provides unified method exports
  • Modular design, allowing crawling methods for different websites to be used as needed

Installation

Requires Python 3.12 or higher.

pip install reycrawler

Modules

reycrawler is divided into multiple modules by functionality and target website, with each module responsible for different website crawling and base functionality.

rall — All import methods

Unified export module.

Provides convenient exports of all reycrawler module methods. It allows the functionality provided by the package to be imported centrally, reducing the need to import from multiple modules separately.


rbaidu — Crawl Baidu website methods

Baidu website crawling module.

Provides data crawling methods for Baidu-related websites.

Mainly includes:

  • Baijiahao website data crawling
  • Baidu Translate website data crawling
  • Other Baidu-related website data crawling

rbase — Base methods

Base methods module.

Provides common base methods and shared dependencies used by other crawler modules.

It is mainly used to support the operation of other modules and provide common base functionality.


rbrowser — Browser methods

Browser crawling module.

Uses human-like browser control to crawl webpage HTML content, allowing anti-crawling mechanisms on some websites to be bypassed.

Browser crawling tasks are provided through database polling, allowing other programs to invoke browser crawling capabilities through database tasks.

Mainly provides:

  • Human-like browser control
  • Webpage HTML content retrieval
  • Browser crawling for anti-crawling scenarios
  • Crawling tasks based on database polling
  • Browser crawling integration methods

rdouban — Crawl Douban website methods

Douban website crawling module.

Provides data crawling methods for Douban.

Mainly includes:

  • Movie and TV ranking data crawling
  • Movie and TV details crawling
  • Other Douban website data crawling

rother — Crawl other website methods

Other website crawling module.

Provides data crawling methods for other websites.

Mainly includes:

  • Chinese lunar calendar date information crawling
  • Other website data crawling

rsina — Crawl Sina website methods

Sina website crawling module.

Provides data crawling methods for Sina.

Mainly includes:

  • Securities market search result crawling
  • Stock detail data crawling
  • Other Sina website data crawling

rtoutiao — Crawl Toutiao website methods

Toutiao website crawling module.

Provides data crawling methods for Toutiao.

Mainly includes:

  • Hot news data crawling
  • Other Toutiao website data crawling

rweibo — Crawl Weibo website methods

Weibo website crawling module.

Provides data crawling methods for Weibo.

Mainly includes:

  • Hot news data crawling
  • Other Weibo website data crawling

Module Overview

Module Function
rall Unified export of all methods
rbase Base methods and shared dependencies
rbrowser Human-like browser control and webpage crawling
rbaidu Baidu-related website crawling
rdouban Douban website crawling
rother Other website crawling
rsina Sina website crawling
rtoutiao Toutiao website crawling
rweibo Weibo website crawling

Dependencies

Main dependencies:

  • bs4
  • fake_useragent
  • reydb
  • reykit
  • selenium

Project Information

Project Information
Name reycrawler
Version 1.0.18
Python >=3.12
Author Rey
Email reyxbo@163.com
Homepage REYXBO
Repository reycrawler-py

Keywords

rey · reyxbo · crawler · crawl-web · browser · baidu · douban · sina · toutiao · weibo

Release files for reycrawler 1.0.20

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reycrawler 1.0.20
File Size Uploaded
reycrawler-1.0.20.tar.gz 154.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for reycrawler 1.0.20
File Interpreter ABI Platform
reycrawler-1.0.20-py3-none-any.whl Python 3 none any Details

Total release size: 174.6 kB

Release files / reycrawler-1.0.20.tar.gz

Download URL reycrawler-1.0.20.tar.gz
Size 154.4 kB
Tags Source
SHA-256 checksum
How to use checksums
bd3c561b0397aa0a9fb22e1f4231f78c0e62b99c2b828abc980c1c31b763285f
BLAKE2b-256 checksum
How to use checksums
55a16b3ce9281cef17097e344a24df81c265f0a3dcdddb29d19108c89eda0c85
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.12

Release files / reycrawler-1.0.20-py3-none-any.whl

Download URL reycrawler-1.0.20-py3-none-any.whl
Size 20.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8b186927e1a4920d1ed44b1c69c37d5ee9b1a8cf431bd1f2a6c33d3793a6a626
BLAKE2b-256 checksum
How to use checksums
c4facb5e672f07d62edfc25f2fe6e61c6ee6c5b044e6a45090beb164f56d7f83
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.12

Release history Release notifications | RSS feed

1.0.26

1 release file

1.0.25

1 release file

This release

1.0.20 This release

2 release files

1.0.18

1 release file

1.0.17

1 release file

1.0.16

1 release file

1.0.15

1 release file

1.0.14

1 release file

1.0.13

1 release file

1.0.12

1 release file

1.0.11

2 release files

1.0.10

1 release file

1.0.9

2 release files

1.0.8

1 release file

1.0.7

2 release files

1.0.6

1 release file

1.0.5

1 release file

1.0.4

1 release file

1.0.3

1 release file

1.0.2

1 release file

1.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page