Skip to main content

中文

reycrawler

reycrawler is a Python method integration package for web crawling.

It provides crawling methods for websites such as Baidu, Douban, Sina, Toutiao, and Weibo, and integrates crawling capabilities for different websites through a modular design.

It also provides browser-based human-like page crawling methods, which can simulate browser behavior to retrieve page HTML content and bypass anti-crawling mechanisms on some websites.

It is suitable for website data collection, information retrieval, and web crawling scenarios.

Features

  • Provides web crawling methods
  • Integrates crawling methods for multiple well-known websites
  • Supports crawling Baidu-related websites
  • Supports crawling Douban
  • Supports crawling Sina
  • Supports crawling Toutiao
  • Supports crawling Weibo
  • Provides crawling methods for other websites
  • Supports human-like page crawling through browsers
  • Supports retrieving page HTML content through browsers
  • Supports browser crawling tasks based on database polling
  • Provides unified method exports
  • Modular design, allowing crawling methods for different websites to be used as needed

Installation

Requires Python 3.12 or higher.

pip install reycrawler

Modules

reycrawler is divided into multiple modules by functionality and target website, with each module responsible for different website crawling and base functionality.

rall — All import methods

Unified export module.

Provides convenient exports of all reycrawler module methods. It allows the functionality provided by the package to be imported centrally, reducing the need to import from multiple modules separately.


rbaidu — Crawl Baidu website methods

Baidu website crawling module.

Provides data crawling methods for Baidu-related websites.

Mainly includes:

  • Baijiahao website data crawling
  • Baidu Translate website data crawling
  • Other Baidu-related website data crawling

rbase — Base methods

Base methods module.

Provides common base methods and shared dependencies used by other crawler modules.

It is mainly used to support the operation of other modules and provide common base functionality.


rbrowser — Browser methods

Browser crawling module.

Uses human-like browser control to crawl webpage HTML content, allowing anti-crawling mechanisms on some websites to be bypassed.

Browser crawling tasks are provided through database polling, allowing other programs to invoke browser crawling capabilities through database tasks.

Mainly provides:

  • Human-like browser control
  • Webpage HTML content retrieval
  • Browser crawling for anti-crawling scenarios
  • Crawling tasks based on database polling
  • Browser crawling integration methods

rdouban — Crawl Douban website methods

Douban website crawling module.

Provides data crawling methods for Douban.

Mainly includes:

  • Movie and TV ranking data crawling
  • Movie and TV details crawling
  • Other Douban website data crawling

rother — Crawl other website methods

Other website crawling module.

Provides data crawling methods for other websites.

Mainly includes:

  • Chinese lunar calendar date information crawling
  • Other website data crawling

rsina — Crawl Sina website methods

Sina website crawling module.

Provides data crawling methods for Sina.

Mainly includes:

  • Securities market search result crawling
  • Stock detail data crawling
  • Other Sina website data crawling

rtoutiao — Crawl Toutiao website methods

Toutiao website crawling module.

Provides data crawling methods for Toutiao.

Mainly includes:

  • Hot news data crawling
  • Other Toutiao website data crawling

rweibo — Crawl Weibo website methods

Weibo website crawling module.

Provides data crawling methods for Weibo.

Mainly includes:

  • Hot news data crawling
  • Other Weibo website data crawling

Module Overview

Module Function
rall Unified export of all methods
rbase Base methods and shared dependencies
rbrowser Human-like browser control and webpage crawling
rbaidu Baidu-related website crawling
rdouban Douban website crawling
rother Other website crawling
rsina Sina website crawling
rtoutiao Toutiao website crawling
rweibo Weibo website crawling

Dependencies

Main dependencies:

  • bs4
  • fake_useragent
  • reydb
  • reykit
  • selenium

Project Information

Project Information
Name reycrawler
Version 1.0.18
Python >=3.12
Author Rey
Email reyxbo@163.com
Homepage REYXBO
Repository reycrawler-py

Keywords

rey · reyxbo · crawler · crawl-web · browser · baidu · douban · sina · toutiao · weibo

Release files for reycrawler 1.0.26

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for reycrawler 1.0.26
File Interpreter ABI Platform
reycrawler-1.0.26-py3-none-any.whl Python 3 none any Details

Release files / reycrawler-1.0.26-py3-none-any.whl

Download URL reycrawler-1.0.26-py3-none-any.whl
Size 20.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b153dd3afe5b0f0d4b39d337da0f738382f7c90fc1c37d880322436aa84bee34
BLAKE2b-256 checksum
How to use checksums
be3ed9dd95f601f18acec9331d6cfd9c671a4e454ba02e9d9042852ccaa109fb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.12

Release history Release notifications | RSS feed

This release

1.0.26 This release

1 release file

1.0.25

1 release file

1.0.18

1 release file

1.0.17

1 release file

1.0.16

1 release file

1.0.15

1 release file

1.0.14

1 release file

1.0.13

1 release file

1.0.12

1 release file

1.0.11

2 release files

1.0.10

1 release file

1.0.9

2 release files

1.0.8

1 release file

1.0.7

2 release files

1.0.6

1 release file

1.0.5

1 release file

1.0.4

1 release file

1.0.3

1 release file

1.0.2

1 release file

1.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page