Skip to main content

Chinese scraping utilities — date parsing, city extraction, SHA256 ID, UA pool, rate limiter, DeepSeek client, web search, hot topic scrapers, LLM extraction pipeline

Project description

Language: English | 简体中文 | 繁體中文 | 日本語

chinese-scraper-utils

PyPI version Python 3.11+ MIT License

Shared Python utilities for Chinese-language web scraping — date parsing, city extraction, stable ID generation, UA rotation, rate limiting, web search, hot topic scrapers, LLM-powered event extraction, and a DeepSeek API client.

Extracted from ComiRadar and weekly-hotspot.


Installation

# Core (zero extra dependencies beyond httpx + openai)
pip install chinese-scraper-utils

# With web search support
pip install chinese-scraper-utils[search]

Quick Start

from chinese_scraper_utils import (
    # Core utilities
    parse_date, extract_date, extract_city, guess_category, stable_id, random_ua,
    # Web search
    search_web,
    # Hot topic scrapers
    scrape_weibo_hot, scrape_zhihu_hot,
    # LLM extraction
    DeepSeekClient, EventExtractor,
)

# Parse Chinese dates
parse_date("2026/05/20")         # "2026-05-20"
extract_date("5月4日上海有漫展")   # "2026-05-04"

# Extract cities (false-positive protected)
extract_city("活动在上海举办")     # "上海"
extract_city("西安路有个活动")     # "" (not a city!)

# Scrape hot topics
weibo = scrape_weibo_hot()       # list[HotTopic]
zhihu = scrape_zhihu_hot()       # list[HotTopic]

# LLM-powered event extraction
client = DeepSeekClient(api_key="sk-xxx")
extractor = EventExtractor(
    client=client,
    event_types=["漫展", "同人展", "演唱会"],
    min_confidence=0.5,
)
events = extractor.extract(["五一北京漫展嘉年华在国家会议中心..."])

# CLI usage
# python -m chinese_scraper_utils search "五一漫展"
# python -m chinese_scraper_utils scrape-weibo
# python -m chinese_scraper_utils extract posts.json -t "漫展,演唱会"

API Reference

Export Type Description
parse_date(s) str → str Structured date parsing
try_parse_date(s) str → str|None Same, returns None on failure
extract_date(text) str → str Chinese text date extraction
CITIES list[str] 50 major Chinese cities
extract_city(text, extra_cities=None) str → str City name extraction (false-positive safe)
normalize_city(city) str → str City name normalization
CATEGORY_ALIASES dict[str,str] Category alias mapping
guess_category(title) str → str Event category guessing (longest-match)
UA_POOL list[str] 21 modern User-Agent strings
random_ua() → str Random UA selection
stable_id(*parts) str → str Deterministic SHA256 short ID
RateLimiter class Async rate limiter with retry + jitter
DeepSeekClient class DeepSeek API client (sync/async, retry)
NEW SearchResult dataclass Web search result (title/url/snippet)
NEW search_web(query, n) → list[SearchResult] DuckDuckGo web search
NEW HotTopic dataclass Unified hot topic (title/summary/url/source)
NEW scrape_weibo_hot() → list[HotTopic] Weibo hot search
NEW scrape_zhihu_hot() → list[HotTopic] Zhihu hot list
NEW scrape_hackernews_top() → list[HotTopic] HN top stories
NEW ExtractedEvent dataclass Structured event (with source tracing)
NEW EventExtractor class 5-stage LLM extraction pipeline + cache
NEW extract_events(texts, ...) → list[ExtractedEvent] Convenience extractor
ScraperError / RateLimitError / etc. Exception classes Typed error hierarchy

Full API docs: API_REFERENCE.md


CLI

python -m chinese_scraper_utils search "五一北京漫展" -n 10
python -m chinese_scraper_utils scrape-weibo
python -m chinese_scraper_utils scrape-zhihu
python -m chinese_scraper_utils scrape-hn
python -m chinese_scraper_utils extract posts.json -t "漫展,演唱会" -c 0.5 -v

Related

License

MIT


Language / 语言

English | 简体中文

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

chinese_scraper_utils-0.3.0.tar.gz (39.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

chinese_scraper_utils-0.3.0-py3-none-any.whl (30.0 kB view details)

Uploaded Python 3

File details

Details for the file chinese_scraper_utils-0.3.0.tar.gz.

File metadata

  • Download URL: chinese_scraper_utils-0.3.0.tar.gz
  • Upload date:
  • Size: 39.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for chinese_scraper_utils-0.3.0.tar.gz
Algorithm Hash digest
SHA256 d9e5d25b933cfd82f7a5e89bcb636e2398d5165a96ee7c7cd34fae1b7be4e6cd
MD5 c4d15c2ed38ec3b7f2d628047ba35209
BLAKE2b-256 6a0cfd78a1babb74ed6f8094a3489c100066283f67a79d59c5403d217d4de287

See more details on using hashes here.

File details

Details for the file chinese_scraper_utils-0.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for chinese_scraper_utils-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 11e0ae378b1c17e3d8413d14b172b0e4cba20cf1183b0492dcd58014d73188c9
MD5 d1f8bf00c17d8430c00dfba7d2c4291a
BLAKE2b-256 2071755d993aa6708b1137683542458abd8674be60d29399e26c2925b94e0382

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page