Skip to main content

fd-cn-gov

PyPI version Python versions License: MIT

Scrapers + a self-contained datasource registry for Chinese central-government ministry open-information archives. Catalog-crawls the public notice / news / data archives of 11 ministries (MOF, PBC, NDRC, MOFCOM, MOHURD, MOT, MOA, SAFE, MNR, MEE, MEM), emitting one JSON record per listed document, and ships a SQLite + JSON registry describing every datasource and its column schema.

Standalone: no monorepo, no MCP server, no daas.db dependency. The MANIFEST at the top of each scraper is the single source of truth; the bundled registry.db / registry.json are derived artifacts.

Ministries

Name Label Seed URL
mee_gsgg_archive MEE Notice Archive (生态环境部公示公告) https://www.mee.gov.cn/ywdt/gsgg/
mem_tzgg_archive MEM Notice Archive (应急管理部通知公告) https://www.mem.gov.cn/gk/tzgg/
mnr_tzgg_archive MNR Notice Archive (自然资源部通知公告) https://www.mnr.gov.cn/gk/tzgg/
moa_govpublic_archive MOA GovPublic Archive (农业农村部 机构分类) https://www.moa.gov.cn/govpublic/
mof_gkml_archive MOF gkml Archive (财政部信息公开) https://www.mof.gov.cn/gkml/
mofcom_xwfb_archive MOFCOM News Preview (商务部新闻发布) https://www.mofcom.gov.cn/xwfb/index.html
mohurd_xinwen_archive MOHURD Xinwen Archive (住建部新闻动态) https://www.mohurd.gov.cn/xinwen/
mot_shuju_archive MOT Data Hub Archive (交通运输部数据) https://www.mot.gov.cn/shuju/index.html
ndrc_tzgg_archive NDRC Notice Archive (发改委通知公告) https://www.ndrc.gov.cn/xwdt/tzgg/
pbc_xinwen_archive PBC News Archive (人民银行新闻发布) https://www.pbc.gov.cn/goutongjiaoliu/113456/113469/index.html
safe_whxw_archive SAFE News Archive (外汇局外汇新闻) https://www.safe.gov.cn/safe/whxw/index.html

Each scraper emits a JSON array of records to stdout with at least title, date (YYYY-MM-DD, from the URL t<YYYYMMDD>_ token with a <span> fallback), and url (absolute, the primary key). Some add section / subsection / doc_type / department. See fd-cn-gov describe <name> for the exact columns of any source.

Install

pip install fd-cn-gov

Requires Python ≥3.10. Dependencies: scrapling (HTTP + adaptive parsing), sqlalchemy.

From source

pip install git+https://github.com/FindDataOfficial/cn-goverment-datasource.git

CLI

# List the 11 registered sources
fd-cn-gov list

# Show one source's identity + full column schema
fd-cn-gov describe mof_gkml_archive

# Crawl one archive (default: 50 pages per sub-archive; prints JSON to stdout)
fd-cn-gov crawl mof_gkml_archive --max-pages 2 > records.json

# Full crawl, no page cap (use sparingly — see Polite crawling below)
fd-cn-gov crawl mof_gkml_archive --all > records.json

# Regenerate the bundled registry from the scripts' MANIFESTs
fd-cn-gov build-registry

Python API

import fd_cn_gov

# Discover datasources
for s in fd_cn_gov.list_sources():
    print(s.name, s.label, s.url)

# One source + its column schema
src = fd_cn_gov.get_source("mof_gkml_archive")
cols = fd_cn_gov.get_columns("mof_gkml_archive")  # -> list[Column]
#   Column(name='url', type='string', primary_key=True, nullable=False,
#          description='absolute document URL (.htm/.html/.pdf)',
#          source_field='a@href', semantic_type='url', ...)

Registry schema

The bundled fd_cn_gov/registry/registry.db is a 3-table SQLite, mirroring the column shapes of the originating daas.db for these sources (no foreign keys, no stale-FK footgun):

  • sources — one row per ministry: id, name, label, url, description, category, category_label, config_json
  • datasource_columns — one row per output column: datasource_id, table_name, column_name, column_type, is_primary_key, is_nullable, description, source_field, unit, semantic_type
  • scraw_configs — one row per scraper: name, url, columns_json

registry.json is a deterministic full dump of the same three tables (sort_keys), so the registry is diff-friendly in code review.

To regenerate after editing a MANIFEST:

fd-cn-gov build-registry

The build is logically idempotent: re-running produces an identical .dump and a byte-identical registry.json. (Only SQLite's file-change-counter header byte differs on each write — unavoidable on any SQLite write; compare via .dump or registry.json for equality.)

Polite crawling

These are .gov.cn hosts. Each scraper paces requests (SLEEP ≈ 0.3s between pages) and defaults to a 50-page cap per sub-archive. Use --max-pages N to bound a crawl further; reserve --all for when you genuinely need full history and can afford the time. Do not run multiple --all crawls in parallel against the same ministry.

Project layout

fd-cn-gov/
├── pyproject.toml
├── README.md
├── LICENSE
└── fd_cn_gov/
    ├── __init__.py            # public read API: list_sources / get_source / get_columns
    ├── cli.py                 # fd-cn-gov CLI
    ├── scraw_contract.py      # ScrawManifest / ScrawColumn dataclasses (vendored, trimmed)
    ├── build_registry.py      # regenerates registry.db + registry.json from MANIFESTs
    ├── scripts/               # 11 ministry scrapers, each with a module-level MANIFEST
    │   ├── mof_gkml_archive.py
    │   └── ...
    └── registry/
        ├── __init__.py        # read API over the bundled DB
        ├── registry.db        # generated — checked in
        └── registry.json      # generated — checked in

License

MIT

Metadata

Release files for fd-cn-gov 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fd-cn-gov 0.1.2
File Size Uploaded
fd_cn_gov-0.1.2.tar.gz 42.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fd-cn-gov 0.1.2
File Interpreter ABI Platform
fd_cn_gov-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 103.6 kB

Release files / fd_cn_gov-0.1.2.tar.gz

Download URL fd_cn_gov-0.1.2.tar.gz
Size 42.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8c5764f0e6c2f455ed8b7a679804538931ee8335e3899e15faf03dcadff1d6f8
BLAKE2b-256 checksum
How to use checksums
74f13d695f5510112d68769554608accc6c6e02f7ee0befd5aa3f1418b6f42d5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / fd_cn_gov-0.1.2-py3-none-any.whl

Download URL fd_cn_gov-0.1.2-py3-none-any.whl
Size 61.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9cdd217979f54e6f6f55a0782608c016489d85e141fc2e76654b6ebcd8005aa7
BLAKE2b-256 checksum
How to use checksums
b0a12ba9e9b1f57c9a09b2a979501701c1c6cff63561d555f390dfa825304364
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.1.3

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page