Skip to main content

ManualForge

Configuration-driven management manual generation framework. 配置驱动的管理手册生成框架。 Define your data sources, fields, and templates in YAML — get a formatted report. 在 YAML 中定义数据源、字段和模板,即可生成格式化报告。

PyPI version Python

Built on Kedro pipelines with Polars for data processing and Typst for document rendering. 基于 Kedro 流水线 + Polars 数据处理 + Typst 文档渲染。

ManualForge is a reusable Python package (pip install manualforge). It can be embedded into domain applications that need config-driven document generation — for example, regulatory manuals with rule-code field mapping and reconciliation reporting. ManualForge 是一个可复用的 Python 包pip install manualforge),可嵌入需要配置驱动文档生成的领域应用——例如带规则代码字段映射与核对报告的监管类手册。

What's New / 版本更新

v0.4.0(2026-09-03 · 正式发布)

  • Data quality engine / 数据质量引擎(manualforge.quality:声明式规则执行(count / unique / null_count / not_in / in_list / custom_func)+ register_custom 注册表 + JSON/MD/历史 CSV 报告;run_dq_on_frame 一键入口;领域检查函数由应用注册,引擎零领域依赖。
  • GB/T 9704 official report / 公文上报报告manualforge.quality.gbt9704 将 DQ 报告 dict 与历史轮次渲染为公文 PDF(redline=false、spacing-theme=relaxed、table-theme=three-line)。
  • Typst reporting base / Typst 报告基座(manualforge.reporting:转义(esc/esc_str)、值格式化(fmt_value/fmt_ts)与 PDF 编译统一实现,收敛跨应用重复代码。
  • 接入方可将 DQ 引擎与公文上报薄化为适配层;V1/V2/preflight/manifest 与公文 .typ 逐字节回归通过。

v0.3.2(2026-09-02)

  • Code-version manual rendering / 手册 code 版渲染:在 report 配置中开启 code_version.enabled 后,条目标题渲染为「名称(规则代码)」(如 财务信息隔离(KGBKQ_1)),并在 .typ 文件头部生成全量 #let warning_type_dict = (…) 代码→名称字典,便于程序化解析与对照;
  • Multi-report selection / 多报告段convert_rules_to_typst_jinja 新增 report_name 参数(默认 rules_manual,向后兼容),同一数据源可并行渲染 多个手册版本(如原版 / 人工认可版 / 全建议预览版);
  • Title normalization / 标题规范化:条目标题统一 trim 尾随空格;
  • Rule-code string-safe variant / 规则代码转义修复规则代码 增加字符串 安全版本(规则代码__str),修复 KGBKQ_1 在标题/字典中被 content 转义成 KGBKQ\_1 的问题。

说明:v0.3.3 为文档修订版(README 功能说明同步),渲染能力与 v0.3.2 一致。

Philosophy / 设计理念

ManualForge separates what you want to produce from how it's produced. ManualForge 将「要生成什么」与「如何生成」解耦。

  • What / 内容: Defined in conf/base/parameters_manualforge.yml — your data sources, expected columns, standardization rules, sort orders, summary dimensions, and report templates. 在配置文件中定义数据源、期望列、标准化规则、排序、汇总维度和报告模板。
  • How / 方法: Implemented by the pipeline nodes — reusable data processing functions that read from your config. 由流水线节点实现——可复用的数据处理函数,读取配置驱动行为。

To create a new manual for a different domain, you only need to edit the config file (and optionally provide new templates). No Python code changes required. 要为新领域创建手册,只需编辑配置文件(可选提供新模板),无需修改 Python 代码。

Features / 功能

Capability / 能力 Description / 说明
Multi-sheet Excel ingestion / 多表 Excel 读取 Auto-detect headers, filter cover sheets, merge into structured DataFrames. 自动检测表头,过滤封面页,合并为结构化 DataFrame。
Field standardization / 字段标准化 Mapping files + exact matching + fuzzy matching (difflib / duckdb). 映射文件 + 精确匹配 + 模糊匹配。
Config-driven summaries / 配置驱动汇总 Define group-by dimensions, sort orders, ability categories, and output paths in YAML. 在 YAML 中定义分组维度、排序、能力类别和输出路径。
Typst report generation / Typst 报告生成 Jinja2 templates → Typst source → PDF compilation. Jinja2 模板 → Typst 源码 → PDF 编译。
Pipeline hooks / 流水线钩子 Shell command hooks at pipeline/node granularity for pre/post processing. 流水线/节点粒度的 shell 命令钩子,用于前后处理。
Auto-backup / 自动备份 Pre-run config snapshot + post-run data backup via hooks. 跑前配置快照 + 跑后数据备份,通过 hooks 自动触发。
Config deploy / 配置部署 cfg-backup / cfg-deploy — backup, restore, and deploy configs from templates. 备份、恢复和从模板部署配置文件。
Code-version manual rendering / 手册 code 版渲染 Per-report code_version toggle: titles with rule codes + header warning_type_dict; multi-report rendering via report_name. 按报告开启代码版:标题带规则代码、文件头带代码→名称字典,支持多报告段并行。
Data quality engine / 数据质量引擎 Declarative YAML rules → Polars checks + custom-function registry (manualforge.quality). 声明式规则执行(count/unique/null/not_in/in_list/custom_func),领域函数由应用注册。
GB/T 9704 official report / 公文上报报告 Quality report rendered as official-style PDF via gbt9704-gongwen (manualforge.quality.gbt9704). DQ 报告 dict + 历史轮次 → 公文 PDF。
Typst reporting base / Typst 报告基座 Shared escaping/formatting/PDF compilation helpers (manualforge.reporting). 统一转义/值格式化/PDF 编译。

Quick Start / 快速开始

# 1. Install dependencies / 安装依赖
pip install -r requirements.txt

# 2. Copy and customize configuration / 复制并自定义配置
#    Option A: interactive deployment / 交互式部署
./scripts/cfg-deploy --from-examples

#    Option B: manual copy / 手动复制
cp conf/examples/parameters_manualforge.yml.example conf/base/parameters_manualforge.yml
cp conf/examples/catalog.yml.example          conf/base/catalog.yml
cp conf/examples/hooks.yml.example            conf/base/hooks.yml
cp conf/examples/parameters.yml.example       conf/base/parameters.yml
cp conf/examples/credentials.yml.example      conf/local/credentials.yml

# 3. Edit the config files to point to your data sources
#    编辑配置文件,指向你的数据源
#    (conf/base/ is gitignored — your real configs stay local)
#    (conf/base/ 已 gitignore — 实际配置保存在本地)

# 4. Run the pipeline / 运行流水线
kedro run

# Run specific node groups / 运行特定节点组
kedro run --tags conversion        # Excel → Parquet only / 仅 Excel → Parquet
kedro run --tags standardization   # Standardization only / 仅标准化
kedro run --tags csv               # Summary tables only / 仅汇总表

Backup & Config Management / 备份与配置管理

Auto-backup via hooks (runs on every kedro run): 通过 hooks 自动备份(每次 kedro run 自动触发):

kedro run
  ├─ [before_pipeline]  cfg-backup      ← snapshot conf/base/
  └─ [after_pipeline]   backup_data.sh  ← snapshot pipeline output data

Manual backup/restore/deploy: 手动备份/恢复/部署:

# Config backup / 配置备份
./scripts/cfg-backup              # backup conf/base/ → conf/.backups/
./scripts/cfg-backup -l           # list existing backups

# Config restore / deploy from examples / 配置恢复 / 从模板部署
./scripts/cfg-deploy                        # interactive menu | 交互菜单
./scripts/cfg-deploy -l                     # list config backups
./scripts/cfg-deploy -r 20260617_105645     # restore specific backup | 恢复指定备份
./scripts/cfg-deploy --from-examples         # deploy fresh templates | 从模板部署
./scripts/cfg-deploy --from-examples --dry-run  # preview | 预览

# Data backup / 数据备份
./scripts/backup_data.sh          # backup pipeline output → data/.backups/
./scripts/backup_data.sh -k 5     # keep only last 5 backups

Project Structure / 项目结构

├── conf/
│   ├── base/                          # ★ Gitignored — copy from examples/ | 从 examples/ 复制
│   │   ├── parameters_manualforge.yml # Central project configuration | 项目中心配置
│   │   ├── catalog.yml                # Kedro data catalog | 数据目录
│   │   ├── hooks.yml                  # Pipeline hooks (shell commands) | 流水线钩子
│   │   └── parameters.yml             # Pipeline parameters | 流水线参数
│   ├── examples/                      # ★ Tracked example templates | 版本追踪的示例模板
│   │   ├── parameters_manualforge.yml.example
│   │   ├── catalog.yml.example
│   │   ├── hooks.yml.example
│   │   ├── parameters.yml.example
│   │   └── credentials.yml.example
│   ├── local/                         # Local-only (gitignored) | 仅本地 (gitignored)
│   │   └── credentials.yml
│   └── logging.yml
├── data/                              # Gitignored except .gitkeep | 除 .gitkeep 外均 gitignored
│   ├── 01_raw/                        # Raw Excel/CSV + mapping files | 原始数据 + 映射文件
│   ├── 02_intermediate/              # Parquet, reconcile reports | 中间数据、核对报告
│   ├── 03_primary/                   # Standardized data | 标准化后数据
│   ├── 04_feature/                   # Summary tables (CSV + Markdown) | 汇总表
│   └── 08_reporting/                 # Typst sources & compiled PDFs | Typst 源码和 PDF
├── scripts/                          # Auxiliary scripts | 辅助脚本
│   ├── backup_data.sh                # ★ Backup pipeline output data | 备份管道输出数据
│   ├── cfg-backup                    # ★ Backup conf/base/ config | 备份配置文件
│   ├── cfg-deploy                    # ★ Deploy/restore configs | 部署/恢复配置
│   ├── main.sh                       # ★ Kedro runner with auto-backup | 带自动备份的启动脚本
│   ├── convert_csv_to_md.py          # CSV → Markdown conversion | 转换
│   ├── extract_rule_field_mapping.py # Rule field extraction | 规则字段提取
│   ├── extract_rule_overview.py      # Rule overview extraction | 规则概览提取
│   └── render_with_forge.py          # Markdown → DOCX/PDF rendering | 渲染
├── src/manualforge/                  # Framework source code | 框架源码
│   ├── config.py                     # Configuration helper utilities | 配置工具
│   ├── hooks.py                      # Kedro pipeline hooks (PipelineHooks base class) | 流水线钩子基类
│   ├── io/                           # Custom Kedro datasets (PolarsExcelDataset) | 自定义数据集
│   ├── quality/                       # Data quality engine + GB/T 9704 report | 数据质量引擎 + 公文报告(0.4.0)
│   ├── reporting/                     # Typst escaping/PDF base | Typst 转义/编译基座(0.4.0)
│   ├── pipelines/
│   │   └── data_processing_pl/       # Core pipeline: 12 reusable nodes | 核心流水线:12 个可复用节点
│   │       ├── nodes.py              #   Node functions | 节点函数
│   │       ├── pipeline.py           #   Pipeline definition | 流水线定义
│   │       ├── rulecsv2typ.py        #   CSV → Typst/Jinja conversion | CSV → Typst/Jinja 转换
│   │       └── standardize_fields.py #   Field standardization engine | 字段标准化引擎
│   ├── pipeline_registry.py          # Pipeline registration | 流水线注册
│   ├── settings.py                   # Kedro project settings | 项目设置
│   └── __main__.py                   # CLI entry point | CLI 入口
├── templates/                        # Jinja2 Typst templates | Jinja2 Typst 模板
│   └── manual.typ.j2
├── pyproject.toml                    # Project metadata & dependencies | 项目元数据和依赖
└── requirements.txt

Configuration Guide / 配置指南

The central configuration file is conf/base/parameters_manualforge.yml. Copy from conf/examples/ and customize. 核心配置文件为 conf/base/parameters_manualforge.yml。从 conf/examples/ 复制后进行自定义。

1. Data Sources / 数据源

Define your Excel files, expected headers, and sheet filtering rules. 定义 Excel 文件、期望表头和 Sheet 过滤规则:

datasources:
  primary_data:
    filepath: "data/01_raw/your_data.xlsx"
    sheet:
      exclude_names: ["封面", "封皮"]
      name_becomes_column: "sheet_name"
    header_detection:
      mode: keyword_match
      expected_headers:
        - "column_a"
        - "column_b"
    cleaning:
      drop_rows_where:
        column_a: ["column_a"]   # drop residual header rows | 删除残留表头行
      fill_null: forward
      deduplicate: true

2. Field Standardization / 字段标准化

Define which fields to standardize, their mapping files, and special corrections. 定义需要标准化的字段、映射文件和特殊修正:

standardization:
  fields:
    - name: "dept_name"
      mapping_file: "data/01_raw/dept_list"
      case_corrections:
        wrong_name: "correct_name"
      special_mappings:
        alias: "canonical_name"
      fuzzy:
        enabled: true
        threshold: 0.8
        method: difflib             # difflib | duckdb

3. Sort Orders / 排序

Define reusable sort order lists referenced by summaries. 定义汇总引用的可复用排序列表:

sort_orders:
  model_names:
    - "Model A"
    - "Model B"
  dep_names:
    - "HR"
    - "Finance"

4. Summaries / 汇总

Define what summary tables to generate. 定义要生成的汇总表:

summaries:
  my_summary:
    description: "Fields grouped by model and department"
    group_by: ["model", "department"]
    struct_columns: ["module", "system", "field_name"]
    sort_by:
      department: dep_names
    output:
      csv: "data/04_feature/my_summary.csv"

5. Reports / 报告

Define report templates and output. 定义报告模板和输出:

reports:
  my_report:
    description: "Rules manual"
    template_source: inline
    data_source: rules_data
    output_typ: "data/08_reporting/output.typ"
    typst_compile:
      enabled: true

Data Layers / 数据分层

Layer / 层级 Directory / 目录 Description / 说明
Raw / 原始 data/01_raw/ Source Excel/CSV files, mapping files / 源文件与映射文件
Intermediate / 中间 data/02_intermediate/ Parquet, reconcile reports / Parquet 与核对报告
Primary / 主数据 data/03_primary/ Standardized data / 标准化后数据
Feature / 特征 data/04_feature/ Summary tables (CSV + Markdown) / 汇总表
Reporting / 报告 data/08_reporting/ Typst sources & PDF output / Typst 源码与 PDF

Requirements / 环境要求

  • Python >= 3.10
  • Typst CLI (for PDF compilation / 用于 PDF 编译)

Recent Changes / 近期变更

2026-09-03 · v0.4.1(文档修订 / docs revision)

  • README 功能说明改纯功能向(去除对下游项目的引用),补充 v0.4.0 manualforge.quality / manualforge.reporting / 公文 PDF 能力介绍;同步 PyPI 描述。

2026-09-03 · v0.4.0(正式发布)

功能与 0.4.0a1 实验版一致;0.4.0a1 为同批预发布(PyPI 保留,仓库标签以 v0.4.0 为准)。

  • manualforge.quality:声明式数据质量引擎(规则执行 + register_custom 注册表 + JSON/MD/历史 CSV 写盘 + run_dq_on_frame 完整入口);领域检查函数由接入方注册,引擎零领域依赖。
  • manualforge.reporting:Typst 转义/值格式化/PDF 编译基座(esc/esc_str/fmt_value/fmt_ts/compile_typst),收敛跨应用重复实现。
  • manualforge.quality.gbt9704:GB/T 9704 公文 PDF 上报生成器(redline=false、spacing-theme=relaxed、table-theme=three-line),读取 DQ 报告 dict + 历史轮次渲染。
  • DQ 引擎/公文上报脚本可薄化为接入方适配层(sql/excel 回归通过)。

2026-09-02 · v0.3.4

  • 维护性发布:版本号 0.3.3 → 0.3.4(无功能变更,同步版本元数据与徽标)。

2026-09-02 · v0.3.3

  • 文档修订:README 补 v0.3.2 功能说明(code 版渲染、多报告段、标题 trim、规则代码转义修复),同步 PyPI 描述;版本号 0.3.2 → 0.3.3。

2026-09-02 · v0.3.2

  • Code-version manual renderingrulecsv2typ.py + templates/manual.typ.j2): convert_rules_to_typst_jinja 支持 report_name / code_version——条目标题 「名称(规则代码)」,头部条件渲染全量 #let warning_type_dict = (…)
  • 规则代码字符串安全版本规则代码__str 修复标题/字典中转义问题;
  • 标题 trim:条目标题统一去除尾随空格。

2026-06-29

  • Template extraction (rulecsv2typ.py): Extracted inline RECIPE_TEMPLATE_LEVEL2 into standalone templates/recipe.typ.j2; loaded via _load_recipe_template() using Path(__file__).parents[4]. Removed orphaned templates/report.typ.j2.

2026-06-26

  • Rule code parsing (nodes.py): Added _parse_rule_codes and _RULE_CODE_RE for GZW rule code extraction; rule-code column saved as List(String) without forward-fill; added "规则代码" to EXPECTED_HEADER in convert_excel_to_parquet_fj1 and process_attachment1_excel.
  • Model name standardization (nodes.py): Added _load_model_mapping() and _normalize_model_name() for fuzzy-matching model names against a reference table; 概览 sheet now uses its own 模型名称 column instead of the sheet name.
  • Department field (nodes.py): Added 主研部门 to fj2 EXPECTED_HEADER.
  • Path resolution (nodes.py, standardize_fields.py): Changed Path(__file__).parents[4] to Path.cwd() so paths resolve correctly when ManualForge is used as an installed package (e.g., from a downstream app).
  • Direct data access (rulecsv2typ.py): convert_rules_to_typst_jinja now reads from the Kedro catalog as a Polars DataFrame directly instead of round-tripping through CSV on disk.

Development / 开发

pip install -e ".[dev]"
ruff check src/
pytest

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

manualforge-0.4.1.tar.gz (62.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

manualforge-0.4.1-py3-none-any.whl (57.2 kB view details)

Uploaded Python 3

File details

Details for the file manualforge-0.4.1.tar.gz.

File metadata

  • Download URL: manualforge-0.4.1.tar.gz
  • Upload date:
  • Size: 62.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for manualforge-0.4.1.tar.gz
Algorithm Hash digest
SHA256 482d9dabff922b32604761e588f24f7efa4c2ac0d32c428a63e7bcbb5900d896
MD5 379c70581fed527e4b129c9f8336a5a1
BLAKE2b-256 d1c64646ef5d81ec5dc86f5974a19bf9a86edc9caaa94e946d63805eeee47bbf

See more details on using hashes here.

File details

Details for the file manualforge-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: manualforge-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 57.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for manualforge-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 c7c6b8fdfdf2504c80e2220c1bc138b7a7759fcde5d3dd112ecd0ee0987a32e4
MD5 ec62fa67f4da0953db099d8ee09a1e66
BLAKE2b-256 2f71ebe80750ef98a8ce509d7e71b09de3fd5f555d262a43cb7a2664e9db161f

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.4

2 files

0.3.3

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page