This release is a pre-release and may not be stable for production use.
🌊 FlexAligner
Robust Speech-Text Alignment from Signal to Symbol
A Neural-Based Forced Alignment Framework for "Wild" Real-World Data
面向真实非受控数据的深度学习强鲁棒性对齐工具
📖 Introduction
FlexAligner is a robust speech-text alignment framework built upon wav2vec 2.0. It is designed for real-world linguistic data, where audio signals and textual transcriptions may contain noise, hesitations, untranscribed events, or local mismatches.
FlexAligner decomposes forced alignment into two stages:
- Macro-Segmentation (CTC Chunking): Uses a CTC acoustic model to locate reliable transcript anchors and divide long-form audio into ordered chunks.
- Micro-Alignment (Local Alignment): Uses a constrained pronunciation graph and two-pass Viterbi decoding to estimate word and phone boundaries within each chunk.
🌟 Key Features
- 🛡️ Tolerance to Mismatch: Uncovered portions of the audio timeline are
represented explicitly as
NULLintervals instead of being silently forced into neighboring words or phones. - 🎯 Word and Phone Boundaries: Produces Praat TextGrid files with continuous
wordsandphonestiers while preserving the input word order. - 🔒 Local and Reproducible: Models and pronunciation dictionaries are supplied as explicit local paths; alignment does not automatically download models or silently generate OOV pronunciations.
- 📦 Python Package and CLI: Provides a typed Python API and a command-line interface for single-file alignment.
The first public preview focuses on English, CPU, and single-file alignment. Mandarin, GPU, batch processing, Web services, automatic model download, multi-format decoding, resampling, default G2P, and confidence calibration are reserved interfaces and are not yet production features.
🌏 简介
FlexAligner 是一个基于 wav2vec 2.0 的语音—文本对齐框架,面向真实语言材料中 常见的噪音、停顿、未转写声音事件以及音频与文本局部不一致等问题。
FlexAligner 将强制对齐分为两个阶段:
- 宏观切分(CTC Chunking): 使用 CTC 声学模型寻找可靠的文本锚点,并将长音频 划分为顺序一致的局部片段。
- 微观对齐(Local Alignment): 在每个片段内构建受约束的发音图,通过两遍 Viterbi 解码估计词和音素边界。
🌟 核心优势
- 🛡️ 容错设计: 对未被词或音素覆盖的时段使用明确的
NULL区间表示,避免将其 静默挤压到相邻标签中。 - 🎯 词与音素边界: 输出 Praat TextGrid,
words和phones两层连续覆盖完整 音频时间轴,同时保持输入词序。 - 🔒 本地与可复现: 模型和发音词典均由用户显式指定;运行时不会自动下载模型, 也不会对 OOV 词静默生成发音。
- 📦 Python 包与 CLI: 提供带类型定义的 Python API 和单文件命令行接口。
首个公开预览版聚焦于英语、CPU、单文件对齐。普通话、GPU、批处理、Web 服务、 自动模型下载、多格式解码、自动重采样、默认 G2P 和置信度校准目前仅保留接口,尚未 作为正式能力开放。
🏗️ Architecture
graph TD
Input[Input: PCM16 WAV + Transcript] --> Lexicon[Local Pronunciation Lexicon];
Lexicon --> B[Stage 1: CTC Chunking];
B --> C{Reliable Ordered Chunks};
C --> D[Stage 2: Pronunciation Graph];
D --> E[Two-pass Viterbi Decoding];
E --> F[Words + Phones + NULL TextGrid];
🚀 Installation
The import-safe core package targets Python 3.10–3.14. The frozen inference extra pins Torch 2.3.1 and Transformers 4.41.2 and is installable on Python 3.10–3.12. Real-model release evidence currently covers only Linux x86_64 with Python 3.10.8; Python 3.13–3.14 are core-only.
After the public alpha is available from PyPI, install the exact preview with:
python -m pip install "flexaligner[inference]==0.1.0a1"
For a CPU-only Linux environment, install the frozen Torch build from its CPU index first:
python -m pip install torch==2.3.1 --index-url https://download.pytorch.org/whl/cpu
python -m pip install "flexaligner[inference]==0.1.0a1"
To work from a reviewed source checkout:
git clone https://github.com/USTCPhonetics/FlexAligner.git
cd FlexAligner
python -m pip install -e ".[inference]"
The package does not include acoustic models. Prepare compatible local Chunker and Aligner model directories and a pronunciation dictionary before alignment.
基础包支持 Python 3.10–3.14。首个 alpha 的推理依赖固定为 Torch 2.3.1 和 Transformers 4.41.2,仅支持在 Python 3.10–3.12 上安装;真实模型发布验证目前仅覆盖 Linux x86_64 与 Python 3.10.8。Python 3.13–3.14 暂时只承诺基础包和接口可用。 wheel 不包含声学模型,也不会自动下载模型;运行前必须准备本地 Chunker、Aligner 模型目录和发音词典。
💻 Usage
1. Command Line Interface (CLI)
Align one English 16 kHz mono PCM16 WAV file:
flexaligner align \
--audio recording.wav \
--text-file transcript.txt \
--lexicon english.dict \
--chunker-model /local/models/en/chunker \
--aligner-model /local/models/en/aligner \
--output recording.TextGrid \
--chunk-metadata recording.alignment.json \
--num-threads 1
Literal transcript text may be passed with --text instead of --text-file.
--num-threads configures Torch's process-global CPU thread count for the
inference lifetime; it is not isolated to a single aligner instance.
Use the capability command to inspect the installed preview:
flexaligner capabilities
flexaligner capabilities --json
2. Python API
from pathlib import Path
from flexaligner import (
AlignmentRequest,
FlexAligner,
LocalModelBundle,
TextGridOutput,
)
models = LocalModelBundle(
chunker_dir=Path("/local/models/en/chunker"),
aligner_dir=Path("/local/models/en/aligner"),
)
with FlexAligner(
models=models,
lexicon_path=Path("english.dict"),
) as aligner:
result = aligner.align(
AlignmentRequest(
audio_path=Path("recording.wav"),
transcript=Path("transcript.txt").read_text(encoding="utf-8"),
output=TextGridOutput(path=Path("recording.TextGrid")),
utterance_id="recording",
)
)
print(result.output_sha256)
Input Requirements / 输入要求
-
Audio must be uncompressed 16 kHz mono PCM16 WAV.
-
The transcript must be UTF-8 text and fully covered by the pronunciation dictionary.
-
Model and dictionary files must be available locally.
-
Output paths use no-clobber semantics: an existing output file is not overwritten.
-
音频必须是未压缩的 16 kHz、单声道、PCM16 WAV。
-
文本必须为 UTF-8,且发音词典需要覆盖全部输入词。
-
模型与词典必须提前保存在本地。
-
输出采用不覆盖已有文件的策略;若目标文件已存在,程序不会将其覆盖。
-
--num-threads会设置 Torch 的进程全局 CPU 线程数,并非仅对单个 aligner 实例生效。
🗓️ Roadmap
- Core Alignment Engine: Two-stage CTC chunking and local alignment.
- English CPU Single-File Alignment: CLI, Python API, and validated TextGrid output.
- Continuous TextGrid Coverage:
wordsandphonestiers useNULLintervals to cover the complete timeline. - Mandarin Alignment: Model, segmentation, and release validation.
- GPU and Batch Processing: Accelerated and high-throughput workflows.
- Audio Frontend: Multi-format decoding and automatic resampling.
- Model Distribution: Documented model acquisition and compatibility validation.
- PyPI Release: Publish the approved public preview as
flexaligner.
👨💻 Authors & Affiliation
Yiming Wang (王一鸣) - University of Science and Technology of China (USTC)
Jiahong Yuan (袁家宏) - University of Science and Technology of China (USTC)
📜 Citation
If you use FlexAligner in your research, please cite:
@misc{flexaligner2026,
title = {FlexAligner: Robust Speech--Text Alignment via CTC Chunking and Local Cross-Entropy Alignment},
author = {Wang, Yiming and Yuan, Jiahong},
year = {2026},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/USTCPhonetics/FlexAligner}},
organization = {University of Science and Technology of China}
}
📄 License
FlexAligner is released under the MIT License. Please refer to the
repository's LICENSE file for the authoritative license and copyright notice.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file flexaligner-0.1.0a1.tar.gz.
File metadata
- Download URL: flexaligner-0.1.0a1.tar.gz
- Upload date:
- Size: 57.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6c771a0fb7ffa295d02d55dba82799dced0e90c4960077dce8441c38af4fff67
|
|
| MD5 |
f7c33c99571c86ebf4981d2e7f9b85e5
|
|
| BLAKE2b-256 |
0d51fb13ffa5fb0bdafb317ad2e6aeea73041a14cd463e002f564b2b5184f6f6
|
Provenance
The following attestation bundles were made for flexaligner-0.1.0a1.tar.gz:
Publisher:
release.yml on USTCPhonetics/FlexAligner
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
flexaligner-0.1.0a1.tar.gz -
Subject digest:
6c771a0fb7ffa295d02d55dba82799dced0e90c4960077dce8441c38af4fff67 - Sigstore transparency entry: 2572134013
- Sigstore integration time:
-
Permalink:
USTCPhonetics/FlexAligner@d2cbd4bc4afa6f507016137006ce6094ce82f132 -
Branch / Tag:
refs/tags/v0.1.0a1 - Owner: https://github.com/USTCPhonetics
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d2cbd4bc4afa6f507016137006ce6094ce82f132 -
Trigger Event:
push
-
Statement type:
File details
Details for the file flexaligner-0.1.0a1-py3-none-any.whl.
File metadata
- Download URL: flexaligner-0.1.0a1-py3-none-any.whl
- Upload date:
- Size: 64.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8a47ccd1268f68c9c03a69b4f6526b95715967b0da45e5c0e45c5a602bbe2fdc
|
|
| MD5 |
75b9917d089f18d123953acaa7de9352
|
|
| BLAKE2b-256 |
ac8ce6f02788cb39625a25374fa0133dd5baa7bace6a45d8ba98a28c36471c85
|
Provenance
The following attestation bundles were made for flexaligner-0.1.0a1-py3-none-any.whl:
Publisher:
release.yml on USTCPhonetics/FlexAligner
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
flexaligner-0.1.0a1-py3-none-any.whl -
Subject digest:
8a47ccd1268f68c9c03a69b4f6526b95715967b0da45e5c0e45c5a602bbe2fdc - Sigstore transparency entry: 2572134071
- Sigstore integration time:
-
Permalink:
USTCPhonetics/FlexAligner@d2cbd4bc4afa6f507016137006ce6094ce82f132 -
Branch / Tag:
refs/tags/v0.1.0a1 - Owner: https://github.com/USTCPhonetics
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d2cbd4bc4afa6f507016137006ce6094ce82f132 -
Trigger Event:
push
-
Statement type: