Formosan Languages Grapheme-to-Phoneme (G2P) Toolkit - 台灣本土語言文字轉音素工具
Project description
FormoSpeech G2P
Grapheme-to-Phoneme (G2P) Toolkit for Taiwanese Languages
Features
- G2P Conversion: Convert text to IPA or Pinyin pronunciation sequences
- Smart Tokenization: Jieba-based segmentation using only dialect-specific lexicons, without mixing in the default Chinese dictionary
- Variant Character Normalization: Automatic conversion of variant characters to standard forms
- Mixed Chinese-English Support: Optional English pronunciation integration
- Extended CJK Support: Full support for CJK Extension B–H and Private Use Area characters commonly used in Taiwanese languages
- Unknown Word Detection: Automatic identification and reporting of out-of-vocabulary words
Supported Hakka Dialects
- 客語_四縣(hak_sx)
- 客語_南四縣(hak_nsx)
- 客語_海陸(hak_hl)
- 客語_大埔(hak_dp)
- 客語_饒平(hak_rp)
- 客語_詔安(hak_za)
Installation
From PyPI
pip install formog2p
From Github
pip install git+https://github.com/hungshinlee/formospeech-g2p.git
Development Installation
git clone https://github.com/hungshinlee/formospeech-g2p.git
cd formospeech-g2p
# Using uv (recommended)
uv sync --all-extras
# Or using pip
pip install -e ".[dev]"
# Install pre-commit hooks (optional)
pre-commit install
Quick Start
from formog2p.hakka import g2p
# Basic G2P conversion
result = g2p("天公落水", "hak_sx", "ipa")
print(result.pronunciations)
# ['tʰ-ien_24 k-uŋ_24', 'l-ok_5 s-ui_31']
# Check for unknown words
if result.has_unknown:
print(f"Unknown words: {result.unknown_words}")
Usage
G2P conversion (Concatenated String)
# To get a simple string output, you can use the result object
# or use text_to_pronunciation
from formog2p.hakka import text_to_pronunciation
pron_str = text_to_pronunciation("天公落水", "hak_sx", "ipa")
# 'tʰ-ien_24 k-uŋ_24 l-ok_5 s-ui_31'
G2P Parameters
result = g2p(
text, # Input text
lang_group="hak_sx", # Dialect name
pronunciation_type="ipa", # Pronunciation format: "ipa" or "pinyin"
unknown_token=None, # Replacement token for unknown words
keep_unknown=True, # Whether to keep unknown words in output
use_variant_map=True, # Whether to apply variant character conversion
include_eng=False, # Whether to include English pronunciations
)
Mixed Chinese-English G2P
from formog2p.hakka import g2p
# Enable English pronunciation (IPA only)
result = g2p("天公落水Hello World", "hak_sx", "ipa", include_eng=True)
print(result.pronunciations)
# ['tʰ-ien_24 k-uŋ_24', 'l-ok_5 s-ui_31', 'h ə l oʊ', 'w ɝ l d']
# Without English (English words treated as unknown)
result = g2p("天公落水Hello", "hak_sx", "ipa", include_eng=False)
print(result.unknown_words)
# ['Hello']
Text Normalization
from formog2p.hakka import normalize, apply_variant_map
# Full normalization (including variant character conversion)
normalize("天公落水!") # '天公落水!' (half-width to full-width)
normalize("台灣真好") # '臺灣真好' (variant character conversion)
normalize("Hello", use_variant_map=True) # 'HELLO' (uppercase conversion)
# Apply variant character conversion only
apply_variant_map("台灣") # '臺灣'
apply_variant_map("温泉") # '溫泉'
Normalization processing steps:
- Unicode NFKC normalization (full-width to half-width)
- Half-width punctuation to full-width (
, ? ! .→,?!。) - Remove unnecessary punctuation (keep
,。?!) - Variant character conversion (optional)
- Uppercase conversion
Punctuation Handling
Punctuation marks ,。?! are treated as known tokens and output directly:
result = g2p("天公落水,好靚!", "hak_sx", "ipa")
print(result.pronunciations)
# ['tʰ-ien_24 k-uŋ_24', 'l-ok_5 s-ui_31', ',', '好靚', '!']
Basic Tokenization
from formog2p.hakka import run_jieba
# Tokenize with specific dialect
words, oovs = run_jieba("天公落水", "hak_sx")
# words: ['天公', '落水']
# Include English dictionary
words, oovs = run_jieba("天公落水ABC", "hak_sx", include_eng=True)
# words: ['天公', '落水', 'ABC']
Pronunciation Lookup
from formog2p.hakka import get_pronunciation
# Query pronunciation for a single word
pron = get_pronunciation("天公", "hak_sx", "ipa")
# ['tʰ-ien_24 k-uŋ_24']
Tokenizer Cache Management
Tokenizers are loaded and cached on first invocation; subsequent calls retrieve from cache:
from formog2p.hakka import get_cached_tokenizers, clear_tokenizer_cache
# View cached tokenizers
get_cached_tokenizers()
# ['hak_sx', ...]
# Clear cache (if dictionary reload is needed)
clear_tokenizer_cache()
Unicode Support
Full support for extended character sets commonly used in Taiwanese languages:
| Range | Description |
|---|---|
U+2E80-U+9FFF |
CJK Radicals, Basic CJK |
U+F900-U+FAFF |
CJK Compatibility Ideographs |
U+20000-U+323AF |
CJK Extension B–H (Taiwanese language characters) |
U+E000-U+F8FF |
Private Use Area (PUA) |
U+F0000-U+10FFFD |
Supplementary Private Use Areas A & B (custom characters) |
API Reference
G2P Functions
| Function | Description |
|---|---|
g2p(text, lang_group, pronunciation_type, ...) |
Full G2P conversion, returns G2PResult |
text_to_pronunciation(text, lang_group, pronunciation_type, ...) |
Returns concatenated pronunciation string |
normalize(text, use_variant_map) |
Text normalization |
apply_variant_map(text) |
Apply variant character conversion |
G2PResult Object
| Attribute | Type | Description |
|---|---|---|
pronunciations |
list[str] |
Pronunciation sequence |
unknown_words |
list[str] |
List of unknown words |
details |
list[dict] |
Detailed word-pronunciation mapping |
has_unknown |
bool |
Whether unknown words exist |
Tokenization Functions
| Function | Description |
|---|---|
run_jieba(text, lang_group, include_eng) |
Tokenize and return (words, unknown_words) |
Pronunciation Lookup
| Function | Description |
|---|---|
get_pronunciation(word, lang_group, pronunciation_type) |
Query single word pronunciation |
segment_with_pronunciation(text, lang_group, include_eng) |
Tokenize with pronunciation |
Statistics and Cache
| Function | Description |
|---|---|
get_cached_tokenizers() |
Get list of cached tokenizers |
clear_tokenizer_cache() |
Clear tokenizer cache |
Project Structure
formospeech-g2p/
├── pyproject.toml
├── README.md
├── README_zh-TW.md
├── lexicon/ # Pronunciation dictionaries
│ ├── ipa/
│ └── pinyin/
├── share/ # Shared resources (e.g. variant map)
├── formog2p/ # Source code
│ ├── __init__.py
│ └── hakka/
│ ├── __init__.py
│ └── g2p.py # Main G2P logic
License
MIT License
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file formog2p-0.1.3.tar.gz.
File metadata
- Download URL: formog2p-0.1.3.tar.gz
- Upload date:
- Size: 6.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8ba2aa274c9e4495286b6a73578d1b0fa0d93b9630bf54f99a0c694c85696d95
|
|
| MD5 |
817f643d7678bfbe04bfe4c62c7a6cc7
|
|
| BLAKE2b-256 |
de134d80252e823b21416ce8617718774842281e417586d96274f42df1e55d87
|
Provenance
The following attestation bundles were made for formog2p-0.1.3.tar.gz:
Publisher:
publish.yml on hungshinlee/formospeech-g2p
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
formog2p-0.1.3.tar.gz -
Subject digest:
8ba2aa274c9e4495286b6a73578d1b0fa0d93b9630bf54f99a0c694c85696d95 - Sigstore transparency entry: 1263236593
- Sigstore integration time:
-
Permalink:
hungshinlee/formospeech-g2p@e8b5183d23d0cae6f8493d5b2d7fcd28dc412dd1 -
Branch / Tag:
refs/tags/v0.1.3 - Owner: https://github.com/hungshinlee
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@e8b5183d23d0cae6f8493d5b2d7fcd28dc412dd1 -
Trigger Event:
push
-
Statement type:
File details
Details for the file formog2p-0.1.3-py3-none-any.whl.
File metadata
- Download URL: formog2p-0.1.3-py3-none-any.whl
- Upload date:
- Size: 6.9 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5d5a20cb9d305ec1bbd63ac7009376998e78ccb6dd618d702ac1f4d67630c399
|
|
| MD5 |
60269c5f21e8f1b828452f5add1bfe76
|
|
| BLAKE2b-256 |
e472a999409fba63bb2aed57f194e504c817141e94e712390ccada9596988683
|
Provenance
The following attestation bundles were made for formog2p-0.1.3-py3-none-any.whl:
Publisher:
publish.yml on hungshinlee/formospeech-g2p
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
formog2p-0.1.3-py3-none-any.whl -
Subject digest:
5d5a20cb9d305ec1bbd63ac7009376998e78ccb6dd618d702ac1f4d67630c399 - Sigstore transparency entry: 1263236599
- Sigstore integration time:
-
Permalink:
hungshinlee/formospeech-g2p@e8b5183d23d0cae6f8493d5b2d7fcd28dc412dd1 -
Branch / Tag:
refs/tags/v0.1.3 - Owner: https://github.com/hungshinlee
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@e8b5183d23d0cae6f8493d5b2d7fcd28dc412dd1 -
Trigger Event:
push
-
Statement type: