Skip to main content

g2p-mix

PyPI License

Mixed Chinese–English grapheme-to-phoneme conversion with source alignment.

The package intentionally supports two modes:

  • Mandarin + English
  • Cantonese + English

Python 3.10 or newer is required.

Installation

pip install g2p-mix

Quick start

from g2p_mix import G2P

g2p = G2P()
result = g2p("你这个 idea,不太 make sense。")

print(result.phones)
('n', 'i3', 'zh', 'e4', 'g', 'e5', 'AY0', 'D', 'IY1', 'AH0', 'b', 'u2', 't', 'ai4', 'M', 'EY1', 'K', 'S', 'EH1', 'N', 'S')

Mandarin is the default. Use Cantonese by changing only the mode:

g2p = G2P("cantonese")
print(g2p("你好 idea").phones)
('n', 'ei5', 'h', 'ou2', 'AY0', 'D', 'IY1', 'AH0')

result.phones is always the final, directly usable output. Numeric tones are attached in native mode; IPA tone letters and English stress marks are attached in IPA mode.

Arabic numbers and other written forms are normalized automatically by WeText:

result = G2P()("版本1.0发布于2026年")
print(result.normalized_text)
版本一点零发布于二零二六年

Unicode compatibility letters and numbers are normalized before TN, so full-width input such as ABC 123 is supported without changing Chinese punctuation. Decomposable English diacritics are folded only for pronunciation lookup, so café retains its spelling and source spans while using the CMUdict entry for cafe. Latin text that cannot be folded safely still raises G2PError instead of silently losing phones.

When a sentence contains a covered POS-dependent homograph, the English backend tags the complete English projection once and shares that context across its tokens. For example, record receives different pronunciations in I record music and This is a record. Chinese islands remain visible to the tagger as one <ZH> placeholder and punctuation is retained. Unambiguous sentences skip POS tagging, while words outside CMUdict continue through segmentation and the g2p-en OOV predictor. Project-reviewed corrections to upstream homograph data live in an external resource rather than being embedded in backend code.

Unknown Chinese characters are strict by default. Use preserve when a partially pronounced result is preferable to rejecting the whole sentence:

result = G2P(unknown="preserve")("你𲎯好")
print(result.phones)
print(result.warnings)

The unknown character remains as a source-aligned unit with is_unknown=True, empty phones, and a warning; no pronunciation is invented. A compatible secondary backend can instead be selected explicitly:

g2p = G2P(
    backend="g2pw",
    fallback_backend="pypinyin",
)

IPA

from g2p_mix import G2P

g2p = G2P(output="ipa", tone_sandhi=False)
result = g2p("中国 idea")

print(result.phones)
('ʈ͡ʂ', 'ʊ', 'ŋ˥˥', 'k', 'w', 'o˧˥', 'a', 'ɪ', 'd', 'ˈi', 'ə')

Base phones without tone or stress remain available separately:

print(result.base_phones)
('ʈ͡ʂ', 'ʊ', 'ŋ', 'k', 'w', 'o', 'a', 'ɪ', 'd', 'i', 'ə')

Backends

Built-in Chinese backends are selected by name:

Mode Default Alternatives
mandarin pypinyin g2pw
cantonese tojyutping pycantonese
g2p = G2P("mandarin", backend="g2pw")

G2PW is optional:

pip install "g2p-mix[g2pw]"

The G2PW backend keeps the upstream ONNX session by default; pass G2PWBackend(onnx_threads=N) to rebuild it with a different intra-op thread count (upstream hardcodes 2). The knob is latency-only and never changes predictions.

Two opt-in switches trade a sliver of accuracy for speed. Both need the extra graph files produced by scripts/prepare_g2pw_split.py (which also verifies that the split graphs are bit-exact on the upstream per-position inputs):

  • G2PWBackend(quantize=True) loads a dynamic-INT8 copy of the model (606 MB -> 159 MB) with the same windowed inference semantics; clean-CPP target accuracy measured 96.0269% vs 96.0716% fp32.
  • G2PWBackend(use_split_model=True) runs the BERT encoder once per sentence instead of once per polyphonic position; combine with quantize=True for a quantized encoder. Sharing the encoder sees the whole sentence instead of the training-time ±16 character window, so predictions drift slightly on longer sentences (clean-CPP 95.7359%).

Both numbers and the default fp32 path's latency are listed in benchmarks/baseline.md. Downloads are scoped to what a configuration actually runs: the default path skips the split/INT8 graphs and the 393 MB tokenizer checkpoint, the split path skips the monolithic graph, and INT8 paths skip their fp32 counterparts (so use_split_model=True, quantize=True pulls about 155 MB).

A custom backend object can be passed through the same argument:

g2p = G2P("mandarin", backend=MyMandarinBackend())

Mandarin phrase lexicon

g2p_mix/dict/phrases.txt pins curated readings for polyphonic phrases, one entry per line (word syllable ..., with trailing tone digits and 5 for the neutral tone). The lexicon is re-matched against the analyzed tokens after each Mandarin backend — left to right, longest match first — so an entry keeps working for both pypinyin and g2pw even when the segmenter splits the phrase apart. It only rewrites listed phrases and never changes anything else.

Detailed results

Most applications only need result.phones. result.base_phones always removes Mandarin and Cantonese tones as well as English stress. Source-aligned units are available when more detail is required:

for unit in result.units:
    print(unit.text, unit.phones, unit.tone, unit.source_spans)

IPA units also retain their original alphabet and phones:

for unit in result.units:
    print(unit.source_alphabet, unit.source_phones)

The input remains losslessly reconstructable:

assert result.reconstruct_original() == "中国 idea"

Phonetic similarity

Install the optional PanPhon backend:

pip install "g2p-mix[similarity]"

Then compare text directly through the same G2P object:

g2p = G2P("mandarin", tone_sandhi=False)

near = g2p.compare("西", "she")
far = g2p.compare("西", "key")

assert near.score > far.score
print(near.score, near.alignment)

Similarity currently uses base phones. Tone and stress remain available in the structured result but are intentionally excluded from the score. PanPhon is loaded lazily and is not required for G2P or IPA output.

CLI

g2p_mix "你这个 idea。"
g2p_mix "你这个 idea。" --mode cantonese
g2p_mix "你这个 idea。" --output ipa
g2p_mix "银行 ATM" --backend g2pw --format json

Advanced usage

The root package exposes only the simple API:

from g2p_mix import G2P, G2PError, G2PResult

Backend protocols, structured models, transcription, projections, and the internal pipeline live in their respective submodules. See Architecture and extension points.

Development

python -m pip install -U pip
python -m pip install -e ".[dev]"
ruff check .
ruff format --check .
python -m pytest

The normal suite uses an injected converter for G2PW tests and does not download a model. Run the real-model smoke test explicitly:

python -m pip install -e ".[g2pw,test]"
G2P_MIX_TEST_G2PW=1 python -m pytest -m g2pw

Quality evaluation is separate from unit tests and the published wheel:

python -m benchmarks
python -m benchmarks --json
python -m benchmarks --fail-under 1.0
python -m benchmarks --corpus cpp --max-cases 100 --seed 42
python -m benchmarks --corpus hkcancor --cantonese-backend tojyutping
python -m benchmarks benchmarks/data/mandarin_normalization_sandhi.json

See benchmarks/README.md for the dataset schema, backend comparison options, reproducible CPP/HKCanCor adapters, metrics, and measured baselines.

Metadata

Release files for g2p-mix 0.7.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for g2p-mix 0.7.1
File Size Uploaded
g2p_mix-0.7.1.tar.gz 5.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for g2p-mix 0.7.1
File Interpreter ABI Platform
g2p_mix-0.7.1-py3-none-any.whl Python 3 none any Details

Total release size: 10.1 MB

Release files / g2p_mix-0.7.1.tar.gz

Download URL g2p_mix-0.7.1.tar.gz
Size 5.0 MB
Tags Source
SHA-256 checksum
How to use checksums
d17769a16184622bf5dd5862444f3e660f8997d5beb00f8fa14767ce0082ad76
BLAKE2b-256 checksum
How to use checksums
249efda563c49332c110f25f705c5fe90b42f1787ec1badc294f0c962c31913a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / g2p_mix-0.7.1-py3-none-any.whl

Download URL g2p_mix-0.7.1-py3-none-any.whl
Size 5.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
2c43eeb885a119863e616646db02b1803518c7be4683e1f1a4bf433f6067eba8
BLAKE2b-256 checksum
How to use checksums
13e866ac6bf3b19b2069459cb09e0f1a75204e7e59dcd466799511eefde53b8b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

0.7.1 This release

2 release files

0.7.0

2 release files

0.6.5

1 release file

0.6.4

1 release file

0.6.3

1 release file

0.6.2

1 release file

0.6.1

1 release file

0.6.0

1 release file

0.5.9

1 release file

0.5.8

1 release file

0.5.7

1 release file

0.5.6

1 release file

0.5.5

1 release file

0.5.4

1 release file

0.5.3

1 release file

0.5.2

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.9

1 release file

0.4.8

1 release file

0.4.2

1 release file

0.4.1

1 release file

0.3.9

1 release file

0.3.8

1 release file

0.3.7

1 release file

0.3.6

1 release file

0.3.5

1 release file

0.3.1

1 release file

0.3.0

1 release file

0.2.9

1 release file

0.2.8

1 release file

0.2.7

1 release file

0.2.6

1 release file

0.2.5

1 release file

0.2.4

1 release file

0.1.7

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page