g2p-en

A Simple Python Module for English Grapheme To Phoneme Conversion

These details have not been verified by PyPI

Project links

Project description

[Update] * We removed TensorFlow from the dependencies. After all, it changes its APIs quite often, and we don’t expect you to have a GPU. Instead, NumPy is used for inference.

This module is designed to convert English graphemes (spelling) to phonemes (pronunciation). It is considered essential in several tasks such as speech synthesis. Unlike many languages like Spanish or German where pronunciation of a word can be inferred from its spelling, English words are often far from people’s expectations. Therefore, it will be the best idea to consult a dictionary if we want to know the pronunciation of some word. However, there are at least two tentative issues in this approach. First, you can’t disambiguate the pronunciation of homographs, words which have multiple pronunciations. (See a below.) Second, you can’t check if the word is not in the dictionary. (See b below.)

a. I refuse to collect the refuse around here. (rɪ|fju:z as verb vs. |refju:s as noun)
b. I am an activationist. (activationist: newly coined word which means n. A person who designs and implements programs of treatment or therapy that use recreation and activities to help people whose functional abilities are affected by illness or disability. from WORD SPY

For the first homograph issue, fortunately many homographs can be disambiguated using their part-of-speech, if not all. When it comes to the words not in the dictionary, however, we should make our best guess using our knowledge. In this project, we employ a deep learning seq2seq framework based on TensorFlow.

Algorithm

Spells out arabic numbers and some currency symbols. (e.g. $200 -> two hundred dollars) (This is borrowed from Keith Ito’s code)
Attempts to retrieve the correct pronunciation for homographs based on their POS)
Looks up The CMU Pronouncing Dictionary for non-homographs.
For OOVs, we predict their pronunciations using our neural net model.

Environment

python 3.x

Dependencies

numpy >= 1.13.1
nltk >= 3.2.4
python -m nltk.downloader “averaged_perceptron_tagger” “cmudict”
inflect >= 0.3.1
Distance >= 0.1.3

Installation

pip install g2p_en

python setup.py install

nltk package will be automatically downloaded at your first run.

Usage

from g2p_en import G2p

texts = ["I have $250 in my pocket.", # number -> spell-out
         "popular pets, e.g. cats and dogs", # e.g. -> for example
         "I refuse to collect the refuse around here.", # homograph
         "I'm an activationist."] # newly coined word
g2p = G2p()
for text in texts:
    out = g2p(text)
    print(out)
>>> ['AY1', ' ', 'HH', 'AE1', 'V', ' ', 'T', 'UW1', ' ', 'HH', 'AH1', 'N', 'D', 'R', 'AH0', 'D', ' ', 'F', 'IH1', 'F', 'T', 'IY0', ' ', 'D', 'AA1', 'L', 'ER0', 'Z', ' ', 'IH0', 'N', ' ', 'M', 'AY1', ' ', 'P', 'AA1', 'K', 'AH0', 'T', ' ', '.']
>>> ['P', 'AA1', 'P', 'Y', 'AH0', 'L', 'ER0', ' ', 'P', 'EH1', 'T', 'S', ' ', ',', ' ', 'F', 'AO1', 'R', ' ', 'IH0', 'G', 'Z', 'AE1', 'M', 'P', 'AH0', 'L', ' ', 'K', 'AE1', 'T', 'S', ' ', 'AH0', 'N', 'D', ' ', 'D', 'AA1', 'G', 'Z']
>>> ['AY1', ' ', 'R', 'IH0', 'F', 'Y', 'UW1', 'Z', ' ', 'T', 'UW1', ' ', 'K', 'AH0', 'L', 'EH1', 'K', 'T', ' ', 'DH', 'AH0', ' ', 'R', 'EH1', 'F', 'Y', 'UW2', 'Z', ' ', 'ER0', 'AW1', 'N', 'D', ' ', 'HH', 'IY1', 'R', ' ', '.']
>>> ['AY1', ' ', 'AH0', 'M', ' ', 'AE1', 'N', ' ', 'AE2', 'K', 'T', 'IH0', 'V', 'EY1', 'SH', 'AH0', 'N', 'IH0', 'S', 'T', ' ', '.']

May, 2018.

Kyubyong Park & Jongseok Kim

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

2.1.0

Dec 31, 2019

2.0.0

Apr 8, 2019

1.0.1

Oct 1, 2018

1.0.0

May 17, 2018

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

g2p_en-2.1.0.tar.gz (3.1 MB view details)

Uploaded Dec 31, 2019 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

g2p_en-2.1.0-py3-none-any.whl (3.1 MB view details)

Uploaded Dec 31, 2019 Python 3

File details

Details for the file g2p_en-2.1.0.tar.gz.

File metadata

Download URL: g2p_en-2.1.0.tar.gz
Upload date: Dec 31, 2019
Size: 3.1 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.21.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.31.1 CPython/3.7.1

File hashes

Hashes for g2p_en-2.1.0.tar.gz
Algorithm	Hash digest
SHA256	`32ecb119827a3b10ea8c1197276f4ea4f44070ae56cbbd01f0f261875f556a58`
MD5	`a2472d72e09d266a3d725a3bf839e5b6`
BLAKE2b-256	`5f222c7acbe6164ed6cfd4301e9ad2dbde69c68d22268a0f9b5b0ee6052ed3ab`

See more details on using hashes here.

File details

Details for the file g2p_en-2.1.0-py3-none-any.whl.

File metadata

Download URL: g2p_en-2.1.0-py3-none-any.whl
Upload date: Dec 31, 2019
Size: 3.1 MB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.21.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.31.1 CPython/3.7.1

File hashes

Hashes for g2p_en-2.1.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`2a7aabf1fc7f270fcc3349881407988c9245173c2413debbe5432f4a4f31319f`
MD5	`c3a482f7940df3d620e0f172a33981c9`
BLAKE2b-256	`d7d9b77dc634a7a0c0c97716ba97dd0a28cbfa6267c96f359c4f27ed71cbd284`

See more details on using hashes here.

g2p-en 2.1.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Project description

Algorithm

Environment

Dependencies

Installation

Usage

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes