Scraping grapheme-to-phoneme data from Wiktionary

These details have not been verified by PyPI

Project links

Homepage

Project description

WikiPron

WikiPron is a command-line tool and Python API for mining multilingual pronunciation data from Wiktionary, as well as a database of pronunciation dictionaries mined using this tool.

Command-line tool
Python API
Data
Models
Development

If you use WikiPron in your research, please cite the following:

Jackson L. Lee, Lucas F.E. Ashby, M. Elizabeth Garza, Yeonju Lee-Sikka, Sean Miller, Alan Wong, Arya D. McCarthy, and Kyle Gorman (2020). Massively multilingual pronunciation mining with WikiPron. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4223-4228. [bibtex]

Command-line tool

Installation

pip install wikipron

Usage

Quick start

After installation, the terminal command wikipron will be available. As a basic example, the following command scrapes G2P data for French:

wikipron fra

Specifying the language

The language is indicated by a three-letter ISO 639-3 language code, e.g., fra for French. For which languages can be scraped, here is the complete list of languages on Wiktionary that have pronunciation entries.

Specifying the dialect

One can optionally specify dialects to target using the --dialect flag. The dialect name can be found together with the transcription on Wiktionary. For example, "(UK, US) IPA: /təˈmɑːtəʊ/". To restrict to the union of dialects use the pipe character '|': e.g., --dialect='General American | US'. Transcriptions which lack a dialect specification are selected regardless of the value of this flag.

Specifying the transcription level

By default, WikiPron selects broad pronunciations in angled brackets /like this/. One can instead select narrow transcriptions written [like this] using the --narrow flag. Note that some languages only have broad or narrow transcriptions (e.g., Russian only has the latter.

Segmentation

By default, the segments library is used to segment the transcription into whitespace. The segmentation tends to place IPA diacritics and modifiers on the "parent" symbol. For instance, [kʰæt] is rendered kʰ æ t. This can be disabled using the --no-segment flag.

Parentheses

Some of transcriptions contain parentheses to indicate alternative pronunciations. The parentheses (but not the content) are discarded in the scrape unless the --no-skip-parens flag is used.

Output

The scraped data is organized with each <word, pronunciation> pair on its own line, where the word and pronunciation are separated by a tab. Note that the pronunciation is in International Phonetic Alphabet (IPA), segmented by spaces that correctly handle the combining and modifier diacritics for modeling purposes, e.g., we have kʰ æ t with the aspirated k instead of k ʰ æ t.

For illustration, here is a snippet of French data scraped by WikiPron:

accrémentitielle    a k ʁ e m ɑ̃ t i t j ɛ l
accrescent  a k ʁ ɛ s ɑ̃
accrétion   a k ʁ e s j ɔ̃
accrétions  a k ʁ e s j ɔ̃

By default, the scraped data appears in the terminal. To save the data in a TSV file, please redirect the standard output to a filename of your choice:

wikipron fra > fra.tsv

Advanced options

The wikipron terminal command has an array of options to configure your scraping run. For a full list of the options, please run wikipron -h.

Python API

The underlying module can also be used from Python. A standard workflow looks like:

import wikipron

config = wikipron.Config(key="fra")  # French, with default options.
for word, pron in wikipron.scrape(config):
    ...

Data

We also make available a database of over 3 million word/pronunciation pairs mined using WikiPron.

Models

We host grapheme-to-phoneme models and modeling software in a separate repository.

Development

Repository

The source code of WikiPron is hosted on GitHub at https://github.com/CUNY-CL/wikipron, where development also happens.

For the latest changes not yet released through pip or working on the codebase yourself, you may obtain the latest source code through GitHub and git:

Create a fork of the wikipron repo on your GitHub account.
Locally, make sure you are in some sort of a virtual environment (venv, virtualenv, conda, etc).

Download and install the library in the "editable" mode together with the core and dev dependencies within the virtual environment:

git clone https://github.com/<your-github-username>/wikipron.git
cd wikipron
pip install -U pip setuptools
pip install -r requirements.txt
pip install --no-deps -e .

We keep track of notable changes in CHANGELOG.md.

Contributing

For questions, bug reports, and feature requests, please file an issue.

If you would like to contribute to the wikipron codebase, please see CONTRIBUTING.md.

License

WikiPron is released under an Apache 2.0 license. Please see LICENSE.txt for details.

Please note that Wiktionary data in the data/ directory has its own licensing terms.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

This version

1.3.3

Jul 27, 2024

1.3.2

Jul 17, 2024

1.3.1

Mar 2, 2024

1.3.0

Nov 28, 2022

1.2.0

Jan 30, 2021

1.1.0

Mar 3, 2020

1.0.0

Nov 29, 2019

0.1.1

Aug 15, 2019

0.1.0

Aug 14, 2019

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wikipron-1.3.3.tar.gz (23.8 kB view details)

Uploaded Jul 27, 2024 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

wikipron-1.3.3-py3-none-any.whl (27.4 kB view details)

Uploaded Jul 27, 2024 Python 3

File details

Details for the file wikipron-1.3.3.tar.gz.

File metadata

Download URL: wikipron-1.3.3.tar.gz
Upload date: Jul 27, 2024
Size: 23.8 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.1 CPython/3.12.2

File hashes

Hashes for wikipron-1.3.3.tar.gz
Algorithm	Hash digest
SHA256	`5eb3d1c1bdba89baa7479d7822dde779e19d12aa9ad49d530984927df058fa15`
MD5	`933e6e1d6566d1feb30ca7f7a447b7c9`
BLAKE2b-256	`d58419a9a932e4c45be859ed9bb51e68ab91620b0ee2cbc6459ec4bae2942111`

See more details on using hashes here.

File details

Details for the file wikipron-1.3.3-py3-none-any.whl.

File metadata

Download URL: wikipron-1.3.3-py3-none-any.whl
Upload date: Jul 27, 2024
Size: 27.4 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.1 CPython/3.12.2

File hashes

Hashes for wikipron-1.3.3-py3-none-any.whl
Algorithm	Hash digest
SHA256	`6c29497dc83ddbbccda8f0d3594c45dba7918a05e322ad9d900daed784b435e8`
MD5	`117c593988e0610124272e81e0159151`
BLAKE2b-256	`669103960973748cc260b3dc555ed1921b0a5b150b96b9c6b609ac8965f749e7`

See more details on using hashes here.

wikipron 1.3.3

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

WikiPron

Command-line tool

Installation

Usage

Quick start

Specifying the language

Specifying the dialect

Specifying the transcription level

Segmentation

Parentheses

Output

Advanced options

Python API

Data

Models

Development

Repository

Contributing

License

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes