Skip to main content
Help the Python Software Foundation raise $60,000 USD by December 31st!  Building the PSF Q4 Fundraiser

query Unicode script metadata

Project description

Simple Python 3 module to query Unicode UCD script metadata (see UAX

| This module is useful for querying if a text is made of Latin
| Arabic, hiragana, kanji (han), and so on. It works for all scripts
| by the Unicode character database.

| This module is dumb and slow. If you need speed, you probably want to
| implement your own functions.

Sample usage:


>>> import uniscripts
>>> uniscripts.is_script('A', 'Latin')

# if you pass it a string, all characters must match
>>> uniscripts.is_script('はるはあけぼの', 'Hiragana')

>>> uniscripts.is_script('はるはAkebono', 'Hiragana')

# ...but by default, it ignores 'Common' characters, such as punctuation.
>>> uniscripts.is_script('はるは:あけぼの', 'Hiragana')

>>> uniscripts.is_script('中華人民共和国', 'Han') # 'Han' = kanji or hànzì

>>> uniscripts.which_scripts('z')

>>> uniscripts.which_scripts('は')

>>> uniscripts.which_scripts('ー') # U+30FC
['Common', 'Katakana', 'Hiragana', 'Hangul', 'Han', 'Bopomofo', 'Yi']

See docstrings for ``is_script()``, ``which_scripts()``.

Project details

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for uniscripts, version 1.0.5
Filename, size File type Python version Upload date Hashes
Filename, size uniscripts-1.0.5.tar.gz (16.1 kB) File type Source Python version None Upload date Hashes View

Supported by

Pingdom Pingdom Monitoring Google Google Object Storage and Download Analytics Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page