Skip to main content

A word-level Language Identification (LID) tool for Tagalog-English (Taglish) text.

Project description

TagLID

A word-level Language Identification (LID) tool for Tagalog-English (Taglish) text

PyPI version Downloads License GitHub stars CI

About

TagLID labels each word in a Taglish (Tagalog-English mix) text by language. It gives either a simple tag (tgl or eng) or detailed frequency info with flags indicating how the word was identified. It is a rule-based, opinionated system that relies mostly on dictionary lookups, with additional logic for skipping numbers, names, and interjections, and for handling slang, abbreviations, contractions, stemming and lemmatizing inflected words, intrawords, and misspelling correction.

Installation

pip install taglid

Usage

TagLID can be used as a library via import taglid, or as a CLI application via python -m taglid.

Library

Textual data

Use the lid module for textual data.

Use lang_identify to identify each word in a text. It takes any string and returns a list of words with their corresponding English and Tagalog scores, flag, and correction.

from taglid.lid import lang_identify

labeled_text = lang_identify("hello, mundo")
print(labeled_text)

Output:

[{'Word': 'hello', 'eng': 1.0, 'tgl': 0.0, 'Flag': 'DICT', 'Correction': None}, {'Word': 'mundo', 'eng': 0.0, 'tgl': 1.0, 'Flag': 'DICT', 'Correction': None}]

Use tabulate to view the output in tabular format.

from tabulate import tabulate

print(tabulate(labeled_text, headers="keys"))

Output:

word      eng    tgl  flag    correction
------  -----  -----  ------  ------------
hello       1      0  DICT
mundo       0      1  DICT

Use simplify to reduce the output to only the words and their language. It takes the return value of lang_identify and returns a list of tuples containing the word and its language.

from taglid.lid import simplify

simplified_text = simplify(labeled_text)
print(simplified_text)

Output:

[('hello', 'eng'), ('mundo', 'tgl')]

Datasets

Use the lid_dataset module for datasets.

Use lang_identify_df to label each word in each cell of a pandas DataFrame. It takes a DataFrame of multiple rows and columns, where each cell contains textual data, and returns a labeled DataFrame where each token is a row labeled by its original row, original column, and token index.

import pandas as pd
from taglid.lid_dataset import lang_identify_df

data = [["hello po", "ano?"], ["mag-aask lang po", "what?"]]

df = pd.DataFrame(data)

labeled_df = lang_identify_df(df)
print(labeled_df)

Output:

     col  token_index      word  eng  tgl  flag correction
row
0      0            1     hello  1.0  0.0  DICT       None
0      0            2        po  0.0  1.0  DICT       None
0      1            1       ano  0.0  1.0  FREQ       None
1      0            1  mag-aask  0.5  0.5  INTW       None
1      0            2      lang  0.0  1.0  FREQ       None
1      0            3        po  0.0  1.0  DICT       None
1      1            1      what  1.0  0.0  DICT       None

CLI

Run TagLID from the terminal.

python -m taglid.lid

Then type a sentence when prompted.

text: hello, mundo

Output:

word      eng    tgl  flag    correction
------  -----  -----  ------  ------------
hello       1      0  DICT
mundo       0      1  DICT

Add --simplify to only show the words and their language.

python -m taglid.lid --simplify --text hello, mundo

Output:

-----  ---
hello  eng
mundo  tgl
-----  ---

Use lid_dataset with Excel files to directly label spreadsheets.

python -m taglid.lid_dataset in_path out_path

Accuracy

The accuracy hasn't been tested yet.

Development

This project uses uv for dependency management.

Clone the repo and sync dependencies (including dev and test groups):

git clone https://github.com/andrianllmm/taglid.git
cd taglid
uv sync --all-groups

Run the tests:

uv run pytest

Contributing

Contributions are welcome! See CONTRIBUTING.md for details.

License

Distributed under the MIT License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

taglid-0.1.1.tar.gz (543.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

taglid-0.1.1-py3-none-any.whl (540.3 kB view details)

Uploaded Python 3

File details

Details for the file taglid-0.1.1.tar.gz.

File metadata

  • Download URL: taglid-0.1.1.tar.gz
  • Upload date:
  • Size: 543.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.33 {"installer":{"name":"uv","version":"0.11.33","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for taglid-0.1.1.tar.gz
Algorithm Hash digest
SHA256 44cc449eb7144fc3b571be150d1e2b2766a90781f50ff0ff0a7cae4107b7cc51
MD5 f92435cd6b13a103262cebba459fac1d
BLAKE2b-256 22b5129a7065ba676b28957757e009d39f4bb7670a2ccc8a5675532cb238ae0c

See more details on using hashes here.

File details

Details for the file taglid-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: taglid-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 540.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.33 {"installer":{"name":"uv","version":"0.11.33","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for taglid-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 49302f2044dfb3a35e125a2827b9978b56656c95fc94d6015a68b87612c1c35a
MD5 5468ac14d66936078fb4d801373fc2b8
BLAKE2b-256 a02bc60b2b73cff1c527bda2f086a39b93e46e29b0ac247dd74327c117dcf8a2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page