Skip to main content

A Word level Language Identification (LID) tool for Tagalog-English (Taglish) text.

Project description

TagLID

A word-level Language Identification (LID) tool for Tagalog-English (Taglish) text

About

TagLID is a library that labels each word in a Taglish (Tagalog-English mix) text by language. It gives either a simple tag (tgl or eng) or detailed frequency info with flags indicating how the word was identified. It is a rule-based and opinionated system that mostly uses dictionary lookups. It also handles cases like skipping numbers, names, and interjections, and includes logic for dealing with slang, abbreviations, contractions, stemming or lemmatizing inflected words, intrawords, and correcting misspellings.

Installation

pip install taglid

Usage

TagLID can act as a standalone library that can be imported via import taglid or as a CLI application via python -m taglid.

Library Mode

Textual data

Use the lid module for textual data.

Use lang_identify to identify each word in a text. This takes any string and returns a list of words and their corresponding English and Tagalog values, flag, and correction.

from taglid.lid import lang_identify

labeled_text = lang_identify("hello, mundo")
print(labeled_text)

Output:

[{'Word': 'hello', 'eng': 1.0, 'tgl': 0.0, 'Flag': 'DICT', 'Correction': None}, {'Word': 'mundo', 'eng': 0.0, 'tgl': 1.0, 'Flag': 'DICT', 'Correction': None}]

Use tabulate to view output in tabular format.

from tabulate import tabulate

print(tabulate(labeled_text, headers="keys"))

Output:

word      eng    tgl  flag    correction
------  -----  -----  ------  ------------
hello       1      0  DICT
mundo       0      1  DICT

Use simplify to only show the words and their language. This takes the return value of lang_identify and returns a list of tuples containing the word and its language.

from taglid.lid import simplify

simplified_text = simplify(labeled_text)
print(simplified_text)

Output:

[('hello', 'eng'), ('mundo', 'tgl')]

Datasets

Use the lid_dataset module for datasets.

Use lang_identify_df to label each word in each cell in a pandas DataFrame. This takes a DataFrame of multiple rows and columns with each cell containing textual data and returns a labeled DataFrame where each token is a row labeled by its original row, original column, and token index.

import pandas as pd
from taglid.lid_dataset import lang_identify_df

data = [["hello po", "ano?"], ["mag-aask lang po", "what?"]]

df = pd.DataFrame(data)

labeled_df = lang_identify_df(df)
print(labeled_df)

Output:

     col  token_index      word  eng  tgl  flag correction
row
0      0            1     hello  1.0  0.0  DICT       None
0      0            2        po  0.0  1.0  DICT       None
0      1            1       ano  0.0  1.0  FREQ       None
1      0            1  mag-aask  0.5  0.5  INTW       None
1      0            2      lang  0.0  1.0  FREQ       None
1      0            3        po  0.0  1.0  DICT       None
1      1            1      what  1.0  0.0  DICT       None

CLI Mode

Run TagLID from the terminal.

python -m taglid.lid

Then type a sentence when prompted.

text: hello, mundo

Output:

word      eng    tgl  flag    correction
------  -----  -----  ------  ------------
hello       1      0  DICT
mundo       0      1  DICT

Add --simplify to only show the words and their language.

python -m taglid.lid --simplify --text hello, mundo

Output:

-----  ---
hello  eng
mundo  tgl
-----  ---

Use lid_dataset with Excel files to directly label spreadsheets.

python -m taglid.lid_dataset in_path out_path

Accuracy

The accuracy hasn't been tested yet.

Development

This project uses uv for dependency management.

Clone the repo and sync dependencies (including dev and test groups):

git clone https://github.com/andrianllmm/taglid.git
cd taglid
uv sync --all-groups

Run the tests:

uv run pytest

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

taglid-0.1.0.tar.gz (543.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

taglid-0.1.0-py3-none-any.whl (540.7 kB view details)

Uploaded Python 3

File details

Details for the file taglid-0.1.0.tar.gz.

File metadata

  • Download URL: taglid-0.1.0.tar.gz
  • Upload date:
  • Size: 543.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.33 {"installer":{"name":"uv","version":"0.11.33","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for taglid-0.1.0.tar.gz
Algorithm Hash digest
SHA256 4ac840ee2fe7948c49e0cefc0f01351195026e72e1c82e2ff010afd6aea274fc
MD5 2849fea110000ec4221a43f6e93d5726
BLAKE2b-256 b369a5c682e77ce83ead2d6597f9eab33c01afcf41dabbef5ed820314e7d10ff

See more details on using hashes here.

File details

Details for the file taglid-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: taglid-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 540.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.33 {"installer":{"name":"uv","version":"0.11.33","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for taglid-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d453d962f0de7289bf25aaf59c19d77f4f9aaa51f5b70310107803da60f6c704
MD5 a4d729b41a6d11b8f075c72ad4ad68bd
BLAKE2b-256 e7950c29a056cd441e4f59de95d4b98ad74adf6f3a7344ab7938161da083e711

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page