Tetun Tokenizer
Tetun tokenizer is a Python package used to tokenize an input text into tokens. It offers several tokenization techniques as follows:
TetunStandardTokenizer: Segments the input text into individual tokens based on word boundaries, punctuation, and special characters.TetunSentenceTokenizer: Splits sentences using ending delimiters such as periods (.), question marks (?), and exclamation marks (!). Titles represented by periods, such as Dr. and Ph.D., are preserved.TetunBlankLineTokenizer: Segments the input text based on the presence of blank lines.TetunSimpleTokenizer: Extracts only strings and numbers from the input text while discarding punctuation and special characters.TetunWordTokenizer: Extracts word units from the input text, excluding numbers, punctuation, and special characters.
Installation
You can install the Tetun tokenizer via pip:
pip install tetun-tokenizer
Usage
To utilize the Tetun tokenizer, simply import the desired tokenizer feature/class from the tokenizer module within the Tetun tokenizer package. Here are some examples demonstrating its usage:
- Utilizing
TetunStandardTokenizerto tokenize the input text.
from tetuntokenizer.tokenizer import TetunStandardTokenizer
# Instantiate Tetun Standard Tokenizer class
tetun_tokenizer = TetunStandardTokenizer()
# Input text
text = "Ha'u mak ita-nia maluk di'ak. Ha'u iha $0.25 atu fó ba ita."
# Tokenize the input text
output = tetun_tokenizer.tokenize(text)
# Print the output
print(output)
Expected output:
["Ha'u", 'mak', 'ita-nia', 'maluk', "di'ak", '.', "Ha'u", 'iha', '$', '0.25', 'atu', 'fó', 'ba', 'ita', '.']
- Using
TetunSentenceTokenizerto tokenize the input text.
from tetuntokenizer.tokenizer import TetunSentenceTokenizer
tetun_tokenizer = TetunSentenceTokenizer()
text = "Ha'u ema-ida ne'ebé baibain de'it. Tebes ga? Ita-nia maluk Dr. ka Ph.D sira husi U.S.A mós dehan!"
output = tetun_tokenizer.tokenize(text)
print(output)
Expected output:
["Ha'u ema-ida ne'ebé baibain de'it.", 'Tebes ga?', 'Ita-nia maluk Dr. ka Ph.D sira husi U.S.A mós dehan!']
- Utilizing
TetunBlankLineTokenizerto tokenize the input text.
from tetuntokenizer.tokenizer import TetunBlankLineTokenizer
tetun_tokenizer = TetunBlankLineTokenizer()
text = """
Ha'u mak ita-nia maluk di'ak.
Ha'u iha $0.25 atu fó ba ita.
"""
output = tetun_tokenizer.tokenize(text)
print(output)
Expected output:
["\n Ha'u mak ita-nia maluk di'ak.\n Ha'u iha $0.25 atu fó ba ita.\n "]
- Using
TetunSimpleTokenizerto tokenize a given text.
from tetuntokenizer.tokenizer import TetunSimpleTokenizer
tetun_tokenizer = TetunSimpleTokenizer()
text = "Ha'u mak ita-nia maluk di'ak. Ha'u iha $0.25 atu fó ba ita."
output = tetun_tokenizer.tokenize(text)
print(output)
Expected output:
["Ha'u", 'mak', 'ita-nia', 'maluk', "di'ak", "Ha'u", 'iha', '0.25', 'atu', 'fó', 'ba', 'ita']
- Using
TetunWordTokenizerto tokenize the input text.
from tetuntokenizer.tokenizer import TetunWordTokenizer
tetun_tokenizer = TetunWordTokenizer()
text = "Ha'u mak ita-nia maluk di'ak. Ha'u iha $0.25 atu fó ba ita."
output = tetun_tokenizer.tokenize(text)
print(output)
Expected output:
["Ha'u", 'mak', 'ita-nia', 'maluk', "di'ak", "Ha'u", 'iha', 'atu', 'fó', 'ba', 'ita']
Citation
If you use this repository or any of its contents for your research, academic work, or publication, we kindly request that you cite it as follows:
@inproceedings{de-jesus-nunes-2024-labadain-crawler,
title = "Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus",
author = "de Jesus, Gabriel and
Nunes, S{\'e}rgio Sobral",
editor = "Calzolari, Nicoletta and
Kan, Min-Yen and
Hoste, Veronique and
Lenci, Alessandro and
Sakti, Sakriani and
Xue, Nianwen",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.390",
pages = "4368--4380"
}
Acknowledgement
This work is financed by National Funds through the Portuguese funding agency, FCT - Fundação para a Ciência e a Tecnologia under the PhD scholarship grant number SFRH/BD/151437/2021 (DOI 10.54499/SFRH/BD/151437/2021).
License
Additional Information
For the source code, visit the GitHub repository for this project.
Release files for tetun-tokenizer 1.2.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tetun_tokenizer-1.2.3.tar.gz | 6.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tetun_tokenizer-1.2.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 6.0 MB
Release files / tetun_tokenizer-1.2.3.tar.gz
| Download URL | tetun_tokenizer-1.2.3.tar.gz |
|---|---|
| Size | 6.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a1b68836b069e22909963dedaf576f8dbdae3727d7791c77115fda66b40837f0
|
|
BLAKE2b-256 checksum How to use checksums |
631a1703483fb515b676a7fa10923098aacff5b07f4c94f83ad215eca5cd1e3b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.0 CPython/3.9.7
|
Release files / tetun_tokenizer-1.2.3-py3-none-any.whl
| Download URL | tetun_tokenizer-1.2.3-py3-none-any.whl |
|---|---|
| Size | 5.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
125d3d69d03507154b71b630374016b08a339f578d0e1629f4eaf5965b92a951
|
|
BLAKE2b-256 checksum How to use checksums |
3d054c4c9cc25177f94dd3f746f2161dcb4fc14fa587b47d9fe277d95f81bee0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.0 CPython/3.9.7
|