A Natural Language Processing toolkit for the Nagamese language.
Project description
NagaNLP: Natural Language Processing for Nagamese
NagaNLP is a comprehensive Natural Language Processing toolkit for the Nagamese language, featuring state-of-the-art models for various NLP tasks. It provides simple and efficient tools for working with Nagamese text data.
Features
- Part-of-Speech Tagging: Transformer-based and NLTK-based POS taggers
- Neural Machine Translation: Seq2Seq model for Nagamese to English translation
- Named Entity Recognition: Pre-trained BERT models for entity recognition
- Word Alignment: Tools for parallel corpus alignment
- Subword Tokenization: Support for handling out-of-vocabulary words
- Easy Integration: Simple Python API for all functionalities
- Pre-trained Models: Ready-to-use models for quick deployment
- Extensible: Easy to customize and extend for specific use cases
Installation
Prerequisites
- Python 3.8 or higher
- pip (Python package manager)
Install from PyPI (recommended)
pip install naganlp
Install from source
git clone https://github.com/AgnivaMaiti/naga-nlp.git
cd naga-nlp
pip install -e .
Install with optional dependencies
# For development with testing and documentation
pip install "naganlp[dev]"
# For using GPU acceleration (requires CUDA-compatible GPU)
pip install "naganlp[gpu]"
Pre-trained Models
All models are available on Hugging Face:
- Machine Translation: agnivamaiti/naganlp-nmt-en
- Named Entity Recognition: agnivamaiti/naganlp-ner-crf-tagger
- POS Tagger: agnivamaiti/naganlp-pos-tagger
Quick Start
Part-of-Speech Tagging
Using the Transformer-based Tagger (Recommended)
from naganlp import PosTagger
# Initialize the tagger (automatically downloads the model on first use)
tagger = PosTagger()
# Tag a sentence
sentence = "moi school te jai" # "I go to school" in Nagamese
result = tagger.tag(sentence)
# Print the results
for token in result:
print(f"{token['word']}: {token['entity_group']}")
Using the NLTK-based Tagger (Lightweight)
from naganlp import NltkPosTagger
# Load the pre-trained NLTK model
tagger = NltkPosTagger("naga_pos_model.pkl")
# Tag a list of tokens
result = tagger.predict(["moi", "school", "te", "jai"])
print(result)
# Output: [('moi', 'PRON'), ('school', 'NOUN'), ('te', 'ADP'), ('jai', 'VERB')]
Machine Translation
from naganlp import Translator
# Initialize the translator (uses 'agnivamaiti/naganlp-nmt-en' by default)
translator = Translator()
# Translate from Nagamese to English
translation = translator.translate("moi school te jai")
print(f"Translation: {translation}")
# Output: "I go to school"
# Get translation with token IDs
translation, token_ids = translator.translate("tumi kiman din ahiba?", return_tokens=True)
print(f"Translation: {translation}")
print(f"Token IDs: {token_ids}")
Named Entity Recognition
from naganlp import NERTagger
# Initialize the NER tagger with the pre-trained model
ner_tagger = NERTagger(model_id="agnivamaiti/naganlp-ner-crf-tagger")
# Extract named entities from text
entities = ner_tagger.tag("Agniva is going to Guwahati tomorrow")
for entity in entities:
print(f"{entity['word']}: {entity['entity_group']} (confidence: {entity['score']:.2f})")
# Training a new NER model (example)
# python main.py train-ner --data-file path/to/ner_data.json --hub-id your-username/naganlp-ner-crf-tagger
Word Alignment
from naganlp import align_parallel_texts
# Align parallel sentences
df = align_parallel_texts(
source_texts=["moi school jai", "tumi kiman din te ahibo?"],
target_texts=["I go to school", "In how many days will you come?"],
config={"alignment_method": "fast_align"} # or 'simalign' for better quality
)
print(df[['source', 'target', 'alignment']])
Subword Tokenization
from naganlp import SubwordTokenizer
# Initialize tokenizer with a pre-trained model
tokenizer = SubwordTokenizer()
# Tokenize a sentence
tokens = tokenizer.tokenize("moi school te jai")
print(f"Tokens: {tokens}")
# Convert tokens to IDs
ids = tokenizer.encode("moi school te jai")
print(f"Token IDs: {ids}")
API Reference
PosTagger
class PosTagger(model_name: str = 'agnivamaiti/naganlp-pos-tagger', use_nltk: bool = False)
A part-of-speech tagger for Nagamese text.
Parameters:
model_name: Hugging Face model ID or path to local modeluse_nltk: If True, uses NLTK-based tagger instead of transformer
Methods:
tag(text: Union[str, List[str]]): Tag a single sentence or list of sentences__call__: Alias fortag
NltkPosTagger
class NltkPosTagger(model_path: str)
Lightweight POS tagger using NLTK's CRF implementation.
Parameters:
model_path: Path to the trained NLTK model file
Methods:
predict(tokens: List[str]): Tag a list of tokensevaluate(test_data): Evaluate the model on test data
Translator
class Translator(model_id: str = 'agnivamaiti/naganlp-nmt-en', device: str = None)
Neural Machine Translation model for Nagamese to English.
Parameters:
model_id: Hugging Face model ID or path to local modeldevice: Device to run the model on ('cuda' or 'cpu')
Methods:
translate(text: str, beam_size: int = 5, max_len: int = 50, length_penalty: float = 0.7, return_tokens: bool = False): Translate textbatch_translate(texts: List[str], **kwargs): Translate a batch of texts
NERTagger
class NERTagger(model_name: str = 'agnivamaiti/naganlp-ner-crf-tagger')
Named Entity Recognition for Nagamese text.
Parameters:
model_name: Hugging Face model ID or path to local model
Methods:
tag(text: str): Extract named entities from textbatch_tag(texts: List[str]): Process multiple texts in a batch
SubwordTokenizer
class SubwordTokenizer(model_path: str = None, vocab_size: int = 8000)
Subword tokenizer for Nagamese text.
Parameters:
model_path: Path to SentencePiece modelvocab_size: Vocabulary size (only used when training a new model)
Methods:
tokenize(text: str): Tokenize text into subwordsencode(text: str): Convert text to token IDsdecode(ids: List[int]): Convert token IDs back to texttrain(input_file: str, model_prefix: str): Train a new tokenizer
Pre-trained Models
NagaNLP provides several pre-trained models available on Hugging Face Hub:
-
Machine Translation (Nagamese to English)
- Model ID:
agnivamaiti/naganlp-nmt-en - Type: Seq2Seq Transformer
- Training Data: Parallel Nagamese-English corpus
- Usage:
translator = Translator(model_id="agnivamaiti/naganlp-nmt-en")
- Model ID:
-
Named Entity Recognition
- Model ID:
agnivamaiti/naganlp-ner-crf-tagger - Type: CRF-based NER Tagger
- Training Data: Annotated Nagamese text
- Usage:
ner_tagger = NERTagger(model_id="agnivamaiti/naganlp-ner-crf-tagger")
- Model ID:
-
Part-of-Speech Tagger
- Model ID:
agnivamaiti/naganlp-pos-tagger - Type: Transformer-based POS Tagger
- Training Data: An annotated Nagamese text
- Usage:
pos_tagger = PosTagger(model_id="agnivamaiti/naganlp-pos-tagger")
- Model ID:
Command Line Interface
{{ ... }} NagaNLP provides a convenient CLI for common tasks:
# Tag text with POS tags
naganlp pos-tag --text "moi school te jai"
# Translate text
naganlp translate --text "tumi kiman din ahiba?" --target-lang en
# Train a new model
naganlp train-pos-tagger --train-file train.conll --output-dir ./model
# Evaluate a model
naganlp evaluate --model agnivamaiti/naga-pos-tagger --eval-file test.conll
Contributing
Contributions are welcome! Please read our Contributing Guidelines for details on how to contribute to the project.
Setting up Development Environment
-
Clone the repository:
git clone https://github.com/AgnivaMaiti/naga-nlp.git cd naga-nlp
-
Install development dependencies:
pip install -e ".[dev]"
-
Run tests:
pytest tests/ -
Run code formatting and linting:
black . flake8 .
License
This project is licensed under the MIT License - see the LICENSE file for details.
Citation
If you use NagaNLP in your research, please cite it as follows:
@software{naganlp2023,
author = {Agniva Maiti},
title = {NagaNLP: Natural Language Processing for Nagamese},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/AgnivaMaiti/naga-nlp}}
}
Contact
For questions, suggestions, or support:
- Open an issue on GitHub Issues
- Email: Agniva Maiti
- Twitter: @AgnivaMaiti
from naganlp import NltkPosTagger
# First train and save the model (only needed once)
from naganlp.nltk_tagger import train_and_save_nltk_tagger
train_and_save_nltk_tagger("path/to/your/conll/file.conll", "naga_pos_model.pkl")
# Then load and use the trained model
tagger = NltkPosTagger("naga_pos_model.pkl")
# Tag a list of pre-tokenized words
result = tagger.predict(["moi", "school", "te", "jai"])
print(result)
# Output: [('moi', 'PRON'), ('school', 'NOUN'), ('te', 'ADP'), ('jai', 'VERB')]
Translation
from naganlp import Translator
# Initialize the translator with the pre-trained model from Hugging Face
translator = Translator(model_id="agnivamaiti/naganlp-nmt-en")
# Translate from Nagamese to English
translation = translator.translate("moi school te jai")
print(translation)
# Output: "I go to school"
Documentation
Data Requirements
- For POS Tagging: CONLL-formatted file with token and POS tag columns
- For Translation: Parallel corpus in CSV format with 'nagamese' and 'english' columns
Model Training
POS Tagger Training
# Training a new POS tagger
python main.py train-tagger --conll-file path/to/train.conll --hub-id agnivamaiti/naganlp-pos-tagger
# Using the pre-trained model
from naganlp import PosTagger
pos_tagger = PosTagger(model_id="agnivamaiti/naganlp-pos-tagger")
NMT Model Training
# Training a new NMT model
python main.py train-translator --data-file path/to/parallel_corpus.csv --hub-id agnivamaiti/naganlp-nmt-en
# Using the pre-trained translation model
from naganlp import Translator
translator = Translator(model_id="agnivamaiti/naganlp-nmt-en")
Advanced Usage
Custom Model Paths
# Load custom models
custom_tagger = PosTagger(model_name_or_path="path/to/custom/model")
custom_translator = Translator(model_path="path/to/translator.pt", vocabs_path="path/to/vocabs.pkl")
## Contributing
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.
## Code of Conduct
This project follows a [Code of Conduct](CODE_OF_CONDUCT.md).
## License
This project is licensed under the MIT License - see the [LICENSE](https://github.com/AgnivaMaiti/naga-nlp/blob/main/LICENSE) file for details.
## Contact
- Agniva Maiti
- Email: agnivamaiti.official@gmail.com
- LinkedIn: [Agniva Maiti](https://linkedin.com/in/agniva-maiti)
## Acknowledgments
- KIIT University for the support and resources
- All contributors and users of this library
## Citation
If you use NagaNLP in your research, please cite:
```bibtex
@software{naganlp,
title={NagaNLP: Natural Language Processing Toolkit for Nagamese},
author={Agniva Maiti},
year={2025},
publisher={GitHub},
journal={GitHub repository},
howpublished={\url{https://github.com/AgnivaMaiti/naga-nlp}}
}
Support
For questions and support, please open an issue on our GitHub repository.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file naganlp-0.1.2.tar.gz.
File metadata
- Download URL: naganlp-0.1.2.tar.gz
- Upload date:
- Size: 537.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
db1d551c5ae44eb3fc4ad56676ac3d879f39c958d4826a533634e144146b6f28
|
|
| MD5 |
1b4fccfaa72fd128e655576c71029f37
|
|
| BLAKE2b-256 |
422cc4d643673820c8c036586e24412ffeaeabf17529e207b00b562559680141
|
File details
Details for the file naganlp-0.1.2-py3-none-any.whl.
File metadata
- Download URL: naganlp-0.1.2-py3-none-any.whl
- Upload date:
- Size: 195.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d2ba103d8b0ed3462f680fb370261b55d1fdf7662b20cd9abe345b508e6840ab
|
|
| MD5 |
190578f7cd51a23ca07cbc1ebb99a635
|
|
| BLAKE2b-256 |
b5c45b4ce7ca88fae7d842fbfae9111a56ee828ad3e3da8867db439449a7dff9
|