Skip to main content

A language identifier for Indian languages using a pretrained Random Forest model.

Project description

langid_indian

A language identifier for Indian languages using a pretrained Machine Learning model.

PyPI version


📌 Overview

langid_indian is a Python package for language identification of Indian languages written in native scripts. It is based on a pretrained Machine Learning model and supports multiple languages listed in the Eighth Schedule of the Indian Constitution.

The library is designed to be:

  • Lightweight and easy to use
  • Suitable for NLP pipelines, chatbots, multilingual datasets, and research
  • Minimal in preprocessing requirements

Recommended citation title (as suggested): "Language Identification using langid_indian library."


🔗 Project Links


📦 Installation

Using Terminal / Command Prompt

pip install langid_indian

Using Jupyter Notebook / Google Colab / Kaggle

!pip install langid_indian

🚀 Quick Start

import langid_indian

# Initialize the language identifier
lid = langid_indian.LanguageIdentifier()

# Predict the language of a single sentence
print(lid.predict("जब मैं छोटा था, मैं हर रोज़ पार्क जाता था।"))
# Output: hin

🔍 Predicting Multiple Sentences

test_texts = [
    "Hello, how are you?",
    "जब मैं छोटा था, मैं हर रोज़ पार्क जाता था।",
    "आनी हो एक गंभीर मूर्खपणा.",
    "ਮਨੁੱਖੀ ਦਿਮਾਗ਼ ਦੀ ਕਾਢ ਨੇ ਭਾਵੇਂ ਸਭ ਕੁਝ ਸੌਖਾ ਕਰ ਦਿੱਤਾ ਹੈ ਪਰ ਫਿਰ ਵੀ ਸਭ ਕੁਝ ਸਮਝਣਾ ਜਾਂ ਕਰਨਾ ਨਿਯਮਾਂ ਵਿੱਚ ਬੱਝਾ ਪਿਆ ਹੈ।",
    "માં વરસાદનું પાણી મોટા જથ્થામાં જમીનની નીચે જ ઉતરી જાય છે।",
    "କିନ୍ତୁ ପୁଅ, ତୁମେ ଛୋଟ।"
]

for text in test_texts:
    print(f"Text: {text} | Detected language: {lid.predict(text)}")

Sample Output

Text: Hello, how are you? | Detected language: eng
Text: जब मैं छोटा था, मैं हर रोज़ पार्क जाता था। | Detected language: hin
Text: आनी हो एक गंभीर मूर्खपणा. | Detected language: gom
Text: ਮਨੁੱਖੀ ਦਿਮਾਗ਼ ਦੀ ਕਾਢ ਨੇ ਭਾਵੇਂ ਸਭ ਕੁਝ ਸੌਖਾ ਕਰ ਦਿੱਤਾ ਹੈ ਪਰ ਫਿਰ ਵੀ ਸਭ ਕੁਝ ਸਮਝਣਾ ਜਾਂ ਕਰਨਾ ਨਿਯਮਾਂ ਵਿੱਚ ਬੱਝਾ ਪਿਆ ਹੈ। | Detected language: pan
Text: માં વરસાદનું પાણી મોટા જથ્થામાં જમીનની નીચે જ ઉતરી જાય છે। | Detected language: guj
Text: କିନ୍ତୁ ପୁଅ, ତୁମେ ଛୋଟ। | Detected language: ory

🧪 Testing the Package

The repository includes a basic test file located at:

tests/test_basic.py

test_basic.py

from langid_indian import LanguageIdentifier

def test_prediction():
    identifier = LanguageIdentifier()
    text = "यह एक परीक्षण है"
    lang = identifier.predict(text)
    print(f"Predicted language: {lang}")

if __name__ == "__main__":
    test_prediction()

Run the Test

From the project root directory:

python -m tests.test_basic

🌐 Supported Languages

The model supports 22 Indian languages:

  • Assamese (অসমীয়া)
  • Bengali (বাংলা)
  • Bodo
  • Dogri
  • English
  • Gujarati (ગુજરાતી)
  • Hindi (हिन्दी)
  • Kannada (ಕನ್ನಡ)
  • Kashmiri
  • Konkani
  • Maithili
  • Malayalam (മലയാളം)
  • Manipuri
  • Marathi (मराठी)
  • Nepali
  • Odia (ଓଡ଼ିଆ)
  • Punjabi (ਪੰਜਾਬੀ)
  • Santali
  • Sindhi
  • Tamil (தமிழ்)
  • Telugu (తెలుగు)
  • Urdu (اردو)

📄 Citation

If you use langid_indian in academic or research work, please cite the following paper:

BibTeX

@misc{ingle2025ilidnativescriptlanguage,
  title={ILID: Native Script Language Identification for Indian Languages},
  author={Yash Ingle and Pruthwik Mishra},
  year={2025},
  eprint={2507.11832},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2507.11832}
}

👤 Author

Yash Ingle B.Tech (AI), SVNIT Surat


📜 License

This project is licensed under the MIT License.


⭐ Acknowledgement

This work is part of ongoing research on Indian Language Identification using Machine Learning and Native Script Text, developed under academic guidance.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

langid_indian-0.1.4.tar.gz (5.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

langid_indian-0.1.4-py3-none-any.whl (5.8 kB view details)

Uploaded Python 3

File details

Details for the file langid_indian-0.1.4.tar.gz.

File metadata

  • Download URL: langid_indian-0.1.4.tar.gz
  • Upload date:
  • Size: 5.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.5

File hashes

Hashes for langid_indian-0.1.4.tar.gz
Algorithm Hash digest
SHA256 121f4e29de355debc3a5fc34b13a29cee01a45d93fc77320782149b5a371ee64
MD5 578bb4fd163f63f5fd203c075326848c
BLAKE2b-256 1b6afac417821f48fb7a6e7e44c37fc974317b1193902d9938c6a92ed87d49cb

See more details on using hashes here.

File details

Details for the file langid_indian-0.1.4-py3-none-any.whl.

File metadata

  • Download URL: langid_indian-0.1.4-py3-none-any.whl
  • Upload date:
  • Size: 5.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.5

File hashes

Hashes for langid_indian-0.1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 5725d85eb1aaf85849fb4173ad79b361d5fd34bc5a1d661c9e9c40316981bdce
MD5 68218c2d0f55798f92c2b028b64b83d7
BLAKE2b-256 6b8325d3eaf1f88f9108507e72cbc3033451bb7ffa33063efd94c46749c13b58

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page