Language Detection AI
A classical machine-learning language detector: feed it a sentence, it tells you which language it is written in. Built with Python + scikit-learn using character n-grams + Naive Bayes / Linear SVC.
Includes detection of Roman-script code-mixed Indian languages (Tanglish, Hinglish, Teluglish, Kanglish, Manglish).
Install & run the chatbot UI
The UI ships as a pip-installable package (langid-chat) that bundles the
trained model, so anyone can run it with a single command — no repo clone
required.
pip install langid-chat
# Launch the chat UI (http://127.0.0.1:8000)
langid-web
# Or use the CLI
langid "Kya kar rahe ho bhai"
langid --interactive
langid-web supports --host and --port (e.g. langid-web --port 9000).
To install from a source checkout instead:
pip install -e .
langid-web
# or without installing at all:
python web/app.py
Quick start
# 0. (Optional) Re-fetch the code-mixed corpora into data/raw (cached by default)
python src/fetch_codemix.py
# 1. Prepare data -- the deployed model is trained on the FULL dataset with
# balanced class weights (keeps all code-mixed signal):
python src/preprocess.py --no-balance
python src/train.py --class-weight balanced
# 2. Test new sentences
python detect.py "Kya kar rahe ho bhai"
python detect.py --interactive
On Windows set
$env:PYTHONIOENCODING='utf-8'first so multilingual text prints correctly.
detect.py also reports a confidence score per guess. Below a calibrated
threshold it says LOW CONFIDENCE instead of guessing — short or ambiguous
input (e.g. ok, vanakkam) gets flagged rather than answered wrongly.
Tune with --min-confidence 0.5, or force an always-guess with 0.
Languages supported (15)
Tamil, Russian, Arabic (distinct scripts) · Spanish, Portuguese, French, Italian (linguistically close) · English, German, Dutch · plus code-mixed: Tanglish, Hinglish, Teluglish, Kanglish, Manglish.
Results
Best model (LinearSVC, full data + balanced class weights) reaches ~98% test accuracy across the 15 classes; script-distinct languages (Arabic, Russian, Tamil) hit ~100%. High-confidence calls (>= 45%) are ~99% correct on the test split; the only reliable misdetections left are 1–2 word Roman-script fragments, which are exactly the ones the confidence gate flags.
Limitations: it only knows the 15 trained languages and still guesses on anything fed to it without confidences if you lower the threshold. Very short text is inherently unreliable and is flagged, not silently misanswered.
See results/results.md for the full write-up and project.md/tasks.md for
scope and tasks.
Metadata
Release files for langid-chat 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langid_chat-0.1.0.tar.gz | 9.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langid_chat-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.1 MB
Release files / langid_chat-0.1.0.tar.gz
| Download URL | langid_chat-0.1.0.tar.gz |
|---|---|
| Size | 9.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cccd6362132b70b13af384c45e8bcce3a8ea9ecc7e7f4d888c4a002f7eef0d55
|
|
BLAKE2b-256 checksum How to use checksums |
01d9d9850f2920b80bff218a916d97c85073cdfc632d67b354284114ef5c899c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / langid_chat-0.1.0-py3-none-any.whl
| Download URL | langid_chat-0.1.0-py3-none-any.whl |
|---|---|
| Size | 9.0 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a9a6f4217eccb236d14bb93ab5de0e176ef24790780fc1b645a263786c911c1a
|
|
BLAKE2b-256 checksum How to use checksums |
06ef56907336a707ff1fd1a6b8ac57b7582a59f6c8027a021948aecd7c9b769d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|