py3langid is a fork of the standalone language identification tool langid.py by Marco Lui.
Original license: BSD-2-Clause. Fork license: BSD-3-Clause.
Changes in this fork
Execution speed has been improved and the code base has been modernized for Python 3.10+:
Import: Loading the package (import py3langid) is about 25% faster
Execution: Language detection with langid.classify is 10x faster on single sentences and 3-4x faster on paragraphs (less on longer texts, about 1.4x at 100 kB)
Startup: Loading the default classification model is 2-3x faster, with a model six times larger
For implementation details see this blog post: How to make language detection with langid.py faster.
The fork also ships a retrained model covering 139 languages (up from 97) and a fully rewritten, reproducible training pipeline (see Training a model).
For version history see the changelog.
Usage
Install: pip install py3langid — use as import py3langid as langid or on the command-line as langid.
With Python
>>> import py3langid as langid
>>> langid.classify('This text is in English.')
('en', -68.562286)
>>> langid.rank('This text is in English.') # all languages, most likely first
>>> from py3langid.langid import LanguageIdentifier, MODEL_FILE
>>> identifier = LanguageIdentifier.from_model_file(MODEL_FILE, norm_probs=True)
>>> identifier.set_languages(['de', 'en', 'fr'])
>>> identifier.classify('This should be enough text.')
('en', 0.9999628)
# abstention: return ('und', confidence) below a threshold
>>> identifier = LanguageIdentifier.from_model_file(MODEL_FILE, norm_probs=True,
... min_confidence=0.2)
>>> identifier.classify('ok')
('und', 0.0140845)
Input can be str or UTF-8 bytes; input is NFC-normalized before classification, and all-uppercase text is case-folded.
On the command-line
# basic usage with probability normalization
$ echo "This should be enough text." | langid -n
('en', 0.9935992)
# define a subset of target languages
$ echo "This won't be recognized properly." | langid -n -l fr,it,tr
('fr', 0.4838270)
Run langid without input to get an interactive prompt, pipe text into it to classify a whole document, or add --line to classify each line separately. langid -u URL downloads and classifies a web page. See langid --help for all options.
Languages
The shipped model knows 139 languages plus zxx (ISO 639 codes):
ace, af, am, an, ar, ary, arz, as, az, ba, bcl, be, bg, bn, br, bs, ca, crh, cs, cy, da, de, dz, el, en, eo, es, et, eu, ext, fa, fi, fo, fr, fuv, fy, ga, gcf, gcr, gd, gl, gom, grc, gu, gug, guw, ha, hbo, he, hi, hr, ht, hu, hy, id, ig, is, it, ja, jv, ka, kab, kik, kk, km, kn, ko, ku, ky, la, lb, lg, lij, ln, lo, lt, ltg, lv, mg, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, nso, oc, om, or, pa, pcm, pl, ps, pt, qu, ro, ru, rw, sa, sdh, se, si, sk, sl, sn, so, sq, sr, st, sv, sw, ta, te, tg, th, tk, tl, tr, tt, ug, uk, ur, uz, uzs, vec, vi, vo, wa, wuu, xh, yo, yue, zh, zu, zxx
zxx is a synthetic “not a language” class that catches numbers, markup, identifiers, and similar non-linguistic content. With min_confidence set, low-confidence predictions are returned as und (undetermined).
Batch mode
langid -b reads file paths from stdin (one per line) and classifies the files in parallel, writing CSV to stdout:
$ find corpus -name "*.txt" | langid -b
corpus/a.txt,en,-127.32
corpus/b.txt,de,-81.15
With -d, the output is one CSV row per file with the full score distribution over all languages (one column per language).
Web service
langid -s serves language identification over HTTP (default port 9008). Use the langid console script; python -m py3langid.langid is not supported. Endpoints /detect and /rank accept GET, POST, and PUT:
$ curl -d "q=This is a test" localhost:9008/detect
For production, use py3langid.server:application under a WSGI server.
Custom models
langid -m FILE (or LanguageIdentifier.from_modelpath(path)) loads a model trained with this package (model.npz.xz). Models from the original langid.py are not supported.
Training a model
python -m py3langid.train.train -m model_dir corpus_dir, run from a clone of the repository (the training code is not part of the PyPI package) — see TRAINING.md for corpus layout, data gathering, hygiene, and pipeline design. Training is deterministic: the same corpus and settings reproduce the model byte for byte.
Read more
Release files for py3langid 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| py3langid-0.4.0.tar.gz | 4.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| py3langid-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 9.2 MB
Release files / py3langid-0.4.0.tar.gz
| Download URL | py3langid-0.4.0.tar.gz |
|---|---|
| Size | 4.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5924159039b5cc282c2edc10a7ea9363b2099834a0e99ed8d5bea3cc435d7871
|
|
BLAKE2b-256 checksum How to use checksums |
16fc51c88e5ef8346878ae6c09e1f7c86c6e6308c57794723bf423a72ddba4dd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.1
|
Release files / py3langid-0.4.0-py3-none-any.whl
| Download URL | py3langid-0.4.0-py3-none-any.whl |
|---|---|
| Size | 4.6 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c25bd038147951ba3a89efe9ec12cc7d2116cb6f311ff8af84ade398a8bc0616
|
|
BLAKE2b-256 checksum How to use checksums |
692d0eeff2727c970b1d553be70d15ba2d84e3830696c09a79d954c5ae3d5212
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.1
|