Skip to main content

py3langid is a fork of the standalone language identification tool langid.py by Marco Lui.

Original license: BSD-2-Clause. Fork license: BSD-3-Clause.

Changes in this fork

Execution speed has been improved and the code base has been modernized for Python 3.10+:

  • Import: Loading the package (import py3langid) is about 25% faster

  • Execution: Language detection with langid.classify is 10x faster on single sentences and 3-4x faster on paragraphs (less on longer texts, about 1.4x at 100 kB)

  • Startup: Loading the default classification model is 2-3x faster, with a model six times larger

For implementation details see this blog post: How to make language detection with langid.py faster.

The fork also ships a retrained model covering 139 languages (up from 97) and a fully rewritten, reproducible training pipeline (see Training a model).

For version history see the changelog.

Usage

Install: pip install py3langid — use as import py3langid as langid or on the command-line as langid.

With Python

>>> import py3langid as langid

>>> langid.classify('This text is in English.')
('en', -68.562286)
>>> langid.rank('This text is in English.')   # all languages, most likely first

>>> from py3langid.langid import LanguageIdentifier, MODEL_FILE
>>> identifier = LanguageIdentifier.from_model_file(MODEL_FILE, norm_probs=True)
>>> identifier.set_languages(['de', 'en', 'fr'])
>>> identifier.classify('This should be enough text.')
('en', 0.9999628)

# abstention: return ('und', confidence) below a threshold
>>> identifier = LanguageIdentifier.from_model_file(MODEL_FILE, norm_probs=True,
...                                                 min_confidence=0.2)
>>> identifier.classify('ok')
('und', 0.0140845)

Input can be str or UTF-8 bytes; input is NFC-normalized before classification, and all-uppercase text is case-folded.

On the command-line

# basic usage with probability normalization
$ echo "This should be enough text." | langid -n
('en', 0.9935992)

# define a subset of target languages
$ echo "This won't be recognized properly." | langid -n -l fr,it,tr
('fr', 0.4838270)

Run langid without input to get an interactive prompt, pipe text into it to classify a whole document, or add --line to classify each line separately. langid -u URL downloads and classifies a web page. See langid --help for all options.

Languages

The shipped model knows 139 languages plus zxx (ISO 639 codes):

ace, af, am, an, ar, ary, arz, as, az, ba, bcl, be, bg, bn, br, bs, ca,
crh, cs, cy, da, de, dz, el, en, eo, es, et, eu, ext, fa, fi, fo, fr,
fuv, fy, ga, gcf, gcr, gd, gl, gom, grc, gu, gug, guw, ha, hbo, he, hi,
hr, ht, hu, hy, id, ig, is, it, ja, jv, ka, kab, kik, kk, km, kn, ko,
ku, ky, la, lb, lg, lij, ln, lo, lt, ltg, lv, mg, mk, ml, mn, mr, ms,
mt, my, ne, nl, nn, no, nso, oc, om, or, pa, pcm, pl, ps, pt, qu, ro,
ru, rw, sa, sdh, se, si, sk, sl, sn, so, sq, sr, st, sv, sw, ta, te, tg,
th, tk, tl, tr, tt, ug, uk, ur, uz, uzs, vec, vi, vo, wa, wuu, xh, yo,
yue, zh, zu, zxx

zxx is a synthetic “not a language” class that catches numbers, markup, identifiers, and similar non-linguistic content. With min_confidence set, low-confidence predictions are returned as und (undetermined).

Batch mode

langid -b reads file paths from stdin (one per line) and classifies the files in parallel, writing CSV to stdout:

$ find corpus -name "*.txt" | langid -b
corpus/a.txt,en,-127.32
corpus/b.txt,de,-81.15

With -d, the output is one CSV row per file with the full score distribution over all languages (one column per language).

Web service

langid -s serves language identification over HTTP (default port 9008). Use the langid console script; python -m py3langid.langid is not supported. Endpoints /detect and /rank accept GET, POST, and PUT:

$ curl -d "q=This is a test" localhost:9008/detect

For production, use py3langid.server:application under a WSGI server.

Custom models

langid -m FILE (or LanguageIdentifier.from_modelpath(path)) loads a model trained with this package (model.npz.xz). Models from the original langid.py are not supported.

Training a model

python -m py3langid.train.train -m model_dir corpus_dir, run from a clone of the repository (the training code is not part of the PyPI package) — see TRAINING.md for corpus layout, data gathering, hygiene, and pipeline design. Training is deterministic: the same corpus and settings reproduce the model byte for byte.

Read more

[1] Lui & Baldwin (2011) Cross-domain Feature Selection for Language Identification, IJCNLP 2011.
[2] Lui & Baldwin (2012) langid.py: An Off-the-shelf Language Identification Tool, ACL 2012 Demo.

Release files for py3langid 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for py3langid 0.4.0
File Size Uploaded
py3langid-0.4.0.tar.gz 4.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for py3langid 0.4.0
File Interpreter ABI Platform
py3langid-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 9.2 MB

Release files / py3langid-0.4.0.tar.gz

Download URL py3langid-0.4.0.tar.gz
Size 4.6 MB
Tags Source
SHA-256 checksum
How to use checksums
5924159039b5cc282c2edc10a7ea9363b2099834a0e99ed8d5bea3cc435d7871
BLAKE2b-256 checksum
How to use checksums
16fc51c88e5ef8346878ae6c09e1f7c86c6e6308c57794723bf423a72ddba4dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.1

Release files / py3langid-0.4.0-py3-none-any.whl

Download URL py3langid-0.4.0-py3-none-any.whl
Size 4.6 MB
Tags Python 3
SHA-256 checksum
How to use checksums
c25bd038147951ba3a89efe9ec12cc7d2116cb6f311ff8af84ade398a8bc0616
BLAKE2b-256 checksum
How to use checksums
692d0eeff2727c970b1d553be70d15ba2d84e3830696c09a79d954c5ae3d5212
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.1

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page