from noentenc import Language
from noentenc.language_detection import LanguageDetector
from noentenc.translation import Translator
LanguageDetector().detect("Bon dia! Com estàs?")
# 'cat'
Translator().translate("The weather is nice today.", Language.SPANISH, Language.ENGLISH)
# 'El tiempo es bueno hoy.'
- Fast. Detects a language in as little as 2.5 µs and translates a sentence in about 52 ms, on a laptop CPU. See benchmarks.
- Faster than the originals. Up to 470× faster than
langid.py, 2× faster than the C++fasttextpackage and 5.8× faster than transformers + PyTorch, with the same accuracy. See the comparison. - Small. The default detection model is 0.9 MB, and detection needs only numpy and tqdm. Translation adds onnxruntime, tokenizers and huggingface-hub. There is no torch, transformers or GPU dependency.
- One API, many models. 12 detection models and 4 translation model families sit behind the same two classes. Swapping one is a one-line change, and every detector returns the same ISO 639-3 labels.
- Built for datasets. You can pass a single string, a list or a pandas or polars column.
New to noentenc? The quickstart covers detection, translation and choosing a model on one page.
Install
noentenc uses uv to manage dependencies.
Requires Python 3.11 or newer, on Linux, macOS or Windows. Install with pip or uv:
uv add noentenc # detection only: fastText and langid, with numpy
uv add 'noentenc[translation]' # detection and every translation model
uv add 'noentenc[onnx]' # the ONNX detection models (bert-openlid, xlm-roberta-lid)
uv add 'noentenc[lingua,cld3]' # extra detection backends, see the table below
uv add 'noentenc[all]' # everything
With pip, the same extras work: pip install 'noentenc[translation]'.
Python 3.10 isn't supported: onnxruntime 1.30, pandas 3 and the current numpy releases have no Python 3.10 wheels.
Weights download on first use, to ~/.cache/noentenc (NOENTENC_CACHE or cache_dir move it). Nothing is bundled in the wheel. To run offline, download them ahead of time with noentenc.prepare(...) and pass only_local_files=True; see Run offline.
Detect a language
from noentenc.language_detection import LanguageDetector
detector = LanguageDetector() # fastText lid.176: 176 languages, 0.9 MB
detector.detect("Bon dia! Com estàs?")
# 'cat'
detector.detect("Bon dia! Com estàs?", with_score=True, top_k=2)
# {'cat': 0.88, 'por': 0.11}
detector.detect_batch(["Hello there", "Hola, ¿qué tal?", "你好", "👍", ""])
# ['eng', 'spa', 'zho', 'zxx', 'und']
# Add a "lang" column to a pandas or polars DataFrame.
detector.detect_dataset(df, "text", "lang")
Text without letters outside URLs and email addresses, such as digits, emoji, punctuation or a bare link, returns zxx (no linguistic content) without running the model. Every other text gets the model's best guess, even lol or a person's name. To get und instead when the model is unsure, set thresholds, and use detailed=True to see why:
detector = LanguageDetector(min_letters=4, min_score=0.5)
detector.detect("lol", detailed=True)
# Detection(language='und', status=<DetectionStatus.INSUFFICIENT_TEXT: 'insufficient_text'>, ...)
LanguageDetector(candidates=["eng", "spa", "cat"]) # only ever answer one of these
Scores aren't calibrated and differ between backends, so a threshold that suits one backend doesn't suit another. Benchmarks lists tested settings and their wrong-label and abstention rates per backend.
Translate
from noentenc import Language
from noentenc.translation import Translator
translator = Translator()
translator.translate_batch(
["Where is the station?", "I love this city."], Language.SPANISH, Language.ENGLISH
)
# ['¿Dónde está la estación?', 'Me encanta esta ciudad.']
# A mixed-language inbox: detect each text's language and translate in one call.
translator.translate_batch(
["Hola, ¿cuándo llega mi pedido?", "Der Link funktioniert nicht.", "Thanks!"],
"en",
"auto",
)
# ['Hey, when does my order arrive?', "The link doesn't work.", 'Thanks!']
translator.translate_dataset(df, "review", "review_en", "en", "auto")
Translator() picks the lightest model for each pair. It uses a dedicated Opus-MT model when one exists for the direction (66 directions, about 75M parameters each). For any other pair it uses SMaLL-100, which covers 100 languages. A model that can't handle a pair raises UnsupportedLanguageError. Languages are Language members or codes and names: "es", "spa", "es-ES" and "spanish" all mean Spanish. noentenc.to_language() does that conversion on its own, which also turns detector labels like "spa" or "cmn" into Language members.
source_language="auto" detects each text's language, groups the texts by language and translates each group with the model for that pair, then returns them in input order. Texts already in the target language and texts without linguistic content come back unchanged. When the detector isn't sure (lol, a name), the text is kept as given with status unknown_source. unknown_source="fallback" translates such texts without a source language instead, and unknown_source="raise" fails the call with SourceLanguageError. detailed=True reports each text's detected language, score, status and model. Without a source language, Translator can't pick an Opus-MT model, so it uses SMaLL-100 for every text.
Text of any length works. It is split into sentences, they are translated in one batch, and the translations are joined back with the original spaces, line breaks and blank lines. A single sentence longer than the model can read (about 500 tokens) raises InputTooLongError instead of being cut silently. Pass truncate=True to translate only its start, and detailed=True to get a Translation that says whether input was dropped (input_truncated) or the output hit its length limit (output_limit_reached).
Translator keeps the models it loads for later calls, at most two at a time by default. A loaded model needs a lot more RAM than its download: about 1.1 GB for an Opus-MT pair at q4 and 1.15 GB for SMaLL-100. Pass max_loaded_models= to change the limit (None for no limit), and call translator.unload() to free them all. See memory.
Links, emails and placeholders survive. URLs, email addresses, mentions, hashtags, numbers, inline code, template placeholders ({name}, {{x}}, ${x}, %s), HTML tags and Markdown link targets come back byte-for-byte while the prose around them is translated:
translator.translate(
"Visit https://example.com/reset?id=42 or email support@acme.io", "es", "en"
)
# 'Visite https://example.com/reset?id=42 o envíe un correo electrónico a support@acme.io'
# Without it (preserve=False), Opus-MT writes https://ejemplo.com/reset?id=42: another site.
See Keep links and placeholders for what is covered and its limits.
For bulk jobs, errors="record" keeps going past a text that fails. That text comes back as given, with its status and error in the detailed result (or in an error_column for translate_dataset), and the rest are still translated. Empty batches, blank texts and same-language requests return without loading a model.
Trade speed for quality
LanguageDetector and Translator take a profile in place of a model: "speed" (the default), "balance" or "quality".
from noentenc import Profile
from noentenc.language_detection import LanguageDetector
from noentenc.translation import Translator
LanguageDetector("quality") # fastText GlotLID
Translator(Profile.BALANCE) # Opus-MT where it exists, otherwise NLLB-200 at int8
| Profile | Detection | Translation, pairs without an Opus-MT model |
|---|---|---|
speed (default) |
lid176: 0.9 MB |
SMaLL-100: 595 MB |
balance |
openlid-v3: 1.2 GB, GPL-3.0 |
NLLB-200 int8: 860 MB, CC-BY-NC-4.0 |
quality |
glotlid: 1.7 GB |
NLLB-200 fp32: 3.5 GB, CC-BY-NC-4.0 |
On FLORES-200, balance raises detection accuracy from 50% to 96% over the 176 languages tested, and NLLB-200 adds up to 19 chrF++ on low-resource pairs such as English to Tamil. See the profiles benchmark for the numbers behind each choice. The models behind a profile may change between releases, so pass a model explicitly when you need reproducible output.
All three translation profiles use the same Opus-MT model for the 66 pairs it covers, and SMaLL-100 for texts without a source language. They only differ on the other pairs; see routing.
Check before you download
Ask what a profile would load, what it costs and whether it's allowed, without downloading anything:
import noentenc
from noentenc.translation import Translator
plan = noentenc.plan("balance", translation=[("ja", "ca"), ("en", "es")])
for model in plan.models:
print(model.name, model.license, model.download_bytes >> 20, "MB", model.cached)
# On a machine that hasn't downloaded anything yet:
# FastTextModel(openlid-v3) GPL-3.0 1175 MB False
# NLLBModel(Xenova/nllb-200-distilled-600M) CC-BY-NC-4.0 869 MB False
# OpusMTModel(Xenova/opus-mt-en-es) Apache-2.0 290 MB False
# Only pick permissively licensed models, and nothing over 700 MB.
translator = Translator(
"balance",
allowed_licenses=["MIT", "Apache-2.0", "CC-BY-4.0"],
max_download_bytes=700 << 20,
)
translator.plan("ca", "ja").models[0].name # 'SMaLL100Model(casawolice/small100-onnx)'
translator.supports("ace", "en") # False: only NLLB-200 writes Acehnese
Each plan reports the model's licence, download size, RAM (measured, or estimated where marked) and cache state. A pair with no allowed model raises ModelConstraintError before anything downloads. prepare() takes the same constraints.
Examples
| Example | Shows how to |
|---|---|
| detect_dataset.py | Tag a DataFrame column and keep only confident English rows. |
| conservative_detection.py | Abstain on names, chat tokens and links instead of guessing, and restrict the candidate languages. |
| translate_to_english.py | Translate a mixed-language inbox and DataFrame into English in one call with source_language="auto". |
| choose_detection_backend.py | Swap backends, restrict candidate languages, collapse macrolanguages. |
| choose_translation_model.py | Pick a model, precision and thread count, and run offline. |
| prepare_offline.py | Download a profile's weights into one directory, then detect and translate without network access. |
| translate_long_text.py | Translate emails and documents, and choose between an error and truncate=True for over-long sentences. |
| manage_memory.py | Bound how many translation models stay loaded, and free them with unload(). |
| plan_and_limit.py | See what a profile would download and what it costs, and restrict it by licence and download size. |
| preserve_literals.py | Translate messages without breaking their links, emails, code, tags and placeholders. |
| translate_bulk.py | Run bulk jobs with errors="record" so one bad row doesn't stop them, and skip model loads for empty, blank and same-language input. |
| choose_profile.py | Trade latency for quality with the speed, balance and quality profiles. |
| custom_models.py | Plug your own detector and translator into the same API. |
Models
Language detection
| Backend | model= |
Install | Languages | Size | Licence (weights) |
|---|---|---|---|---|---|
FastTextModel |
lid176 (default) |
core | 176 | 0.9 MB | CC-BY-SA-3.0 |
lid176-bin |
core | 176 | 126 MB | CC-BY-SA-3.0 | |
openlid-v2 |
core | 200 | 1.2 GB | GPL-3.0 | |
openlid-v3 |
core | 195 | 1.2 GB | GPL-3.0 | |
glotlid |
core | 2102 | 1.7 GB | Apache-2.0 | |
nllb-lid218e |
core | 218 | 1.2 GB | CC-BY-NC-4.0 | |
path to a .bin/.ftz |
core | ||||
OnnxClassifierModel |
bert-openlid (int8) |
noentenc[onnx] |
201 | 25 MB | MIT |
xlm-roberta-lid (int8) |
noentenc[onnx] |
20 | 279 MB | MIT | |
LangidModel |
core | 97 | 1.9 MB | BSD-2-Clause | |
LinguaModel |
noentenc[lingua] |
75 | ~300 MB wheel | Apache-2.0 | |
Cld3Model |
noentenc[cld3] |
107 | 1 MB | Apache-2.0 | |
HeliportModel |
noentenc[heliport] (no Windows) |
220 | ~130 MB wheel | GPL-3.0 |
from noentenc.language_detection import FastTextModel, LanguageDetector, LinguaModel
LanguageDetector(FastTextModel("glotlid")) # 2102 languages
LanguageDetector(LinguaModel(languages=["cat", "spa", "eng"])) # only these candidates
- Labels are ISO 639-3 codes (
eng,cat,zho). Empty and whitespace-only texts returnund, and texts with no letters outside URLs and email addresses returnzxx, for every backend and without running it.normalize_labels=Falsereturns each model's native codes instead.collapse_macrolanguages=Truefolds individual languages into their macrolanguage (arb→ara,cmn→zho), so results from different models line up.
FastTextModelis our own numpy implementation of fastText inference. It needs neither thefasttextpackage nor onnxruntime, and it matchesfasttext's output to within 1e-6. Large.binmodels are memory-mapped, so they open instantly.LangidModelis our own numpy implementation of langid.py. It reads the weights from the langid 1.1.6 source release on PyPI, without installing or importing thelangidpackage, and matches its probabilities to within 1e-9.
Translation
Every translation model needs noentenc[translation].
| Model | Languages | Download (default precision) | Licence (weights) |
|---|---|---|---|
OpusMTModel.from_pair(src, tgt) |
1 direction each, 66 available (OPUS_MT_PAIRS) |
287 MB (q4), 107 MB (int8) | CC-BY-4.0 or Apache-2.0, per pair |
SMaLL100Model |
any source → 100 targets | 595 MB (int8) | MIT |
M2M100Model |
100 ↔ 100 | 603 MB (int8) | MIT |
NLLBModel |
196 ↔ 196 | 860 MB (int8) | CC-BY-NC-4.0, used by the balance and quality profiles |
Every translation model takes precision= (fp32, int8, q4, where the export has it), num_threads= and a Hugging Face repo id or local directory as model=.
Weights and licences
Weights are pinned to a revision and, where we download them ourselves, checked against a sha256. cache_dir= or NOENTENC_CACHE sets one directory for detection and translation weights. noentenc.prepare(profile, translation=[(source, target), ...]) downloads a profile's models without loading them. With only_local_files=True on LanguageDetector, Translator or any model, a model that isn't cached raises at once instead of downloading, which is what you want on an air-gapped server.
The weights' licences apply to your use of the weights. They don't affect this package's licence.
Bring your own model
Load your own fastText, ONNX classifier or seq2seq weights into an existing backend, or wrap any detector or translator by subclassing a base class with a few methods. See docs/custom-models.md.
from noentenc.language_detection import FastTextModel, LanguageDetector
LanguageDetector(FastTextModel("models/my-domain-lid.ftz"))
Development
make install # uv sync with every extra + dev tools
make lint
make unit-tests
make integration-tests # tests/integrations; downloads real weights
scripts/benchmark_lid.py and scripts/benchmark_translation.py reproduce the speed numbers, and scripts/benchmark_vs_reference.py the comparison with the original implementations. scripts/make_fasttext_fixtures.py regenerates the fastText parity fixtures with the reference fasttext package. That package needs Python 3.12, and the script's docstring has the command. scripts/make_langid_fixtures.py does the same for langid with the reference langid package.
Metadata
Release files for noentenc 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| noentenc-0.3.0.tar.gz | 91.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| noentenc-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 200.8 kB
Release files / noentenc-0.3.0.tar.gz
| Download URL | noentenc-0.3.0.tar.gz |
|---|---|
| Size | 91.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b730db07d174b2c9b2342eca868a1b6b24b3105efaaf9f978e6d9fb7ca6e70c2
|
|
BLAKE2b-256 checksum How to use checksums |
aabe6717c7f3380ff99e0744bd3593781fd4b5f92e11758c9f2777b4b75a449d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / noentenc-0.3.0-py3-none-any.whl
| Download URL | noentenc-0.3.0-py3-none-any.whl |
|---|---|
| Size | 109.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9603f999db7fb0b7ded5807545d934b178abed5cd04e3eb2a6647d56152c292f
|
|
BLAKE2b-256 checksum How to use checksums |
b909826a3c24a8b0b453edd52e366389221f5793a81f26a982bfb8b68d85611f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|