Pantogloss
Pantogloss is a TensorFlow/Keras library that translates text from many languages into English. It provides a Python API, command-line interface, and optional persistent local server with a browser UI.
The default public model is
pantogloss-500-en-compact,
a 539 MiB model validated on NVIDIA CUDA and Apple Metal. Model weights are
downloaded separately from Hugging Face and are never bundled in the Python
wheel. Public models normally require no Hugging Face login.
What it can do
- Translate individual strings, batches, files, and bounded input streams.
- Run on CPU, NVIDIA CUDA, or Apple Silicon Metal.
- Keep a model resident behind a local HTTP API for fast repeated requests.
- Provide a simple Any language → English browser translator.
- Offer fast greedy and higher-quality beam-search decoding presets.
- Return plain strings or structured results with timing and runtime metadata.
- Attach caller-supplied language quality guidance and conservative warnings.
- Preserve input order and isolate individual failures in long-running streams.
Pantogloss translates text; it does not detect the source language or parse document formats. See Project scope.
Installation
Pantogloss supports Python 3.10–3.12. Use the extra for your accelerator:
# CPU
python -m pip install pantogloss
# NVIDIA GPU on Linux
python -m pip install "pantogloss[cuda]"
# Apple Silicon Metal
python -m pip install "pantogloss[metal]"
# Persistent API and browser UI; combine with cuda or metal when needed
python -m pip install "pantogloss[server,metal]"
The first use downloads the selected model. Later loads use the local Hugging Face cache.
Python API
Load one translator and reuse it:
from pantogloss import Translator
translator = Translator.from_pretrained(device="auto")
print(translator.translate("Comment allez-vous ?"))
Translate a batch or request structured results:
translations = translator.translate([
"Hola señor",
"Wie geht es Ihnen?",
])
result = translator.translate_detailed("Bonjour le monde.")
print(result.text)
print(result.elapsed_seconds, result.execution_device)
For a large iterable, translate_iter() and translate_iter_detailed() process
bounded batches without retaining the complete input or output collection.
with open("source.txt", encoding="utf-8") as source:
for translation in translator.translate_iter(source, batch_size=16):
print(translation)
Use the quality preset when latency is less important than beam-search quality:
print(translator.translate("Comment allez-vous ?", preset="quality"))
Command line
pantogloss translate "Comment allez-vous ?"
pantogloss translate --preset quality "Comment allez-vous ?"
pantogloss translate --input source.txt --output english.txt --batch-size 16
pantogloss models
pantogloss info --device gpu --json
pantogloss doctor
Translation output stays clean on stdout. Use --verbose, --report, or
--tensorflow-logs when diagnostics are wanted. Use --offline to require an
already cached model.
Persistent server and browser UI
Each standalone CLI invocation reloads the model. For repeated requests, keep one warmed translator in memory:
python -m pip install "pantogloss[server,metal]" # or server,cuda
pantogloss serve --device gpu
Open http://127.0.0.1:8765/ for the two-pane browser translator. The same
process exposes:
GET /healthGET /infoGET /metricsPOST /translate
curl http://127.0.0.1:8765/translate \
-H 'Content-Type: application/json' \
-d '{"text":"Comment allez-vous ?"}'
The server binds to loopback, warms the model before reporting readiness, and
serializes inference by default. Queuing is bounded, shutdown drains active
TensorFlow work, and optional request logs contain timing/count metadata rather
than source or translated text. On Apple Silicon, the server defaults to the
bounded-memory eager decoder because TensorFlow Metal retains memory during
compiled decoding; --metal-compiled-decode is an explicit faster but
memory-growing opt-in. See the
Server and Web UI wiki page
for deployment, authentication, limits, API schemas, benchmarks, and the
pantogloss balance front end for independently resident workers.
The balancer supports explicit graceful draining, connection-only safe retries,
passive circuit breaking, and Prometheus-compatible reliability metrics.
Models
| Name | Role | Approximate artifact size |
|---|---|---|
pantogloss-500-en-compact |
Recommended default | 539 MiB |
pantogloss-500-en-compact-v6 |
Previous V6-derived compact default | 538 MiB |
pantogloss-500-en-compact-v1 |
Previous compact model | 538 MiB |
pantogloss-500-en-fp16 |
Accelerator-oriented FP16 reference | 1.02 GiB |
pantogloss-500-en |
Historical FP32 numerical reference | 2.03 GiB |
pantogloss-500-en-int8 |
Smaller experimental alternative | 539 MiB |
pantogloss-500-en-v7 |
Opt-in FP32 successor with conversational Spanish improvements | 2.03 GiB |
pantogloss-500-en-v6 |
Opt-in FP32 fine-tuned successor | 2.03 GiB |
Select a model explicitly when needed:
translator = Translator.from_pretrained("pantogloss-500-en-fp16", device="gpu")
successor = Translator.from_pretrained("pantogloss-500-en-v7", device="gpu")
V7 builds on V6 with a bounded conversational Spanish-to-English fine-tune.
It improved chrF by 2.04–2.78 on two sealed Fisher/CALLHOME confirmation splits
while improving aggregate chrF by 0.217 on the 50-language development suite;
no evaluated language crossed the frozen −0.5 chrF retention floor. V7 remains
available as an opt-in full-size model and now supplies the recommended compact
default. Compact V7 retains the measured conversational Spanish gains while
remaining approximately 539 MiB. It slightly improves aggregate chrF over the
previous V6-derived compact default, though individual languages can vary. The
full-size FP32 V6 model remains available as an opt-in reference.
Caller-supplied language-quality guidance is tied to the evaluated model
revision: the default and original FP32 reference have separate measured
catalogs; other model variants return unmeasured rather than borrowing scores.
The INT8 alternative is not the default because it is substantially slower.
Detailed evidence and backend caveats live in the model cards, wiki, and
checked-in experiment reports.
Documentation
- Installation and Quickstart
- Server and Web UI
- Decoding Presets
- Platform Validation
- Language Quality Catalog
- Quality Evaluation
- Development and Testing
- Tika Document Translation
The repository also retains reproducible model-conversion, parity, compression,
and evaluation artifacts under docs and evaluation.
Project scope
Pantogloss translates text and ordered text segments. It intentionally does not
identify languages, detect file types, or parse documents. DocumentTranslator
is a neutral text-segmentation helper. pantogloss-tika remains a compatibility
example and is not a direction for new core dependencies. See the
architecture boundary.
Translation quality varies by language, domain, and input. Pantogloss does not provide calibrated confidence and should not be relied on without review for medical, legal, safety-critical, or other high-stakes decisions.
Provenance and license
The original model was described by Thamme Gowda, Zhao Zhang, Chris A. Mattmann, and Jonathan May in Many-to-English Machine Translation Tools, Data, and Pretrained Models, ACL-IJCNLP 2021 System Demonstrations, DOI 10.18653/v1/2021.acl-demo.37.
Pantogloss and its converted models are licensed under the Apache License, Version 2.0. See NOTICE for attribution.
Release files for pantogloss 0.25.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pantogloss-0.25.0.tar.gz | 94.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pantogloss-0.25.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 204.8 kB
Release files / pantogloss-0.25.0.tar.gz
| Download URL | pantogloss-0.25.0.tar.gz |
|---|---|
| Size | 94.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d0c6770bd789df5e864ff046dce21bcd441f562e03d10dc325e8742dd5be88fb
|
|
BLAKE2b-256 checksum How to use checksums |
354b8c71d60e15b79ee173cfeeb30121c9ed2fd40149678b86f19b12c409e220
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency logRelease files / pantogloss-0.25.0-py3-none-any.whl
| Download URL | pantogloss-0.25.0-py3-none-any.whl |
|---|---|
| Size | 109.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b4dacab3ad2b497e098e7b75d0dee64f807ace7a30d7eba63095e0bce9b38198
|
|
BLAKE2b-256 checksum How to use checksums |
8943d6d4c21289663d46ea03ee11b5d63a99b01d0933f6fd48a5d200a62aef0d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency log