Skip to main content

Pantogloss

Pantogloss is a TensorFlow/Keras many-to-English machine-translation library. Its first model, pantogloss-500-en, was converted and numerically validated from the model described in Many-to-English Machine Translation Tools, Data, and Pretrained Models (ACL-IJCNLP 2021).

The Python package is distributed through PyPI, while the initial model is kept in a separate private Hugging Face repository. Installing Pantogloss does not grant model access; users must be authorized for chrismattmann/pantogloss-500-en and authenticate with hf auth login.

The codebase and converted model are licensed under Apache-2.0. This repository is private during initial development.

Intended API

from pantogloss import Translator

translator = Translator.from_pretrained("pantogloss-500-en")
print(translator.translate("Comment allez-vous ?"))

RTG-compatible beam search is available without changing the return type:

print(
    translator.translate(
        "Comment allez-vous ?",
        beam_size=4,
        length_penalty=0.6,
    )
)

Pantogloss selects the first TensorFlow GPU automatically and enables memory growth. Device choice can also be made explicit:

translator = Translator.from_pretrained("pantogloss-500-en", device="gpu")
print(translator.device_info)

Using device="gpu" fails clearly if TensorFlow cannot see a GPU; use device="cpu" to force CPU inference.

Install the accelerator backend for the machine:

# Linux with an NVIDIA GPU
pip install 'pantogloss[cuda]'

# Apple Silicon
pip install 'pantogloss[metal]'

Both use the same device="auto" or device="gpu" Python API. The CUDA extra does not install or replace the host NVIDIA driver. The Metal extra uses Apple's TensorFlow PluggableDevice and the TensorFlow 2.18 runtime combination validated by the Bytewise project.

The model is stored separately in the private Hugging Face repository chrismattmann/pantogloss-500-en; it is never included in the Python wheel.

Command line

The pantogloss command loads the model once and supports arguments, files, and line-oriented Unix pipelines:

pantogloss info
pantogloss translate "Comment allez-vous ?"
printf 'Hola señor\nWie geht es Ihnen?\n' | pantogloss translate --device gpu
pantogloss translate --input source.txt --output english.txt --batch-size 16
pantogloss translate --beam-size 4 --length-penalty 0.6 "Hola señor"

Use --json for JSON Lines output and --offline to require an already cached model snapshot. Translation data goes to stdout (or --output); model and device diagnostics are suppressed by default so pipelines remain clean. Use --verbose for Pantogloss loading progress or --tensorflow-logs for TensorFlow, CUDA, and Metal startup diagnostics.

Development status

The complete 307-variable Keras model has been converted locally from all 308 learned PyTorch tensors (the target embedding and output projection are tied). Greedy parity against the archived RTG implementation passes across a ten-language batch: token IDs and translations match exactly, while final logits have a maximum absolute error of 1.24e-5. With the original beam size 4 and length penalty 0.6, all decoded four-best candidate sets match. One near-tied example changes top rank because of framework floating-point ordering. Model version 0.1.0 is released in the private Hugging Face repository at an immutable commit.

The source model and generated artifacts stay under the ignored artifacts/ directory. To reproduce conversion after acquiring the source archive:

python tools/convert_rtg_checkpoint.py \
  artifacts/source/rtg500eng-tfm9L6L768d-bsz720k-stp200k-ens05 \
  artifacts/converted/pantogloss-500-en-candidate

Run the reference parity harness with:

CUDA_VISIBLE_DEVICES=-1 python tools/check_parity.py \
  artifacts/source/rtg500eng-tfm9L6L768d-bsz720k-stp200k-ens05 \
  artifacts/converted/pantogloss-500-en-candidate

To require and verify real GPU placement:

python tools/check_gpu.py artifacts/converted/pantogloss-500-en-candidate

Apple Silicon validation

Pantogloss uses the same hardware-neutral GPU API for CUDA and Metal. On an M-series Mac with Python 3.12 and Xcode command-line tools installed:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[metal,test]'
hf auth login
python tools/check_platform.py --device cpu
python tools/check_platform.py --device gpu
python tools/benchmark_inference.py --device gpu --runs 5

The portable platform report identifies the selected backend as cpu, cuda, or metal, verifies the first model variable's actual TensorFlow placement, and runs a real translation. Metal placement and inference are validated on an Apple M3 Max with TensorFlow 2.18.1. For a short batch-one sentence, its warmed median was 0.545 seconds on CPU and 0.633 seconds on Metal; accelerator benefits are expected primarily from batching and future compiled decoding.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pantogloss-0.1.2.tar.gz (31.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pantogloss-0.1.2-py3-none-any.whl (25.5 kB view details)

Uploaded Python 3

File details

Details for the file pantogloss-0.1.2.tar.gz.

File metadata

  • Download URL: pantogloss-0.1.2.tar.gz
  • Upload date:
  • Size: 31.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pantogloss-0.1.2.tar.gz
Algorithm Hash digest
SHA256 3523d35ac06d6a8e95a72b5b42b7f04f5865ad1ebd6dba0abeda7dedecf7e122
MD5 6b7aba30be652301131d88f88b1cfd1c
BLAKE2b-256 28b96a7f6ff8d55c35b47ffdc8f3274e78821d1335f155ecfe1eb70b9d54c292

See more details on using hashes here.

Provenance

The following attestation bundles were made for pantogloss-0.1.2.tar.gz:

Publisher: release.yml on chrismattmann/pantogloss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pantogloss-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: pantogloss-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 25.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pantogloss-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 2a9120a1d61754b3dde2c9fe8d558487e1ca73d135e807a1c4c786f9dda70277
MD5 64a723d83c18532b0dbcd6ee62aff593
BLAKE2b-256 cc48f64fe517a7afa15735b598c5d1d485d9f764ba1785d6bf5ac9fe56c954a9

See more details on using hashes here.

Provenance

The following attestation bundles were made for pantogloss-0.1.2-py3-none-any.whl:

Publisher: release.yml on chrismattmann/pantogloss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

This release

0.1.2 This release

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page