Skip to main content

nmt-forge

Train NMT models without fooling yourself. nmt-forge is a general-purpose training suite that makes the classic training/eval mistakes — leaked test sets, test-driven checkpoint selection, scores without error bars, structural gaps hidden by volume — structurally hard to commit. It doesn't warn; it refuses, and every refusal says what happened, why it corrupts results, and the exact fix.

Every guard mechanizes a real, measured failure from Champollion's Plains Cree work (the 2026-07-12 mistake ledger). Scoring is delegated entirely to mt-eval-harness — forge implements zero metrics.

Install

python3 -m pip install nmt-forge              # the guards, splits, audits, scoring (via mt-eval-harness)
python3 -m pip install 'nmt-forge[hf]'        # + training, export and serving (torch, transformers,
                                   #   accelerate, sentencepiece, peft — CPU wheels are fine)

mt-eval-harness comes in as a dependency: it is forge's scorer AND its language-card resolver, so discover/init work from a plain pip install — cards come from --cards-dir, $MT_EVAL_CARDS_DIR, a checkout or node_modules/champollion above you, or the public card index (cached, reused offline). Offline with no cache: npx champollion network card crk --json > cards/crk.json, then --cards-dir cards.

From pip install to a served model — on a laptop

The north star: "a Cree model for our school" — ~1,600 parallel pairs and a private, teacher-checked test set — or "an Atya model for the hospital" — ~900 pairs and a sensitive test set. No GPU required:

nmt-forge init crk --dir school-mt          # card → workspace + config + NEXT_STEPS.md
cd school-mt                                # run everything from the project dir
nmt-forge registry add project-test ~/teacher-test.jsonl --role test   # YOUR test set
nmt-forge leak-audit ~/pairs.jsonl --clean-to pairs.clean.jsonl        # screen against it
#   …says most test rows have a near-twin in your pairs? a second, twin-free model:
#   --clean-to pairs.notwins.jsonl --drop-test-twins (its OWN file — it writes
#   config-notwins.json too) — one prereg per model, named after it
nmt-forge prereg template --out predictions.json      # write predictions BEFORE any score —
nmt-forge prereg new all-data --eval-set project-test --predictions predictions.json
#   …named after the model it predicts (`prereg new notwins …` for the twin-free one),
#   and before any benchmark (mt-eval run) on the test set: a read blocks a later prereg
#   …then read the training guardrails, once, before the split: the "Train a Model
#   Honestly" page (agents: the MCP tool get_training_guardrails)
nmt-forge split pairs.clean.jsonl --test 0 --dev 100 --seed 42 \
    --out data/split --register project     # train/dev only — the test set stays yours
nmt-forge preflight run --config config.json          # the checks run makes, incl. the [hf] extra
nmt-forge run config.json                   # cpu-tiny: minutes on a CPU, no download
nmt-forge export .forge/runs/<run>/run-manifest.json --prereg all-data --out export/
#   …the twin-free model: --prereg notwins --out export-<its run>/ (with two preregs
#   on one test set, export refuses to guess which one predicted which model)
nmt-forge serve export/model                # http://127.0.0.1:8378 — champollion can call it

That order — register the test set, screen the corpus, one preregistration per model, (benchmarks), read the guardrails, split, train — is the one nmt-forge init, NEXT_STEPS.md and nmt-forge status give. --prereg <id> on export names the prediction that judges each model. Pinning is the other way: prereg new <id> … --config-hash <hash> binds a prediction to one config — the full hash nmt-forge preflight run --config <its config> prints — but any later edit of that config (a time budget, say) changes the hash and drops the pin, so naming it on export is the simpler route.

export scores the test set once (prereg-gated, 95% CIs), writes an mt-eval RunLog + TestReport (built by the harness itself with the metric battery mt-eval run loads for your language, so mt-eval compare works on it — anything it could not compute is listed with the command that computes it) into export/evaluation/, and packages the model into export/model/: weights, tokenizer, a champollion plugin manifest and a DEPLOY.md (with the per-pair fallback config for strings the model cannot do). Every caveat the harness writes on the score — a near-constant output (one of a few sentences for many different inputs), length, copies of the source — is passed on in its own words beside the score: in the export summary, forge-model.json, DEPLOY.md, status, report, compare and lint. Scores follow the harness's scoring standard (standard/1): the headline — the number to quote — is corpus chrF++ with its 95% bootstrap CI, written chrF++ 47.5 [45.9, 49.0], with its sacreBLEU signature wherever a full record is kept (forge-model.json, DEPLOY.md, the export summary). BLEU, spBLEU, TER and COMET are shown beside it, never blended into it; exact match and the referee lanes are labelled diagnostics. forge prints no composite and no quality label: only a speaker's review certifies quality. nmt-forge lint takes the battery manifest (export/evaluation/battery-hyps-battery.json); given a run's run-manifest.json it lints every scored export of that run, or tells you the nmt-forge export command to run first. export/model/ is what you deploy — it holds no test sentence; export/evaluation/ holds your test set's text and never leaves your machine (marked with the test set's terms when it has any). serve speaks the champollion api-method contract (POST /translate — "method": "api") and an OpenAI-compatible /v1/chat/completions (champollion sync --method local) — app strings and Markdown content files alike. The model never sees placeholders, ICU plural/select syntax, tags or Markdown markup: the server copies them around the model's output and hands it only the text between them, sentence by sentence. export prints the verdict on each preregistered prediction and how many test sentences have a near-twin in training (when most do, it says the score measures recall, not translation) — and split, leak-audit and preflight run say the same count BEFORE training, while the test set is still unspent, with the fix: split --near-dupe 0.6 holds out whole templates when forge carves the test set; for a FIXED test set like the teacher's, leak-audit <pairs> --clean-to <pairs.notwins.jsonl> --drop-test-twins drops the training rows that are near-twins of it (same measure and threshold as the count), says how many go and what the strict subset becomes, and refuses to leave nothing to train on. It writes to its own file: leak-audit refuses to replace a file a config, a run or a split reads, or another audit's output (--overwrite only on purpose). At every step nmt-forge status names the next command; while a run is training it says training — wait — and nmt-forge run refuses to start a second run in the same workspace. With several exported models, the user picks the one to deploy: nmt-forge choose <export>/model records it (serving a model to try it is recorded as served, never as the choice).

Enter it in a sovereign contest. A contest's declarative lane (Lane A, mt-eval contest submit-model) takes a model as data: weights, config and tokenizer. export/model/DEPLOY.md §6 names exactly those files (never forge-model.json — this model's scores on your test set and local paths — nor DEPLOY.md or champollion-plugin/; submit-model leaves all three out by itself), the architecture, and the parameter count the lane checks, read from the weights file's header, with the command. The contest side: Run a Sovereign Contest.

A private test set stays private. Mark it the way a data steward does — champollion network register-corpus --tier local-only --role test, or a sidecar next to the file: echo '{"transmission": "local-only"}' > teacher-test.tsv.champollion.json — and forge, like mt-eval, never prints its sentences. leak-audit names the rows that matched it by line number (a row that matched a test answer IS most of that answer), says once why, and shows the text only with --show-text — for a person at the terminal; an AI agent passes whatever it reads to its model provider. --json never carries sentences. The same goes for sealed and consent-required corpora (the harness decides which), for a training corpus that is itself marked, and for an error message that quotes such a row ([sentence withheld]). Files forge writes into your folders keep the text, and files it carves from a marked corpus (split sides, leak-audit --clean-to, sample) get the same mark. When the sidecar names a registered corpora card whose sha256 matches the file, the card's id is the set's dataset id in forge's registry listing, reports, export and mt-eval RunLog — the name mt-eval runs on the file use — with your registry add name (project-test) kept beside it.

Model presets (nmt-forge init --model …, written out as explicit numbers in config.json):

preset what it is needs honest expectation
cpu-tiny (default) a ~6M-parameter Marian trained from scratch; BPE vocabulary fit on the TRAIN rows only a CPU, no download — 1,600 pairs train in ~2–3 minutes on a laptop weak: on 1–2k pairs, chrF++ roughly 5–30 (the top end only for highly templated data) — your data's phrases and templates, not general translation
cpu-finetune --base <hf-id> fine-tune a small pretrained Marian/opus-mt (you name a RELATED pair) a CPU, ~300 MB download usually better than cpu-tiny when a related pair exists — measure it, don't assume
nllb-600m NLLB-200 distilled 600M + LoRA a GPU, ~2.5 GB download the strongest start; the wall-clock gate refuses it on a CPU in minutes

The point of cpu-tiny is not the score: it makes the WHOLE loop real — the fence, the audits, the preregistered test, an exportable model the CLI can call — so a better model later drops into the same project and is measured the same way.

Sixty seconds, end to end

# carve an honest split: pairs sharing a source OR target stay together
$ nmt-forge split corpus.jsonl --test 150 --dev 42 --seed 42 \
      --out data/split --register textbook
split corpus.jsonl: 1240 rows in 1187 share-groups (largest 4)
  train 1048 · dev 42 · test 150  → data/split/
  verified: 0 shared canonical source/target keys across sides

# train: one command, config-hashed, dev-fenced
$ nmt-forge run config.json
dev report (95% CIs — there is no bare-score rendering):
n=42 · set=textbook-dev
  chrf++       44.31  [41.20, 47.15] 95% CI

# score the test set? not without predictions written down first
$ nmt-forge score --eval-set textbook-test --hyps decoded.txt
[preregister] no preregistration for eval set 'textbook-test' ...
  why: results looked at without written-down expectations become post-hoc stories
  fix: write one FIRST: ... — then score

$ nmt-forge prereg template --out preds.json     # the ONE format; edit it
$ nmt-forge prereg new e2 --eval-set textbook-test --predictions preds.json
$ nmt-forge score --eval-set textbook-test --hyps decoded.txt
  chrf++       46.02  [43.11, 48.87] 95% CI

That's the whole philosophy: the honest path is the easy path, and the dishonest paths are closed with actionable messages.

Agents: every command takes --json — exactly one JSON document on stdout, refusals as {"error": {…, "why", "fix"}} with exit 2. Shapes: docs/JSON_OUTPUT.md. Call nmt-forge status --json first, always.

Any language: start from the card

forge is general-purpose across all ~7,900 SSOT language cards. discover answers "what does this language actually have?" honestly (absence on a card = unknown, never zero), and init scaffolds a project from it:

$ nmt-forge discover nav
Navajo (nav) · ltr
WHAT THE CARD SAYS EXISTS (absence = unknown, not zero):
  OPUS: 5 corpora, 36533 aligned pairs
  unknown (card is silent): analyzers, dictionaries, eval datasets
THE ASSET LADDER — what this language can do TODAY:
  ✓ rung 1: parallel text → train with every guard (no pack needed)
  ? rung 2: monolingual text → the tagged backtranslation lane
  ? rung 3: dictionary (+ grammar) → a cited template pack is worth building
  ? rung 4: morphological analyzer → round-trip-VERIFIED synthesis
  ? rung 5: LYSS referee → the language's own metric in selection
  note: no analyzer on the card → synthesis is off the menu until one
  exists; every guard and the training loop work regardless

$ nmt-forge init nav --dir my-navajo-mt
# → .forge/ workspace + starter config.json + NEXT_STEPS.md (the agent brief)

# a language the index doesn't have yet? still trainable — nothing invented:
$ nmt-forge init qzx --no-card --name "My Language" --dir my-mt

A rich card wires more for free: Plains Cree's card carries its eval-dataset ids (cross-checked against the mt-eval registry — all flagged NEVER TRAIN ON THIS) and its LYSS referee, so discover crk emits ready-to-paste --plugin champollion_lyss... lanes. The referee is an OPTIONAL add-on with its own license, never a requirement: init crk wires its lanes into the starter config only when the package is installed, and otherwise names it in NEXT_STEPS.md — the config passes its own preflight on a plain install. Same tool, no special cases — the card decides.

What's inside (start here, dig later)

you want to… use it kills the mistake of…
split a corpus split / verify-split test answers hiding in training via shared sources/targets
pick checkpoints the run's dev-fence the test set choosing the model — and training on the dev set's own rows (a file written before the split, such as an early twin-free corpus, is refused with the re-audit that fixes it)
screen any corpus/harvest leak-audit training on eval text (exact, reworded, or whole-file) — while KEEPING template siblings ("I see the dog" vs "I see the cat") and similar-prompt/different-answer pairs (an identical prompt is dropped whatever its answer), and saying which is which, with examples (by line number only for a private — local-only, sealed, consent-required — set or corpus); --drop-test-twins also drops the near-twins of a FIXED test set when they would turn its score into recall
generate training data synth + a language pack unverified forms, uncited templates, invisible gaps
sample synthetic data sample two template kinds hogging half the signal
report numbers score / compare / evaluate scores with no error bars, no preregistration
hand the model on export / serve a result nobody else can compare (mt-eval TestReport) or deploy (champollion api / OpenAI-compatible endpoint)
see how spent an eval set is ledger show invisible adaptive use; sealed sets are one-shot

Each guard is also a library call under nmt_forge.guards.*. The full mistake→mechanism map, with the measured numbers behind each guard, is in DESIGN.md.

Language packs plug in — forge ships none

forge is general-purpose; language-specific code lives in the language's own home and plugs in through the pack interface (analyzer + dictionary adapter + orthography + grammar-cited templates + checklist):

# from any checkout (no install needed):
nmt-forge synth nmt_forge_crk.pack:get_pack --out data/synth.jsonl
# or, once the pack's package is installed (entry point):  nmt-forge synth crk

The engine enforces the emit law on every pack: every generated word must round-trip through the language's analyzer, every closed-class literal must be cited, every filter is named and counted, and every row is stamped synthetic: true — which is exactly why the registry refuses synthetic rows in test sets (tests are real data only). The Plains Cree reference pack lives in crk-translate (nmt_forge_crk); FST models and dictionaries stay separate, user-fetched tools under their own licenses — never bundled.

The full harness referee stack — neural metrics included

forge speaks everything the eval harness speaks, by delegation: the deterministic lanes (chrF++/BLEU/exact-match), the neural lanes — COMET, COMET-QE, MetricX — and the harness's own plugin discovery (FST word-validity, behavioral linters, card-declared metrics):

nmt-forge score --eval-set project-test --hyps decoded.txt \
    --metric chrf++ --metric comet --target-lang iku --card-plugins iku
  chrf++    32.10  [29.4, 34.9] 95% CI
  comet      0.71  [ 0.66,  0.75] 95% CI
  metricx    3.20  [ 2.9,  3.6] 95% CI  (lower = better)

Neural inference runs once; the bootstrap re-averages cached per-entry scores (the harness's own CI pattern). Missing extras report an install fix — never a fabricated number — and checkpoint selection refuses rather than silently switching metrics. Direction is first-class: MetricX's lower-is-better rides the score through rendering, selection, and A/B winners. And discover tells you which lane to believe, from the WMT meta-evaluations:

$ nmt-forge discover iku
  metric trust (Eskimo-Aleut, WMT meta-eval): comet_score r=0.86, ... bleu r=0.163

— for Inuktitut, BLEU barely tracks human judgment while COMET does; for other families it's the reverse; for most low-resource families the honest answer is UNMEASURED. forge surfaces that before you select on anything.

LYSS referees plug in too

LYSS eval-standard linters (harness MetricPlugin protocol — Plains Cree today, more languages later) drop into every scoring surface:

nmt-forge score --eval-set textbook-test --hyps decoded.txt \
    --plugin champollion_lyss.crk.metrics:CrkLinterMetric
  chrf++                            46.02  [43.11, 48.87] 95% CI
  crk_linter:equivalent_match_rate   0.31  [ 0.24,  0.38] 95% CI
  crk_linter · variant_class_counts: LONG_VOWEL_MACRON=9, WORD_ORDER=3

Every numeric aggregate a plugin reports gets a bootstrap CI; per-entry computation runs once (a thousand resamples never re-run an FST); a plugin that says available: false is shown unavailable, never fabricated. And checkpoint selection can use the language's own referee:

"selection": {"metric": "generation:crk_linter:equivalent_match_rate",
              "plugins": ["champollion_lyss.crk.metrics:CrkLinterMetric"]}

A worked example: the half-epoch death

The first CLEAN-protocol Cree run — honest group-disjoint dev, leak-audited mix — died at epoch 0.52 of a 115,000-step plan. Mechanism: the mix was 97.5% tagged synthetic; early in training the model fits the synthetic mass, so dev loss on the 42 real dev sentences bottomed at step ~8k and drifted upward, and patience-6 declared convergence at half an epoch. Every earlier run had hidden this bug by (illegitimately) using the test set as dev. The honest setup is what surfaced it — that's the point of the suite.

forge makes the fix the default, not a flag (schedule-sanity):

[schedule-sanity] train: regime=synthetic-heavy (auto-detected), floor=38,205 of 114,614 planned steps
[schedule-sanity] early stopping ASKED to stop at step 14,000 but the
  schedule floor (38,205) held training, because: the mix is 97.5% synthetic
  and the dev set is REAL: early in training the model fits the synthetic
  mass, so dev loss on real sentences bottoms fast and drifts up — that
  pattern is EXPECTED, not convergence …

The floor is derived from the config (max of one full pass over the mix and 30% of planned steps, capped at 60%) and activates only in the synthetic-heavy regime — auto-detected from the mix, overridable with one word ("regime": "balanced"), never ten flags. Every intervention prints the dev-loss trajectory and the why; nobody should diagnose this from raw logs again.

Training defaults that encode the ledger

Tagged synthetic lanes (Caswell et al. 2019) with gold untagged; gold upweighting with the exposure math written into the manifest; per-kind sampling caps; curriculum stages; a backtranslation lane that leak-audits mono text before spending translation; generation-headroom checks; and a do-not-train gate — datasets the mt-eval registry protects never enter a mix. Details and the config schema: DESIGN.md §6.

What forge refuses to be

Not an evaluator (the harness scores), not a corpus host (manifests are content-free — hashes and counts, never text), not a language-card writer, not a leaderboard.

License

PolyForm Noncommercial 1.0.0 — source-available, free for noncommercial use; using it for a commercial purpose is not covered by this license (relicensed from AGPL-3.0-or-later on 2026-08-17, before any release shipped). Who is covered, in plain words with examples: Who may use this (a summary, not legal advice; LICENSE governs). The harness it uses stays open-source AGPL-3.0-or-later. Analyzer models and dictionaries consumed by packs are upstream artifacts fetched by the user under the upstream's terms.

Metadata

Release files for nmt-forge 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nmt-forge 0.2.0
File Size Uploaded
nmt_forge-0.2.0.tar.gz 510.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nmt-forge 0.2.0
File Interpreter ABI Platform
nmt_forge-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 883.7 kB

Release files / nmt_forge-0.2.0.tar.gz

Download URL nmt_forge-0.2.0.tar.gz
Size 510.0 kB
Tags Source
SHA-256 checksum
How to use checksums
ab8313b2ac4174e73d7bdab03cdc9298e7e447be7f936ad68eeeb8e68cdca903
BLAKE2b-256 checksum
How to use checksums
d1b3fa75bab8b7178ebdeb283facd651046ef686fb6044a5f658d6a28961665b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / nmt_forge-0.2.0-py3-none-any.whl

Download URL nmt_forge-0.2.0-py3-none-any.whl
Size 373.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f579248072c2c308fdc1e0a4dc7844f1d69cb512fc6364b82e03150805597acb
BLAKE2b-256 checksum
How to use checksums
024d2e65bdcc70e4f280470c1df9fdad6fa13339e034f99425c3d8e3487ea850
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page