MT Eval Harness
Most translation tools evaluate Google Translate and DeepL. This harness exists for the languages they leave unverified.
MT Eval Harness is an open-source evaluation framework for developing, benchmarking, and deploying novel machine translation methods — especially for low-resource languages where commercial tools either don't exist or claim coverage that hasn't been independently validated.
Anyone who speaks both languages can contribute a translation method. Prove it works, export it, deploy it. Every method is welcome, human and machine — we care about getting languages translated, not about which tool wins.
This harness is the proving ground inside Champollion: source-available, singly stewarded infrastructure (this harness itself is open source) to create and trust translation test sets for as many language pairs as possible, and to make the whole field navigable — who can translate what, how good each method is on each kind of text, and where the gaps are. It stands on four pillars:
- Solutions-biased pragmatism — every method is welcome, human and machine; the goal is translated languages, not a winning tool.
- Languages as biodata — language data is treated like biodata: precious, personal, and not ours to take.
- Sovereignty is non-negotiable — designed to work with professionals and communities, never hosting their corpora; community ownership and control of language data is a hard constraint, not a courtesy.
- Two tiers of benchmark, the community in control — public benchmarks on open data map and rank every method cheaply and openly; sovereign benchmarks are secret test sets that communities create, own, and control, and that we never see — the gold standard. The infrastructure is source-available and singly stewarded (this harness is open source); the test sets and the methods for a community's language belong to that community, which holds the keys and can revoke them.
Data sovereignty
This project is designed to work with professionals and communities, and it
never hosts their corpora. We treat language data as biodata: the people who provide a corpus hold
the keys to it — and to anything measured against it. Sovereignty is
non-negotiable, not a courtesy. Corpus content is fetched from source with metadata
cards only, never re-hosted by us; non-commercial and community-property datasets
stay out of any prize, API, or commercial path; and a community can revoke access on
its own timeline. The harness itself ships no community-owned data — the
relevant language-validation standard is a separate package you install yourself
when you want it (see
champollion-LYSS); mt-eval run
installs nothing and names the command when something is missing.
Learn more about the wider network at champollion.dev/docs/network.
Why This Exists
There are ~7,000 living languages. Meta's OMT-1600 claims translation coverage for 1,600 of them — but for the ~1,200 in its long tail (our arithmetic: 1,600 minus the 400+ its authors report the models "understand sufficiently well"), quality is below usable thresholds and the model weights are not currently available. For the remaining ~5,400, translation technology doesn't exist at all. Independent evaluation infrastructure is the missing piece.
This harness provides the infrastructure to crowdsource that work:
- Develop a translation method — an LLM prompt, a coached pipeline, a deterministic process, or any combination
- Benchmark it against a reference corpus with standardized metrics (chrF++, exact match, code-switching detection, hallucination detection, terminology adherence, FST acceptance for morphologically-rich languages)
- Export validated methods as champollion plugins
- Deploy to production websites via champollion's translation CLI
┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Develop │───▶│ Benchmark │───▶│ Export │───▶│ Deploy │
│ method.py │ │ mt-eval run │ │ mt-eval │ │ champollion│
│ │ │ mt-eval test│ │ export │ │ translate │
└─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘
Quick Start
PyPI package name: the harness installs as
mt-eval-harness; the command it puts on your PATH ismt-eval. Neither is the unrelatedmtevalpackage on PyPI.
# Install
python3 -m pip install mt-eval-harness
# Set your API key (supports OpenRouter — any model)
export OPENROUTER_API_KEY=sk-or-...
# Or use a direct provider API key
export OPENAI_API_KEY=sk-... # for --provider openai
export ANTHROPIC_API_KEY=sk-ant-... # for --provider anthropic
export GEMINI_API_KEY=AIza-... # for --provider gemini
# ── Zero setup: no corpus files needed ───────────────────────────────
# The harness ships a dataset registry and downloads corpora on demand.
# 1. See what's available (hundreds of language pairs)
mt-eval list datasets
# 2. Run by dataset id — the corpus is fetched from its upstream source,
# rebuilt + cached locally, and licence-gated (--yes accepts the terms).
# A run scores and writes a report; publishing is a SEPARATE step. Add
# --publish to score and publish in one command, or publish later with
# `mt-eval publish <report>` (see QUICKSTART.md §5).
mt-eval run --corpus eval-amh-fra-globalvoices-test-v1 --model gemini-pro --yes
mt-eval run --corpus eval-amh-fra-globalvoices-test-v1 --model gemini-pro --yes --publish
# 3. Contribute compute on the highest-value pairs. `queue` reads the live
# queue (champollion.dev/queue.json, ranked by expected chain value),
# fetching each corpus automatically. Spend only what you choose:
mt-eval queue --budget 2.00 # run from the top until ~$2 of spend
mt-eval queue --top 5 --dry-run # preview the 5 best open items
# ── Or bring your own corpus file ────────────────────────────────────
# Run a translation experiment with optimal defaults
# (batch_size=25, max_tokens=32768, concurrency=8, cache=on)
mt-eval run --corpus data/corpus.json --model gemini-pro
# Multi-model parallel run — all models execute simultaneously
mt-eval run --corpus data/corpus.json \
-m gemini-pro,claude-opus-4.7,gpt-5.5,deepseek-v4-pro
# Direct provider (skip OpenRouter proxy)
mt-eval run --corpus data/corpus.json \
--model openai/gpt-5.5 --provider openai
# Use a standard parallel text corpus (FLORES+, WMT, NTREX)
mt-eval run \
--source-file flores200/dev/eng_Latn.dev \
--reference-file flores200/dev/fra_Latn.dev \
--target-lang French
# Create a contest — public or private
mt-eval contest create --name "EN→CRK Open" \
--corpus edtekla-v1.json --language-pair "en>crk" \
--visibility public
# Analyze the results
mt-eval test eval/logs/harness/run_*.json
# Generate a comparison dashboard
mt-eval dashboard eval/logs/harness/*_report.json
Performance Defaults
The harness is "fast by default, safe by design." Do NOT lower these values unless you have a specific reason.
All defaults are defined as HARNESS_DEFAULTS constants in config.py. Change them in one place and they propagate everywhere.
| Setting | Default | Why |
|---|---|---|
batch_size |
25 | Groups entries into numbered-list prompts. 25× fewer API calls. Proven reliable across all frontier models. Tool-calling auto-overrides to 1. |
max_tokens |
32768 | Generous headroom eliminates truncation risk. Translation outputs are short (1-30 words), so unused tokens cost nothing. |
concurrency |
8 | Parallel batch calls within a single model. Bounded by asyncio.Semaphore for rate limit safety. |
cache_enabled |
True | File-backed cache prevents redundant API calls. Keyed on model + prompt + temperature + language pair. Almost never a reason to disable. |
temperature |
0.0 | Deterministic output for reproducibility. |
Multi-Model Parallelism
For benchmarks, use execute_multi_run() — not a for-loop over execute_run():
from mt_eval_harness.runner import execute_multi_run
from mt_eval_harness.config import RunConfig
configs = [
RunConfig(model="google/gemini-3.1-pro-preview", corpus_path="data.json", ...),
RunConfig(model="anthropic/claude-opus-4.7", corpus_path="data.json", ...),
RunConfig(model="openai/gpt-5.5", corpus_path="data.json", ...),
]
# All models run in parallel — wall-clock = slowest single model
results = await execute_multi_run(configs)
Each model gets its own aiohttp session and semaphore. A 14-model benchmark runs in ~15 minutes parallel vs ~3.5 hours sequential.
What Makes This Different
| Feature | MT Eval Harness | Other MT Eval Tools |
|---|---|---|
| Language-agnostic | Any pair, any script — metrics resolved per language card | Hardcoded for major languages |
| Plugin architecture | Bring your own methods, metrics, tools | Fixed evaluation pipeline |
| Export to production | Direct champollion plugin export | Evaluation only |
| Crowdsource-ready | Prove your method is better, share it | Researcher-only |
| Model-agnostic | Any OpenRouter model (100+), or direct OpenAI/Anthropic/Gemini | Single-vendor |
| Fast by default | batch=25, cache=on, parallel multi-model | Manual optimization |
| COMET with bootstrap CIs | Cached per-entry bootstrap — no redundant neural inference | CIs rarely computed |
| AfriCOMET auto-selection | Auto-selects masakhane/africomet-mtl for 35 African languages |
One model fits all |
| Per-difficulty-tier analysis | Metrics + CIs per translation difficulty level (Tier 1–5) | Corpus-level only |
| Contest infrastructure | Public, private, or team contests with blind evaluation | No contest support |
| Writing style benchmarking | Custom style metrics + brand voice prompt tuning | Quality metrics only |
Core Architecture
mt_eval_harness/
├── runner.py # Orchestrator — strategy-based execution
├── corpus_loader.py # Multi-format dataset loading (JSON/JSONL/TSV/parallel text)
├── champollion_config.py # language-card lookup (the config lane was retired in 0.2.0)
├── pipeline.py # Shared: cache, hooks, enrichment, logging
├── strategies/ # Execution backends
│ ├── single.py # One entry per API call
│ ├── batch.py # Multiple entries per call
│ ├── tool_call.py # Multi-round tool-calling
│ └── method_strategy.py # Custom TranslationMethod plugins
├── providers/ # Multi-provider LLM abstraction
│ ├── base.py # LLMProvider ABC — uniform interface
│ ├── registry.py # get_provider() factory
│ ├── openrouter.py # OpenRouter (default — proxies any model)
│ ├── openai_provider.py # Direct OpenAI API
│ ├── anthropic_provider.py # Direct Anthropic Messages API
│ └── gemini_provider.py # Direct Google Gemini API
├── tester.py # Offline metric computation
├── exporter.py # champollion plugin packaging
├── api.py # OpenRouter HTTP client (used by openrouter provider)
├── cache.py # Deterministic result caching
├── config.py # Typed configuration + protocols
├── language_cards.py # Language card loader + validation
├── cli.py # Command-line interface
├── dashboard.py # Interactive HTML report generator
│ # NOTE: language-specific eval standards (e.g. the Plains Cree LYSS linter +
│ # semantic validator) are NOT bundled here. They live in the separate
│ # champollion-lyss package, installed by the user from the language card's
│ # evalStandard — the core wheel ships no language-specific scorer code.
└── plugins/ # Extension protocols
├── prompts.py # PromptProvider
├── champollion_prompts.py # retired ChampollionPromptProvider stub (0.2.0)
├── metrics.py # MetricPlugin
├── hooks.py # PostTranslationHook
├── tools.py # ToolProvider
├── giellalt_fst.py # GiellaLT FST morphological validity
├── code_switching.py # Code-switching detection
├── hallucination.py # Hallucination detection
├── terminology.py # Terminology adherence
├── double_pass_compliance.py # Placeholder/quote/casing compliance (planned; no run loads it)
├── writing_style.py # Writing style consistency
└── fst_installer.py # FST binary installer
Extending the Harness
The harness exposes four plugin protocols. If your class has the right method signatures, it works — no inheritance required.
TranslationMethod has three required members — a name attribute and a method_card() method alongside translate (the runner reads both for run IDs, logs, and provenance):
from mt_eval_harness.config import TranslationMethod
class MyTranslationPipeline:
"""Custom pipeline — implements TranslationMethod protocol."""
name = "My Translation Pipeline" # required — run IDs and logs
def method_card(self) -> dict | None:
# required — provenance metadata (or None for no card)
return {"method_id": "my-pipeline-v1", "name": self.name,
"class": "pipeline"}
async def translate(self, entries: list[dict], config) -> list[dict]:
# Your translation logic here
return [{"id": e["id"], "predicted": "..."} for e in entries]
See GUIDE.md for full plugin documentation.
Installation
# Install the harness (PyPI dist: mt-eval-harness; the command is mt-eval)
python3 -m pip install mt-eval-harness
# Interactive setup — installs optional deps with explanations
mt-eval setup
# Or install everything at once, no prompts
mt-eval setup --all
# Check what's installed
mt-eval setup --status
Requirements: Python 3.11+ · At least one API key:
| Provider | Env Var | Flag |
|---|---|---|
| OpenRouter (default) | OPENROUTER_API_KEY |
--provider openrouter |
| OpenAI (direct) | OPENAI_API_KEY |
--provider openai |
| Anthropic (direct) | ANTHROPIC_API_KEY |
--provider anthropic |
| Gemini (direct) | GEMINI_API_KEY or GOOGLE_API_KEY |
--provider gemini |
Ship lean, install on consent. The harness core has minimal dependencies. Optional capabilities (COMET neural metric, FST morphological validation) install interactively via
mt-eval setup— or on-the-fly when the harness detects they'd improve your eval. You never need to know specific pip commands.
Manual pip install (if you prefer)
python3 -m pip install 'mt-eval-harness[comet]' # COMET neural metric + AfriCOMET
python3 -m pip install 'mt-eval-harness[fst]' # FST morphological validation
python3 -m pip install -e ".[dev]" # Development
Documentation
- GUIDE.md — Full user guide and API reference
- CHANGELOG.md — Versioned change log
- CONTRIBUTING.md — Development standards and contribution workflow
- Plugin Specification — champollion plugin export format (§9 of benchmark spec)
- Scoring Specification — SSOT for metrics and the scoring standard (chrF++ headline; the retired composite and quality tiers kept only for verifying legacy cards)
Contests & Leaderboards
A contest is sovereign hosting (founder ruling, 2026-09-06): an entry is a MODEL or a METHOD handed to the organizer's own node, which executes it against a sealed set that never leaves that machine. The open leaderboard — self- reported cards indexed by corpus × pair direction — is a different thing and is not a contest.
# ORGANIZER: split, seal and register in one command. --prize-disposition
# declares what happens to a winning entry — pass_to_holders (it passes to the
# benchmark holders, who keep it) | retain_ip (the entrant keeps ownership) |
# release_open (the entrant must publish it openly); no disposition = no prize.
mt-eval contest prepare --corpus master.json --slug en-crk-open \
--name "EN→CRK Open" --pair "eng>crk" \
--dev-size 200 --secret-size 300 --seed 42 --qualifier-threshold 35 \
--license <SPDX id the rights-holder grants> \
--custodian-group <opaque-id> --threshold-pubkey key.pub.json \
--out ./contest --self-serve \
--prize-disposition retain_ip --results-visibility hidden_until_close \
--anonymize-until-close
# ENTRANT: qualify in public first — the node re-executes this claim itself
mt-eval contest qualify en-crk-open --dev dev-hyps.txt \
--dev-corpus ./contest/public/qual-en-crk-open-2026.json \
--system "my-nmt" --method-class pipeline
# ENTRANT: hand over the entry (weights, or code the node runs offline)
mt-eval contest submit-model en-crk-open --model-dir ./model …
mt-eval contest submit-method en-crk-open --method-dir ./method --dockerfile ./Dockerfile …
# ORGANIZER: a node config (fill in its <...> values), the custodian
# decision, then execute one authorized entry
mt-eval node init
mt-eval node list
mt-eval node approve <authreq-id> --actor "custodian-1"
mt-eval node run-method <authreq-id>
# Rank (verified-only by default; ties by per-segment AR test → CI overlap →
# point equality, competition numbering 1,1,3), then close and export
mt-eval contest rank en-crk-open --json
mt-eval contest close en-crk-open # publishes withheld results, then freezes
mt-eval contest export en-crk-open --format csv --out results.csv
RETIRED 2026-09-06 (founder ruling R2): contest submit, which linked a self-reported card.
RETIRED 2026-09-06 (founder ruling R2): contest submit-hypotheses, which uploaded translations. Both verbs were deleted.
Visibility modes: public (anyone), private (invite-only, blind), team (org-scoped). The primary metric is recorded per contest (--primary-metric chrf_plus_plus|bleu|comet_score|…, default chrF++; the retired composite is refused for a new contest). See GUIDE.md § 14 for the full ranking algorithm and lifecycle.
Currently In Development
We're actively using this harness to develop and evaluate Plains Cree (crk) translation methods — including our own FST-gated pipeline and external systems like Meta's OMT-1600 (which includes CRK at R1 tier). The harness provides independent evaluation with morphological validation that standard metrics cannot.
License
AGPL-3.0-or-later (see LICENSE). The AGPL allows commercial use on its own
terms, including its network-use clause (§13): if you change the harness and let
people use it over a network, you must offer them your changed source. How this
license sits beside the noncommercial ones on the CLI, the MCP server and
nmt-forge: Who may use this (a summary, not legal advice; the license
text governs).
Eval-Standard Plugin exception: as an additional permission under AGPL-3.0 §7,
the Harness may be combined with separately-licensed eval-standard plugins (e.g.
champollion-lyss) that interoperate only through its public plugin interface
(the champollion.eval_standards entry point, the MetricPlugin protocol, the
language-card eval-metric loader, and the documented FST-installer helpers). Such
plugins may use licenses incompatible with the AGPL, including noncommercial ones —
the Harness itself stays AGPL. Full terms: LICENSE-EXCEPTION.md.
Metadata
Release files for mt-eval-harness 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mt_eval_harness-0.2.0.tar.gz | 2.8 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mt_eval_harness-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 4.9 MB
Release files / mt_eval_harness-0.2.0.tar.gz
| Download URL | mt_eval_harness-0.2.0.tar.gz |
|---|---|
| Size | 2.8 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8281fbdb7e9e835888afc06a7aaa66cfd7af3a120dcd32ec511237667ae6cde7
|
|
BLAKE2b-256 checksum How to use checksums |
8003933b80f0c6db9aa83706ef30a736df189aea53a5c586589910f02b01b706
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / mt_eval_harness-0.2.0-py3-none-any.whl
| Download URL | mt_eval_harness-0.2.0-py3-none-any.whl |
|---|---|
| Size | 2.2 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a7f9e8d6ee02f6abd39c70375e66f6346ba3991596f8f0c2d7c7da6359322804
|
|
BLAKE2b-256 checksum How to use checksums |
d9fb1f96cd8f103e74e0b741286f3ba0849da2200eda962fe293918d23c26e02
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|