Skip to main content

MLCommons Endpoints Submission Tools

A Python package with two tools for managing MLPerf Endpoints benchmark submissions:

  • endpoints-submission-cli — registers benchmark runs, assembles submission packages, runs compliance checks, and opens GitHub pull requests via the PRISM API.
  • submission-checker — validates a submission folder against the §9.1 automated compliance rules before or after upload.

Installation

With pip:

pip install endpoints-submission-cli

From source (editable):

pip install -e ".[dev]"

With uv:

uv sync --extra dev

endpoints-submission-cli

Requirements

  • Python 3.10 or later
  • gh CLI — required for creating, updating, and withdrawing submissions

Authentication

Every command requires a PRISM API token in mlc_… format. Supply it as an env var or pass --token per command:

# Persistent (add to shell profile)
export PRISM_USER_API_TOKEN=mlc_your_token_here

# Per-command override
endpoints-submission-cli runs list --token mlc_your_token_here

Configuration

Environment variable Default Description
PRISM_USER_API_TOKEN API key. Required unless --token is passed.

Add to your shell profile for a persistent setup:

export PRISM_USER_API_TOKEN=mlc_your_token_here

Quick start

# 1. Verify connectivity
endpoints-submission-cli runs list

# 2. Register a benchmark run from a local result folder
endpoints-submission-cli runs create --path /results/llama3_h100_c4
# → Run created: d5d9873e-5eca-4f8d-a487-4be1cb8b440c
RUN_ID=d5d9873e-5eca-4f8d-a487-4be1cb8b440c

# 3. Create a submission (assembles, checks, uploads, opens PR)
endpoints-submission-cli submissions create \
  --division standardized \
  --availability available \
  --run-ids $RUN_ID
# → Submission created: a1b2c3d4-…
SUB_ID=a1b2c3d4-e5f6-7890-abcd-ef1234567890

# 4. Withdraw if needed
endpoints-submission-cli submissions withdraw --submission-id $SUB_ID

Command reference

endpoints-submission-cli
├── runs
│   ├── list        List all runs
│   ├── create      Register a run from a local folder
│   ├── get         Fetch run details
│   ├── delete      Delete a run and its archive
│   ├── pin         Pin a run (prevent expiry)
│   └── unpin       Restore normal expiry
└── submissions
    ├── list        List all submissions
    ├── create      Create a submission from runs (full pipeline)
    ├── get         Fetch submission details
    ├── update      Update run list or metadata
    ├── withdraw    Withdraw a submission
    └── remove-run  Remove a run from a submission

Use --help on any command for full flag details:

endpoints-submission-cli submissions create --help

submission-checker

CLI tool for validating MLPerf Endpoints submissions against the §9.1 automated compliance checks.

Usage

Check a submission

submission-checker check /path/to/submission

The path may be the submitting organisation's directory or a <submission_id>/ directory below it; a submission root is the level holding results/ and docs/ (§8.1).

Options:

Flag Description
--strict Treat warnings as errors (exit 1 on any warning)
--quiet / -q Suppress INFO-level passing checks
--output FILE / -o FILE Write full results as JSON to FILE
--seed-sets FILE Published seed sets to check against (§4.6). Defaults to the bundled set; also settable via $MLPERF_ENDPOINTS_SEED_SETS.
--approved-drafters FILE Published approved drafters (§2.9.4). Defaults to the bundled list; also settable via $MLPERF_ENDPOINTS_APPROVED_DRAFTERS.

Exit codes: 0 = all checks passed, 1 = one or more errors (or warnings with --strict).

Show region boundaries

submission-checker regions --max-concurrency 1024 --min-concurrency 16

Prints the concurrency range for each region given a (C_max, C_min) pair, using the §5.5 reference algorithm. --min-concurrency defaults to 32; in a real submission C_min is derived from the lowest measurement point rather than declared (§5.4).

Required Files in submission structure

Layout as of mlcommons/endpoints_policies PR #119: there is no per-system file — every Pareto point carries its own system_desc.json.

<submitting_organization>/
└── <submission_id>/
    ├── src/
    │   └── <implementation>/         # §2.2.1 — README.md + endpoint interface code
    │       └── README.md
    ├── docs/                         # calibration, software disclosure, …
    └── results/
        └── <system>/
            ├── system_power.json      # §4.5.2 — REQUIRED, one per system
            └── <model_name>/
                └── r<N>/             # one directory per concurrency level
                    ├── point.yaml            # §8.3 measurement-point disclosure
                    ├── system_desc.json      # §8.2 — per point since PR #119
                    ├── result_summary.json   # aggregate metrics
                    ├── accuracy_results.json # §6.6 accuracy results
                    ├── config.yaml           # OPTIONAL as of v1.0
                    └── server_configs/       # OPTIONAL, submitter-defined

src/ and docs/ are shared across the whole submission; each point.yaml names them via shared_src and shared_docs, which must resolve to a directory under the submission root (§9.1).

What gets checked

Structure

Rule Spec Description
path-exists §1 Submission root directory exists
required-dir §1 results/ and docs/ present
src-dir §2.2.1 src/ present with at least one implementation directory
src-readme §2.2.1 Each src/<implementation>/ has a README.md
system-results-dir §1 At least one results/<system>/ directory exists
benchmark-model-dir §1 At least one benchmark-model directory per system
point-dirs §1 At least one r<N>/ Pareto-point directory per model
measurement-points-present §1 Every r<N>/ carries a point.yaml
point-dirname-concurrency §1 r<N>/ name matches the declared concurrency (warn)
result-summary-present §1 result_summary.json exists for each point
shared-path-resolution §9.1 shared_src / shared_docs resolve under the submission root

System description (§8.2)

Rule Spec Description
system-description-present §8.2 Every point has a system_desc.json
system-description-valid §8.2 It parses against the SystemDescription schema
system-description-consistency §8.5 Every point of a curve describes the same system
model-name-valid §2 model_name is one of the round's supported models
model-name-consistency §16 It matches the results directory name
max-concurrency-declared §7 max_supported_concurrency (C_max) present and > 32
tps-utilization §8.2 Equals system_tps / max(system_tps) over the point's own curve
power-descriptor §4.5.2 system_power.json present per system and states a derivable power
power-estimated §4.5.2 Flags component groups left for MLCommons to auto-populate (warn)

Regions (§5)

C_min is derived from the submission's own points in v1.0, not declared, so the boundaries differ per curve. The 10 % margin above C_max is its own region and does not satisfy High Concurrency coverage.

Rule Spec Description
region-basis §5.4 Reports the derived C_min and how many points it came from
region-computation §5.5 (C_max, C_min) is a valid input to the reference algorithm
concurrency-in-range §9.1 Each concurrency falls in a valid region, margin included
region-declared §8.3 Declared region is one of the spec's values
region-placement §8.3 Declared region matches the computed one (warn)
offline-declared §5.7 offline is dedicated, elected, or none
offline-point-present §5.7 Exactly one Offline point; elected sits on the C_max point
offline-ordering §5.7.2 Offline beats C_max on throughput (2% tolerance) and concurrency (warn)
ultra-low-concurrency-coverage §5.4 At least one point at concurrency ≤ 32
low-concurrency-coverage §9.1 At least one point in the Low Concurrency region
med-concurrency-coverage §9.1 At least one point in the Medium Concurrency region
high-concurrency-coverage §9.1 At least one point in the High Concurrency region
point-count §5.3 7–32 measurement points; 8 with a dedicated Offline run
point-cap §2, §8 Point count does not exceed 32

Measurement points (§8.3, §6)

Rule Spec Description
point-config-valid §8.3 point.yaml parses against the PointConfig schema
point-disclosure-complete §8.3 Every required §8.3 disclosure field is present
load-pattern §6.1 load_pattern is concurrency with a positive level
streaming-config §6.5 stream_all_chunks is True
point-duration §6.2 Steady-state window's issue-time span meets the region minimum (warn)
steady-state-valid §4.4 status, verdict and gating state use the spec's vocabulary
steady-state-consistency §4.4 The reported status agrees with the window it describes
steady-state-basis §4.4 Which basis supplies the official result; flags fallbacks and drift (warn)
min-query-count §6.4 n_samples_completed meets the dataset minimum
warmup-present §6.3.3 Warmup declaration present
warmup-logs-retained §6.3.2 Warmup log retention declared (warn)
warmup-salt §6.3.3 Warns when the warmup salt is enabled
config-consistency-dataset §16 All points use the same dataset

Seed binding (§4.6)

Rule Spec Description
seed-set-consistency §9.1 Every point records the same seed set
seed-set-membership §9.1 The bound set is one MLCommons published
seed-runtime-match §2.1.1 The RNG seeds equal the bound set's values
target-cohort §4.6 target_cohort matches YYYY-MM-C0 / YYYY-MM-C1
seed-set-adoption §4.6 target_cohort falls inside the set's four-cohort adoption window
seed-config-legacy §4.6 v0.7 fallback: seeds == 42 when no seed_set is declared
seed-set-registry §4.6 Warns when the seed-set file itself cannot be read

Speculative decoding (§2.9.4)

Rule Spec Description
approved-drafter §2.9.4 The declared drafter is on the benchmark's published list
drafter-approval-lead-time §2.9.4 Approved at least two cohorts before target_cohort
drafter-list-registry §2.9.4 Warns when the drafter list itself cannot be read

The approved list ships as data (src/submission_checker/data/approved_drafters.yaml) and is empty — §2.9.4's list has not been published yet, and an empty registry means speculative decoding is not permitted for any benchmark, which is §2.9.4's own rule for a benchmark with no approved drafter. Point --approved-drafters FILE or $MLPERF_ENDPOINTS_APPROVED_DRAFTERS at a published list.

The published sets ship as data (src/submission_checker/data/seed_sets.yaml), mirrored from the policies repo's seedset.yaml. The file's cohort-id is the cohort its sets were published for; §4.6's four-cohort adoption window is derived from it. Point --seed-sets FILE or $MLPERF_ENDPOINTS_SEED_SETS at a newer file to check against a set published after this release.

Metrics (§9.1, §14)

Rule Spec Description
result-file-valid §8.3 result_summary.json parses against PointSummary
metric-consistency-duration §14 duration_ns > 0
metric-consistency-accounting §14 completed + failed == issued
metric-consistency-output-tokens §14 total_output_tokens ≥ 0
metric-consistency-system-tps §9.1 Stored system_tps matches the derived value
metric-consistency-tpot-p90 §9.1 Reported TPOT P90 present, finite, strictly positive
metric-consistency-tps-per-user §9.1 Stored tps_per_user matches 1000 / tpot_p90_ms
metric-consistency-tps-per-kw §4.5.3 Stored system_tps_per_kw matches system_tps / provisioned_power_kw
agentic-metric-consistency §4.1 e2e_avg_interactivity is derivable from its reported inputs

Accuracy (§15)

Rule Spec Description
accuracy-present §15 At least one model in the submission carries accuracy results
accuracy-coverage §5.3 Accuracy at each of the four mandatory bands, plus the Offline point
accuracy-valid §15 accuracy_results.json parses correctly
accuracy-sample-count §15 Issued sample count meets the model's minimum
accuracy-gate §15 Score meets the benchmark quality target

Programmatic API

from submission_checker import SubmissionChecker, Report

checker = SubmissionChecker(Path("/submissions/acme_corp"))
report = checker.run()

if report.passed:
    print("All checks passed")
else:
    for result in report.errors:
        print(f"[{result.rule}] {result.message}")

The Report object also exposes report.warnings and serialises cleanly via report.model_dump_json().


Development

uv run pytest                          # run all tests
uv run pytest --no-cov -x             # fast fail on first error
uv run ruff check src/ tests/          # lint
uv run ruff format src/ tests/         # auto-format

Release files for endpoints-submission-cli 1.0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for endpoints-submission-cli 1.0.1.0
File Size Uploaded
endpoints_submission_cli-1.0.1.0.tar.gz 345.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for endpoints-submission-cli 1.0.1.0
File Interpreter ABI Platform
endpoints_submission_cli-1.0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 484.7 kB

Release files / endpoints_submission_cli-1.0.1.0.tar.gz

Download URL endpoints_submission_cli-1.0.1.0.tar.gz
Size 345.3 kB
Tags Source
SHA-256 checksum
How to use checksums
839de89197017614dfef34d549f78365c7df7a684215cdc117c48f6080de3eab
BLAKE2b-256 checksum
How to use checksums
77309e056c8114a5da2a1802768935a12db1d8de23ac376198ec98ca4410c5ff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / endpoints_submission_cli-1.0.1.0-py3-none-any.whl

Download URL endpoints_submission_cli-1.0.1.0-py3-none-any.whl
Size 139.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3aed474274e633575a215383351f319beb46ad102933f7cfed9ef61680dc08cd
BLAKE2b-256 checksum
How to use checksums
ce052d64e931537ab314722fecb3c70540ac326436a031e21cf9c2cfcf616a69
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page