Skip to main content

ES-MoE Toolkit

Drop-in expert-sparse MoE block for Ultralytics YOLO


Python PyPI Docs Colab DeepWiki License

One call adds the block, the router loss reaches the optimiser, and every number here has a run record behind it.

Docs: English · 中文 · Quick start in Colab: notebooks/quickstart.ipynb · Ask questions about the code: DeepWiki.


Install

pip install esmoe

The distribution, the import and the CLI are all esmoe; the project is written ES-MoE in prose.

Requires ultralytics.

Use

One call covers register, graft, build and wire:

import esmoe

model = esmoe.equip("yolo11n.yaml", weight=0.01)
model.train(data="coco8.yaml", epochs=10)

Or take the steps apart when you need control over each one:

from ultralytics import YOLO
import esmoe

esmoe.inject_esmoe()                                    # make `ESMoE` resolvable in model.yaml
esmoe.graft("yolov8n.yaml", out="v8-esmoe.yaml", at=[4, 6])  # insert blocks, renumber the head
model = YOLO("v8-esmoe.yaml")
esmoe.attach_aux_loss(model, weight=0.01)               # router loss joins the training loss

From the shell:

esmoe graft yolo11n.yaml -o yolo11n-esmoe.yaml -e 4 -k 2 --at backbone_end

attach_aux_loss adds an esmoe_aux entry to the trainer's loss table, so a non-zero, back-propagated auxiliary term shows up in results.csv rather than merely in a config.

Written by hand, a grafted config layer is just:

[-1, 1, ESMoE, [4, 2]]                              # num_experts, top_k
[-1, 1, ESMoE, [4, 2, null, {out_norm: true}]]      # ... and settings the trainer must keep
[-1, 1, ESMoE, [4, 2, null, {out_channels: 320}]]   # widened ...
[-1, 1, Index, [320, 0]]                            # ... and the official layer that declares it

The block is channel preserving and infers its width on the first forward, which is what lets stock parse_model size it without a patch. A widened block instead hands a one-element list to the official Index layer after it, whose declared width parse_model does read; graft(out_channels=...) writes both. Settings belong in the config because the trainer rebuilds the model from it, dropping anything set on the blocks beforehand.

Extend

Experts and the balancing objective are plain callables, so a variant is a few lines:

esmoe.ESMoE(num_experts=4, top_k=2, expert=MyExpert, balance=my_balance_fn)

MyExpert(c1, c2, k) -> Module, my_balance_fn(probs, gate) -> scalar. To train with them, pass them to graft or equip: a function or class defined at module level goes into the config as module:qualname, and every rebuild of the model -- the trainer's, each DDP worker's -- imports it back from that name. A lambda or anything defined in __main__ is refused when grafting. esmoe.blocks(model) walks every block in a model, esmoe.collect_aux_loss(model) returns the current step's router loss for custom training loops, and block.spec() reports the settings a block is holding.

Compatibility

backbone build + forward grafted config aux loss in training protocol runs
YOLOv5 yes yes yes yes
YOLOv8 yes yes yes yes
YOLOv9 yes yes yes yes
YOLOv10 yes yes yes yes
YOLO11 yes yes yes yes
YOLO12 yes yes yes yes
YOLO26 yes yes yes yes
YOLO-Master (fork) yes yes yes no

The fork row has no protocol runs of this package: the same-configuration comparison trains upstream's own blocks on the fork and this package's blocks on official ultralytics.

Verified by tests/test_ultralytics.py on ultralytics 8.4.101 and 8.4.132, which report loss items in two different shapes; both are handled. The training column is backed by real 1-epoch VisDrone runs on four generations (results/*-compat-*.json) and by the 120-epoch protocol runs on all seven, each logging a non-zero train/esmoe_aux.

Graft and forward are exercised on every row in CI. The last column separates "the block builds and trains" from "we ran the full budget-fair protocol on it". The YOLO-Master row runs against the fork's vendored ultralytics: scripts/fork_smoke.py grafts their yolo-master-n.yaml, trains one epoch with a non-zero esmoe_aux, and builds their own ES_MOE config alongside ours.

DDP works: attach_aux_loss routes model.train() through a trainer class that lives in esmoe.trainer, so the worker processes ultralytics spawns register the block and the auxiliary loss before they build. Verified by scripts/verify.py (the real worker file in a fresh interpreter; two gloo ranks with agreeing router gradients). Inside a process group an expert that no image routed to joins the graph at zero weight, so the settings ultralytics uses under compile=True (find_unused_parameters=False, static_graph=True) train as well; tests/test_distributed.py checks that with two gloo ranks, and that the block compiles and agrees with eager.

To compare against YOLO-Master under one configuration, attach_aux_loss(model, weight=1.0, recipe="upstream") trains the way its trainer trains a routed model. The auxiliary term is normalised by its running magnitude and capped at 3.0, routers get half the learning rate outside Muon, and experts stay frozen for three epochs. scripts/train.py --upstream runs the same protocol on the fork itself, and scripts/same_config.py pairs the two.

Selected default

ESMoE(num_experts=4, top_k=2) with attach_aux_loss(weight=0.01), chosen under one budget over 2/4/8-expert and top-1 variants. Under the repository protocol (VisDrone, imgsz 800, 120 epochs, three seeds) the matrix runs to seven backbone generations × three arms × three seeds and more, 142 runs, the last twenty-nine of them the same-configuration comparison against YOLO-Master's own fork: twelve in each of two rounds, and five FP32 repeats that measure the noise floor. What separates a positive cell from a negative one is what the backbone ends in, not how new it is: the default wiring is positive on the SPPF family (+0.0055 v5n, +0.0025 v8n, +0.0025 v9t), sits on zero once the end is an attention block (−0.0002 v10n, +0.0013 11n), and is negative on area attention and the E2E head (−0.0018 12n, −0.0034 26n). The block costs +10.4% parameters and about 9% more wall-clock per epoch.

Paired mAP50 delta by backbone generation

Each dot is one seed, each bar the mean of three. The seeds routinely straddle zero even where the mean does not, which is as far as a three-seed protocol can read: an interactive version carries the per-seed values and the second metric. Where the damage lands depends on the backbone: v8n loses large objects (APl −0.010, 0/3), 26n loses small ones (APs −0.0045, 0/3), 12n is direction-unstable.

Further arms ask what upstream's own settings are worth. Two of them are internal to the block and both help; upstream's layout of four blocks per backbone is negative on both metrics at 0/3 on both backbones it ran on, and stays negative with each block's weight cut to a quarter so the auxiliary total matches one block — the count is what costs, not the pressure. That holds under this package's recipe: with upstream's whole recipe the same four blocks are positive on every seed in both frameworks (same-configuration rounds seven and eight).

Paired mAP50 delta for the upstream-alignment arms

How concentrated the dispatch is does not predict accuracy: r = +0.044 over the 81 runs that have both a routing analysis and a paired delta. What the balancing term does secure is that no expert dies. Without it all six checkpoints lose two of four experts; the Switch term at 0.01 leaves none dead in 66. The paper's objective and upstream's read the gate, which is renormalised over the top-K and so has no gradient for an expert outside it: five of six such checkpoints have a dead expert, even at more pressure than Switch, and so do all six same-configuration B checkpoints, trained with upstream's recipe.

The default graft leaves consumers that name the old backbone end by index — YOLOv8's P5 lateral among them — reading the pre-block tensor; graft(..., rewire=True) retargets them. That arm is the only 3/3 one on v8n (+0.0036) and pulls 12n and 26n back to near parity (+0.0001 and −0.0005); only on 11n does it trail the default. Verdicts against the pre-registered lines: docs/JUDGMENT.md. Full tables: docs/SELECTION.en.md, results/buckets.md, results/routing.md, results/report.md.

Develop and reproduce

uv sync --group dev
uv run pytest -q
uv run python scripts/capture_env.py                  # freeze environment into results/env/
EPOCHS=20 FRACTION=0.25 SEEDS="0 1 2" uv run bash scripts/sweep.sh
uv run python scripts/report.py                       # results/summary.md

Every run writes one machine-readable record to results/ (config, dataset, hardware, budget, seed, metrics, artifact, status, limitation). Read the limitations before quoting any number. The 142 protocol checkpoints, with each run's arguments and per-epoch curve, live on the checkpoints branch (Git LFS, orphan — main stays small), flat-named to match the run records.

Linked projects

License

AGPL-3.0-only, matching the Ultralytics ecosystem it builds on.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

esmoe-1.0.0.tar.gz (68.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

esmoe-1.0.0-py3-none-any.whl (38.3 kB view details)

Uploaded Python 3

File details

Details for the file esmoe-1.0.0.tar.gz.

File metadata

  • Download URL: esmoe-1.0.0.tar.gz
  • Upload date:
  • Size: 68.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for esmoe-1.0.0.tar.gz
Algorithm Hash digest
SHA256 7110d7c9e2994a95d24f7760d8938f4e40dd565436ec7ccf415e8618c964d9e8
MD5 f470a974f725227688a38949f8a93ec3
BLAKE2b-256 13f6b0b05bd4e6974a6203bac6f104e9fa60d84eae3a0fcca60fd871fb604b01

See more details on using hashes here.

Provenance

The following attestation bundles were made for esmoe-1.0.0.tar.gz:

Publisher: release.yml on Lfan-ke/ES-MoE

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file esmoe-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: esmoe-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 38.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for esmoe-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b8d7c2fdca8c0d01e210fee3d43209a6676e2de8890ef4e6a9518dc522e40cc2
MD5 53b551180773d2f3fcdc2c206c00578a
BLAKE2b-256 270a853d9756602780e7808fbc8de929cf455c9461703db4ee31aa50c3e32da5

See more details on using hashes here.

Provenance

The following attestation bundles were made for esmoe-1.0.0-py3-none-any.whl:

Publisher: release.yml on Lfan-ke/ES-MoE

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page