Skip to main content

mobius

CI L4: Golden Checkpoint Parity (GPU) L5: End-to-End Generation (GPU) Nightly L2 Architecture Validation

ONNX model definitions for GenAI using the onnxscript.nn API.

Overview

This package provides model definitions for generative AI architectures — LLMs, MoE, multimodal, encoder-only, encoder-decoder, vision, audio, and diffusion models — built directly as ONNX graphs using onnxscript.nn.Module. Rather than tracing or exporting PyTorch models, it constructs the ONNX graph declaratively, then applies pretrained HuggingFace weights.

Supports building ONNX models from HuggingFace model IDs with automatic weight downloading, dtype casting (including bfloat16 via ir.LazyTensor), and multi-component export for pipelines.

📖 Documentation · 📦 Supported Models

Highlighted Models

Category Examples
Text Generation Llama 2/3/4, Mistral, Qwen 2/2.5/3/3.5/3.6, Phi-3/3.5, Gemma 1/2/3/4, Granite, GPT-2, OPT, OLMo, SmolLM3, and many more
Mixture of Experts PhiMoE, GPTOSS, Mixtral, OLMoE, DeepSeek-V2/V3, Qwen2-MoE, Qwen3-MoE, Qwen3-Next, GLM-4-MoE, Arctic, DBRX, Jamba
Multimodal Gemma 3/4, Phi-4MM (vision + audio + LoRA), LLaVA, InternVL2, Qwen2.5-VL, Qwen3-VL, Qwen3.5/3.6-VL, Pixtral
Encoder-only BERT, RoBERTa, ALBERT, DeBERTa, DistilBERT, ELECTRA, XLNet
Encoder-Decoder BART, T5/mT5, Marian, M2M-100, Pegasus, BigBird-Pegasus
Speech-to-Text Whisper, FastConformer-RNNT, FunASR, Qwen3-ASR, SenseVoice
Audio Wav2Vec2, HuBERT, WavLM, SpeechT5
Vision ViT, BEiT, DeiT, DINOv2, Swin, CLIP, SigLIP
Diffusion Stable Diffusion (UNet + VAE + ControlNet), Flux, SD3, DiT, QwenImage, HunyuanDiT, CogVideoX
Adapters T2I-Adapter, IP-Adapter

Supports 290+ Transformers model types and 10 Diffusers component types across 40+ task types and 100+ reusable components.

See the model documentation for the complete list.

Installation

pip install -e .

For running tests:

pip install -e ".[testing]"

Quick Start

Python API

from mobius import build

# Build a model package with weights
pkg = build("meta-llama/Llama-3.2-1B")
pkg.save("output/llama-3.2-1b/")

Static cache (opt-in) pre-allocates fixed-size KV cache buffers, which is useful when you know the maximum sequence length up front:

from mobius import build, CausalLMTask

task = CausalLMTask(static_cache=True, max_seq_len=2048)
pkg = build("meta-llama/Llama-3.2-1B", task=task)
pkg.save("output/llama-3.2-1b-static/")

EP-aware optimization generates graphs tuned for a specific runtime execution provider. Pass execution_provider to target CUDA, DirectML, WebGPU, and more — each with the right set of fused kernels and lowering passes applied automatically:

from mobius import build

# CUDA: GQA fusion, SkipLayerNorm, PackQKV
pkg = build("meta-llama/Llama-3.2-1B",
            execution_provider="cuda", dtype="f16")

# WebGPU: GQA fusion, Shape ops replaced with portable alternatives
pkg = build("meta-llama/Llama-3.2-1B",
            execution_provider="webgpu", dtype="f16")

See the EP quickstart and full EP reference for all supported EPs and options.

CLI

mobius build --model Qwen/Qwen2.5-0.5B output_dir/

# Build for CUDA with f16
mobius build --model meta-llama/Llama-3.2-1B output_dir/ --ep cuda --dtype f16

# Build a diffusers pipeline (all components)
mobius build --model Qwen/Qwen-Image-2512 output_dir/

# Build encoder-decoder model (produces encoder/model.onnx + decoder/model.onnx)
mobius build --model openai/whisper-tiny output_dir/

See the CLI Reference for all subcommands and flags.

Examples

Architecture

HuggingFace Hub
       │
       ▼
 ArchitectureConfig ◄── from_transformers() / from_diffusers()
       │
       ▼
 Model Module ◄── Reusable Components (Attention, MLP, RMSNorm, RoPE, …)
       │
       ▼
 Task ◄── CausalLMTask, VisionLanguageTask, VAETask, DenoisingTask, …
       │
       ▼
 ONNX Model ◄── preprocess_weights() + apply_weights()

The package is organised into four layers:

  • Components — onnxscript.nn.Module building blocks (Attention, MLP, DecoderLayer, RoPE, VisionEncoder, MoELayer, …)
  • Models — Full architectures composed from components
  • Tasks — Define the ONNX graph I/O contract (inputs, outputs, KV cache)
  • Registry — Maps HuggingFace model_type strings to model classes

See the design document for details.

Development

# Unit tests (fast, no network needed)
pytest tests/build_graph_test.py -v

# Integration tests (downloads models)
pytest tests/integration_test.py -m integration -v

# All unit tests (components, configs, tasks, models)
pytest src tests -m "not integration" -v

# Linting
lintrunner f --all-files

Adding a new model

See the AI-assisted model support strategy and the developer skills in .agents/skills/:

Skill Use when
adding-a-new-model Adding any new HuggingFace model architecture
reusable-components Creating or extending components
moe-models Adding a Mixture-of-Experts model
multimodal-models Adding a vision-language model
writing-tests Writing unit or integration tests
writing-rewrite-rules Adding ONNX graph rewrite rules

Metadata

Release files for mobius-onnx 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mobius-onnx 0.1.0
File Size Uploaded
mobius_onnx-0.1.0.tar.gz 1.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for mobius-onnx 0.1.0
File Interpreter ABI Platform
mobius_onnx-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.4 MB

Release files / mobius_onnx-0.1.0.tar.gz

Download URL mobius_onnx-0.1.0.tar.gz
Size 1.1 MB
Tags Source
SHA-256 checksum
How to use checksums
11344ea0c69aa5a2cc3e008b3bbc5703a194cdc37d383a3df51fe69f802412de
BLAKE2b-256 checksum
How to use checksums
cee368b75c84acc4cd69c78946b27fc6fd921014f18a8fc30c1e87bb586ce5ec
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release files / mobius_onnx-0.1.0-py3-none-any.whl

Download URL mobius_onnx-0.1.0-py3-none-any.whl
Size 1.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
9ee9d3d66737abe78f9f9313c25ea08da6c3bf388a070bd8217648a368ace055
BLAKE2b-256 checksum
How to use checksums
11b62eee2093fc2285ab0709a8d487f2fd28745368a7c99ae1c4756a22115c22
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page