Skip to main content

Airborne Muon: Spectral Matrix Orthogonal Optimizer for LLMs

PyPI version License PyTorch

Muon is a production-grade, drop-in replacement for torch.optim.AdamW designed specifically for modern Large Language Models and Transformer architectures.

By replacing Adam's coordinate-wise normalization with Spectral Matrix Sign Orthogonalization via Quintic Newton-Schulz iterations, Muon enables models to converge in half the steps of AdamW while maintaining 100% compute efficiency (zero extra backward passes).


Benchmark Highlights: The 6-Way LLM Optimizer Tournament

Tested head-to-head on autoregressive Transformer pre-training under identical seeds, architecture (MiniGPT), and cosine decay schedules:

Rank Optimizer Optimizer Class Final Val Loss Val Perplexity (PPL) Steps to Beat AdamW
🥇 Airborne Muon Spectral Matrix Orthogonal 2.0141 7.49 PPL 250 steps (50% Compute Reduction!)
🥈 Sophia-FD Second-Order Hessian Probe 2.0794 8.00 PPL 300 steps
🥉 AdamW Industry Standard Baseline 2.1864 8.90 PPL Baseline (500 steps)
4 Adan Adaptive Nesterov Momentum 2.2643 9.62 PPL Never beat AdamW
5 diffGrad Friction-based Gradient Diff 2.3228 10.20 PPL Never beat AdamW
6 Polaris Single-LR Soft Polar 2.7632 15.85 PPL Under-converged

Benchmark Comparison

Key Takeaway: Airborne Muon reached 8.42 PPL at Step 250, defeating AdamW's final 500-step score in half the optimization steps. On full 124M GPT-2 on an NVIDIA A100-80GB, Airborne Muon matched AdamW's final loss in $\le 180$ steps (64% step reduction).


Why Muon Outperforms AdamW on LLMs

  1. Matrix-Aware Geometry: Adam normalizes each parameter scalar independently: $u_i = m_i / \sqrt{v_i}$. For 2D linear weight matrices ($W_q, W_k, W_v, W_{\text{out}}$), this is blind to coordinate rotations and causes severe zigzagging across anisotropic valley walls.
  2. Spectral Orthogonalization (Newton-Schulz-5): Muon computes the approximate matrix sign of the momentum tensor: $$O = M (M^T M)^{-1/2}$$ This normalizes all singular values to $1.0$, ensuring that all principal feature directions update at a uniform, optimal velocity.
  3. Automatic Hybrid Routing:
    • 2D Weight Matrices $\to$ Spectral Orthogonal Momentum ($O$)
    • 1D Vectors (LayerNorm, RMSNorm, Biases) & Token Embeddings $\to$ Decoupled AdamW
  4. Zero Compute Overhead: Unlike second-order curvature methods (AdaHessian, Sophia) that require extra backward passes or double-backward autograd graphs, Muon runs in zero extra forward/backward passes.

Quickstart

Installation

pip install airborne-muon
# Or from source:
pip install git+https://github.com/AirBorneAI/airborne-muon.git

Usage (Drop-in Replacement for AdamW)

import torch
from muon import Muon

# Initialize your Transformer / LLM
model = MyTransformer().to('cuda')

# Initialize Muon with hybrid routing
optimizer = Muon(
    model.parameters(),
    lr=1e-3,           # Learning rate for 1D vectors / embeddings (AdamW)
    matrix_lr=2e-2,    # Spectral learning rate for 2D weight matrices (Muon)
    momentum=0.95,     # Momentum coefficient
    weight_decay=0.01  # Decoupled weight decay
)

# Standard training loop (no closures, no extra passes!)
for x, y in dataloader:
    optimizer.zero_grad(set_to_none=True)
    with torch.autocast('cuda', dtype=torch.bfloat16):
        loss = criterion(model(x), y)
    loss.backward()
    optimizer.step()

Empirical Research Artifacts in this Repository


Citation & Acknowledgments

Built by the research team at AirBorne AI. Inspired by spectral matrix momentum discoveries by Keller Jordan and the Sophia optimization framework by Liu et al. (Stanford).

@software{airborne_muon_2026,
  author = {Singh, Suryaansh Prithvijit and AirBorne AI},
  title = {Airborne Muon: Production Spectral Matrix Orthogonal Optimizer for LLMs},
  year = {2026},
  url = {https://github.com/AirBorneAI/airborne-muon}
}

Metadata

Release files for omnimuon 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for omnimuon 0.1.0
File Size Uploaded
omnimuon-0.1.0.tar.gz 11.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for omnimuon 0.1.0
File Interpreter ABI Platform
omnimuon-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 23.9 kB

Release files / omnimuon-0.1.0.tar.gz

Download URL omnimuon-0.1.0.tar.gz
Size 11.6 kB
Tags Source
SHA-256 checksum
How to use checksums
03f0066ebd1d1157b713940395e21b1fa9c26c9be148b1002deef88191f97cd6
BLAKE2b-256 checksum
How to use checksums
849f39330aa54b97ffa849a83f1b4d99fa6658d7a81be269e858c54e6ba11a06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / omnimuon-0.1.0-py3-none-any.whl

Download URL omnimuon-0.1.0-py3-none-any.whl
Size 12.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
604175f0177c6aacf440398266c0e9eeaa6f9ea7288451e98b8bc59938976e29
BLAKE2b-256 checksum
How to use checksums
6ce38dd9c8b0c21d9fcb1b18c7111f1e4401b2778304f213594cd822b8685ecb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page