Skip to main content

Airborne Muon: Spectral Matrix Orthogonal Optimizer for LLMs

PyPI version License PyTorch

Muon is a production-grade, drop-in replacement for torch.optim.AdamW designed specifically for modern Large Language Models and Transformer architectures.

By replacing Adam's coordinate-wise normalization with Spectral Matrix Sign Orthogonalization via Quintic Newton-Schulz iterations, Muon enables models to converge in half the steps of AdamW while maintaining 100% compute efficiency (zero extra backward passes).


Benchmark Highlights: The 6-Way LLM Optimizer Tournament

Tested head-to-head on autoregressive Transformer pre-training under identical seeds, architecture (MiniGPT), and cosine decay schedules:

Rank Optimizer Optimizer Class Final Val Loss Val Perplexity (PPL) Steps to Beat AdamW
🥇 Airborne Muon Spectral Matrix Orthogonal 2.0141 7.49 PPL 250 steps (50% Compute Reduction!)
🥈 Sophia-FD Second-Order Hessian Probe 2.0794 8.00 PPL 300 steps
🥉 AdamW Industry Standard Baseline 2.1864 8.90 PPL Baseline (500 steps)
4 Adan Adaptive Nesterov Momentum 2.2643 9.62 PPL Never beat AdamW
5 diffGrad Friction-based Gradient Diff 2.3228 10.20 PPL Never beat AdamW
6 Polaris Single-LR Soft Polar 2.7632 15.85 PPL Under-converged

Benchmark Comparison

Key Takeaway: Airborne Muon reached 8.42 PPL at Step 250, defeating AdamW's final 500-step score in half the optimization steps. On full 124M GPT-2 on an NVIDIA A100-80GB, Airborne Muon matched AdamW's final loss in $\le 180$ steps (64% step reduction).


Why Muon Outperforms AdamW on LLMs

  1. Matrix-Aware Geometry: Adam normalizes each parameter scalar independently: $u_i = m_i / \sqrt{v_i}$. For 2D linear weight matrices ($W_q, W_k, W_v, W_{\text{out}}$), this is blind to coordinate rotations and causes severe zigzagging across anisotropic valley walls.
  2. Spectral Orthogonalization (Newton-Schulz-5): Muon computes the approximate matrix sign of the momentum tensor: $$O = M (M^T M)^{-1/2}$$ This normalizes all singular values to $1.0$, ensuring that all principal feature directions update at a uniform, optimal velocity.
  3. Automatic Hybrid Routing:
    • 2D Weight Matrices $\to$ Spectral Orthogonal Momentum ($O$)
    • 1D Vectors (LayerNorm, RMSNorm, Biases) & Token Embeddings $\to$ Decoupled AdamW
  4. Zero Compute Overhead: Unlike second-order curvature methods (AdaHessian, Sophia) that require extra backward passes or double-backward autograd graphs, Muon runs in zero extra forward/backward passes.

Quickstart

Installation

pip install airborne-muon
# Or from source:
pip install git+https://github.com/AirBorneAI/airborne-muon.git

Usage (Drop-in Replacement for AdamW)

import torch
from muon import Muon

# Initialize your Transformer / LLM
model = MyTransformer().to('cuda')

# Initialize Muon with hybrid routing
optimizer = Muon(
    model.parameters(),
    lr=1e-3,           # Learning rate for 1D vectors / embeddings (AdamW)
    matrix_lr=2e-2,    # Spectral learning rate for 2D weight matrices (Muon)
    momentum=0.95,     # Momentum coefficient
    weight_decay=0.01  # Decoupled weight decay
)

# Standard training loop (no closures, no extra passes!)
for x, y in dataloader:
    optimizer.zero_grad(set_to_none=True)
    with torch.autocast('cuda', dtype=torch.bfloat16):
        loss = criterion(model(x), y)
    loss.backward()
    optimizer.step()

Empirical Research Artifacts in this Repository


Citation & Acknowledgments

Built by the research team at AirBorne AI. Inspired by spectral matrix momentum discoveries by Keller Jordan and the Sophia optimization framework by Liu et al. (Stanford).

@software{airborne_muon_2026,
  author = {Singh, Suryaansh Prithvijit and AirBorne AI},
  title = {Airborne Muon: Production Spectral Matrix Orthogonal Optimizer for LLMs},
  year = {2026},
  url = {https://github.com/AirBorneAI/airborne-muon}
}

Metadata

Release files for omnimuon 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for omnimuon 0.1.1
File Size Uploaded
omnimuon-0.1.1.tar.gz 13.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for omnimuon 0.1.1
File Interpreter ABI Platform
omnimuon-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 27.5 kB

Release files / omnimuon-0.1.1.tar.gz

Download URL omnimuon-0.1.1.tar.gz
Size 13.0 kB
Tags Source
SHA-256 checksum
How to use checksums
b86d69b292c8e155daa741dd59a9d5ced123183ac273d3689d36b1f16ab67bf6
BLAKE2b-256 checksum
How to use checksums
e461dec0309e97f81dada7dc1ad4de8f66816e79cfaade2ae9c3c8628d5c0a0b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / omnimuon-0.1.1-py3-none-any.whl

Download URL omnimuon-0.1.1-py3-none-any.whl
Size 14.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0670f1806f17dcfab9db6d0288ebf4f0d26df6bdd85f7280578e0dd99a5fbef5
BLAKE2b-256 checksum
How to use checksums
fe5af8a3548268c3e9332389a87c310f659b452ec818c8b6ad5668174141b115
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page