Airborne Muon: Spectral Matrix Orthogonal Optimizer for LLMs
Muon is a production-grade, drop-in replacement for torch.optim.AdamW designed specifically for modern Large Language Models and Transformer architectures.
By replacing Adam's coordinate-wise normalization with Spectral Matrix Sign Orthogonalization via Quintic Newton-Schulz iterations, Muon enables models to converge in half the steps of AdamW while maintaining 100% compute efficiency (zero extra backward passes).
Benchmark Highlights: The 6-Way LLM Optimizer Tournament
Tested head-to-head on autoregressive Transformer pre-training under identical seeds, architecture (MiniGPT), and cosine decay schedules:
| Rank | Optimizer | Optimizer Class | Final Val Loss | Val Perplexity (PPL) | Steps to Beat AdamW |
|---|---|---|---|---|---|
| 🥇 | Airborne Muon | Spectral Matrix Orthogonal | 2.0141 | 7.49 PPL | 250 steps (50% Compute Reduction!) |
| 🥈 | Sophia-FD | Second-Order Hessian Probe | 2.0794 | 8.00 PPL | 300 steps |
| 🥉 | AdamW | Industry Standard Baseline | 2.1864 | 8.90 PPL | Baseline (500 steps) |
| 4 | Adan | Adaptive Nesterov Momentum | 2.2643 | 9.62 PPL | Never beat AdamW |
| 5 | diffGrad | Friction-based Gradient Diff | 2.3228 | 10.20 PPL | Never beat AdamW |
| 6 | Polaris | Single-LR Soft Polar | 2.7632 | 15.85 PPL | Under-converged |
Key Takeaway: Airborne Muon reached 8.42 PPL at Step 250, defeating AdamW's final 500-step score in half the optimization steps. On full 124M GPT-2 on an NVIDIA A100-80GB, Airborne Muon matched AdamW's final loss in $\le 180$ steps (64% step reduction).
Why Muon Outperforms AdamW on LLMs
- Matrix-Aware Geometry: Adam normalizes each parameter scalar independently: $u_i = m_i / \sqrt{v_i}$. For 2D linear weight matrices ($W_q, W_k, W_v, W_{\text{out}}$), this is blind to coordinate rotations and causes severe zigzagging across anisotropic valley walls.
- Spectral Orthogonalization (Newton-Schulz-5): Muon computes the approximate matrix sign of the momentum tensor: $$O = M (M^T M)^{-1/2}$$ This normalizes all singular values to $1.0$, ensuring that all principal feature directions update at a uniform, optimal velocity.
- Automatic Hybrid Routing:
- 2D Weight Matrices $\to$ Spectral Orthogonal Momentum ($O$)
- 1D Vectors (LayerNorm, RMSNorm, Biases) & Token Embeddings $\to$ Decoupled AdamW
- Zero Compute Overhead: Unlike second-order curvature methods (AdaHessian, Sophia) that require extra backward passes or double-backward autograd graphs, Muon runs in zero extra forward/backward passes.
Quickstart
Installation
pip install airborne-muon
# Or from source:
pip install git+https://github.com/AirBorneAI/airborne-muon.git
Usage (Drop-in Replacement for AdamW)
import torch
from muon import Muon
# Initialize your Transformer / LLM
model = MyTransformer().to('cuda')
# Initialize Muon with hybrid routing
optimizer = Muon(
model.parameters(),
lr=1e-3, # Learning rate for 1D vectors / embeddings (AdamW)
matrix_lr=2e-2, # Spectral learning rate for 2D weight matrices (Muon)
momentum=0.95, # Momentum coefficient
weight_decay=0.01 # Decoupled weight decay
)
# Standard training loop (no closures, no extra passes!)
for x, y in dataloader:
optimizer.zero_grad(set_to_none=True)
with torch.autocast('cuda', dtype=torch.bfloat16):
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
Empirical Research Artifacts in this Repository
muon.py: Production-grade Muon optimizer class with integrated AdamW routing.cs_adam.py: Curvature-Steered Adam implementing directional Gram-Schmidt projection.sophia_fd.py: Sophia-FD optimizer with finite-difference Hutchinson probing.benchmark_grand_showdown.py: Deterministic pretraining benchmark script on TinyShakespeare.models/architectures.py: Causal Transformer (MiniGPT) and ConvNet architectures.test_curvature_unit.py: Mathematical proof of the $1/\eta$ minibatch secant collapse.
Citation & Acknowledgments
Built by the research team at AirBorne AI. Inspired by spectral matrix momentum discoveries by Keller Jordan and the Sophia optimization framework by Liu et al. (Stanford).
@software{airborne_muon_2026,
author = {Singh, Suryaansh Prithvijit and AirBorne AI},
title = {Airborne Muon: Production Spectral Matrix Orthogonal Optimizer for LLMs},
year = {2026},
url = {https://github.com/AirBorneAI/airborne-muon}
}
Metadata
Release files for omnimuon 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| omnimuon-0.1.1.tar.gz | 13.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| omnimuon-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 27.5 kB
Release files / omnimuon-0.1.1.tar.gz
| Download URL | omnimuon-0.1.1.tar.gz |
|---|---|
| Size | 13.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b86d69b292c8e155daa741dd59a9d5ced123183ac273d3689d36b1f16ab67bf6
|
|
BLAKE2b-256 checksum How to use checksums |
e461dec0309e97f81dada7dc1ad4de8f66816e79cfaade2ae9c3c8628d5c0a0b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / omnimuon-0.1.1-py3-none-any.whl
| Download URL | omnimuon-0.1.1-py3-none-any.whl |
|---|---|
| Size | 14.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0670f1806f17dcfab9db6d0288ebf4f0d26df6bdd85f7280578e0dd99a5fbef5
|
|
BLAKE2b-256 checksum How to use checksums |
fe5af8a3548268c3e9332389a87c310f659b452ec818c8b6ad5668174141b115
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|