⚡ OMNIMUON
Production-Grade Spectral & Second-Order Optimizer Suite for Deep Learning & LLMs
OmniMuon delivers high-velocity, mathematically principled optimization for modern deep neural networks. By moving beyond coordinate-wise diagonal heuristics (AdamW), OmniMuon provides drop-in optimizers that cut training steps by up to 64% while preserving 100% compute efficiency (zero extra backward passes).
Installation • Quickstart • Optimizer Architecture Guide • Empirical Benchmarks • Production Runbooks • Citation
📌 Executive Summary
Modern large-scale models are bottlenecked by standard first-order adaptive gradient descent ($g / \sqrt{v}$). OmniMuon provides a unified toolkit implementing both Spectral Matrix Momentum and Finite-Difference Second-Order Curvature Probing:
Muon: Specialized for Autoregressive LLMs & Transformers. Replaces coordinate-wise scaling with Quintic Newton-Schulz matrix sign orthogonalization on 2D weight matrices, paired with decoupled AdamW on 1D vectors and embedding tables.OmniMuon: The Universal Multi-Modal formulation. Adds Canonical Tensor Matricization (unfolding 3D/4D/5D convolution tensors) and Dynamic RMS Energy Matching for Vision Transformers (ViT) and ConvNets.SophiaFD: Second-Order Stochastic optimization utilizing 2-pass Hutchinson diagonal Hessian probing with scale-aware numerical stability.Polaris: First-principles Riemannian Soft-Polar geodesic optimizer with parameter-free self-consistent regularization $\tau = \text{RMS}(M)$, running on a single unified learning rate.
🚀 Installation
Install the production package directly from PyPI:
pip install omnimuon
Or install the bleeding-edge source from GitHub:
pip install git+https://github.com/AirBorneAI/airborne-muon.git
System Requirements
- Python $\ge 3.9$
- PyTorch $\ge 2.0.0$
- Recommended: CUDA $\ge 11.8$ with native
torch.bfloat16ortorch.float16support.
⚡ Quickstart
1. Training Large Language Models with Muon (Drop-in for AdamW)
import torch
import torch.nn as nn
from omnimuon import Muon
model = nn.TransformerEncoder(
nn.TransformerEncoderLayer(d_model=768, nhead=12, batch_first=True),
num_layers=12
).cuda()
# Zero boilerplate! Defaults automatically use matrix_lr=0.02, lr=1e-3, momentum=0.95, weight_decay=0.01:
optimizer = Muon(model.parameters())
# Or customize any parameter whenever needed:
# optimizer = Muon(model.parameters(), lr=1e-3, matrix_lr=0.02, momentum=0.95, weight_decay=0.01)
# Standard training loop (no closures, zero extra backward passes!)
for tokens, targets in dataloader:
tokens, targets = tokens.cuda(), targets.cuda()
optimizer.zero_grad(set_to_none=True)
with torch.autocast('cuda', dtype=torch.bfloat16):
loss = criterion(model(tokens), targets)
loss.backward()
optimizer.step()
2. Multi-Modal Vision & ConvNet Training with OmniMuon
from omnimuon import OmniMuon
# Works seamlessly across Conv2D, Vision Transformers, and Multi-Modal Models
optimizer = OmniMuon(
model.parameters(),
lr=3e-4, # Vector learning rate
matrix_lr=0.02, # Spectral tensor learning rate
rms_scaling=True, # Scales updates by parameter RMS (prevents patch-embed overshoot)
spectral_blend=0.8, # 80% Orthogonal Matrix Sign + 20% AdamW residual
weight_decay=0.01
)
3. Second-Order Curvature Probing with SophiaFD
from omnimuon import SophiaFD
optimizer = SophiaFD(
model.parameters(),
lr=3e-3,
betas=(0.96, 0.99),
gamma=5.0, # Curvature clipping threshold
rho=1.0, # Maximum parameter step bound
delta=1e-3, # Finite-difference perturbation magnitude
k=10 # Re-compute diagonal Hessian every 10 steps
)
for step, (inputs, targets) in enumerate(dataloader):
optimizer.zero_grad(set_to_none=True)
outputs = model(inputs)
loss = criterion(outputs, targets)
loss.backward()
# Sophia-FD updates diagonal Hessian via 2-pass Hutchinson probing every k steps
if step % optimizer.param_groups[0]['k'] == 0:
base_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]
probes = optimizer.sample_probe_vectors()
delta = optimizer.param_groups[0]['delta']
with torch.no_grad():
for p, u in zip(model.parameters(), probes):
if u is not None:
p.add_(u, alpha=delta)
model.zero_grad(set_to_none=True)
criterion(model(inputs), targets).backward()
pert_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]
with torch.no_grad():
for p, u in zip(model.parameters(), probes):
if u is not None:
p.sub_(u, alpha=delta)
for p, bg in zip(model.parameters(), base_grads):
if bg is not None:
p.grad = bg
optimizer.update_hessian(pert_grads, probes)
optimizer.step()
🏛️ Optimizer Architecture Guide
| Optimizer | Primary Domain | Core Mathematical Mechanism | Compute Overhead | Best For |
|---|---|---|---|---|
Muon |
LLMs / Causal Transformers | Quintic Newton-Schulz Spectral Orthogonalization ($O = M(M^TM)^{-1/2}$) + AdamW hybrid routing | 0% (Exact match to AdamW) | Pretraining GPT, Llama, Mistral, BERT architectures. |
OmniMuon |
Universal (Vision, Conv, Diffusion) | Tensor Matricization ($C_{\text{out}} \times C_{\text{in}} \cdot K_h \cdot K_w$) + Parameter RMS Energy Matching | 0% | ViT, ResNet, ConvNeXt, DiT, Multi-modal models. |
SophiaFD |
Non-Convex Surface Navigation | Finite-Difference Hutchinson Diagonal Hessian Preconditioning ($m / \max(\gamma h, \epsilon)$) | ~10% (1 extra pass every $k=10$ steps) | Highly ill-conditioned non-convex landscapes. |
Polaris |
First-Principles Research | Continuous Soft Polar Geodesic ($\Phi_\tau(M) = M(M^TM + \tau^2 I)^{-1/2}$) with $\tau = \text{RMS}(M)$ | 0% | Single-learning-rate parameter-free optimization. |
📊 Empirical Benchmarks
1. The 6-Way LLM Pre-Training Tournament (MiniGPT on TinyShakespeare)
All optimizers evaluated under identical deterministic parameter initializations, token streams, and cosine decay learning rate schedules:
| Rank | Optimizer | Class | Final Val Loss | Val Perplexity (PPL) | Steps to Match AdamW Final Score | Relative Efficiency |
|---|---|---|---|---|---|---|
| 🥇 | Airborne Muon | Spectral Matrix Orthogonal | 2.0141 |
7.49 PPL |
250 steps | 2.0x Faster (50% Fewer Steps) |
| 🥈 | Sophia-FD | Second-Order Hessian Probe | 2.0794 |
8.00 PPL |
300 steps | 1.6x Faster |
| 🥉 | AdamW | Industry Standard Baseline | 2.1864 |
8.90 PPL |
500 steps (Baseline) | 1.0x Baseline |
| 4 | Adan | Adaptive Nesterov Momentum | 2.2643 |
9.62 PPL |
Did not match | - |
| 5 | diffGrad | Friction Gradient Difference | 2.3228 |
10.20 PPL |
Did not match | - |
| 6 | Polaris | Single-LR Soft Polar | 2.7632 |
15.85 PPL |
Under-converged | Requires warm-up tuning |
2. Full-Scale 124M GPT-2 on NVIDIA A100-80GB (FineWeb-Edu)
Tested under native torch.bfloat16 FlashAttention on real streaming educational web text:
| Optimizer | Final Val Loss (500 steps) | Final Val Perplexity | Steps to Match AdamW Final Loss | Total Step Reduction |
|---|---|---|---|---|
| AdamW Baseline | 2.8104 |
16.62 PPL | 500 steps | Baseline |
| Airborne Muon | 2.2677 |
9.66 PPL |
$\mathbf{\le 180}$ steps | 64% Compute Reduction |
🛠️ Production Runbooks & Best Practices
Recommended Learning Rates & Schedules
-
Cosine Decay with Warmup: Always use a short linear warmup ($2-5%$ of total training steps) followed by half-period cosine decay down to $10%$ of peak learning rate.
Model Family Muon(matrix_lr)Muon(lr- AdamW)Weight Decay Small LLMs (< 500M params) 0.021e-30.01Medium LLMs (1B - 7B params) 0.0156e-40.01Large LLMs (7B - 70B params) 0.013e-40.05Vision Transformers (ViT) 0.01(OmniMuon)3e-40.05 -
Distributed Training (DDP / FSDP / DeepSpeed):
MuonandOmniMuonoperate entirely locally on the parameter gradients computed during backpropagation. They require no global cross-GPU communication during the Newton-Schulz polynomial step. You can use standardDistributedDataParallelor PyTorch FSDP without modifying your distributed training harness. -
Mixed Precision: Newton-Schulz iterations automatically run in native
torch.bfloat16on Ampere, Hopper, and Blackwell architectures for maximum throughput.
🔬 Scientific Foundations & Prior Art
- Matrix Orthogonalization: Utilizes the degree-5 quintic Newton-Schulz iteration with optimal convergence coefficients $(a=3.4445, b=-4.7750, c=2.0315)$ discovered by Keller Jordan (2024).
- Curvature Optimization: Builds upon finite-difference Hutchinson stochastic Hessian probing formulated by Liu et al. (Stanford, 2023).
- The $1/\eta$ Minibatch Secant Collapse: Proved by AirBorne AI Research, demonstrating why consecutive-batch secant estimations collapse into learning rate artifacts.
📜 Citation
If you use OmniMuon in your academic research or production infrastructure, please cite:
@software{omnimuon2026,
author = {Singh, Suryaansh Prithvijit and AirBorne AI Research},
title = {OmniMuon: Universal Spectral Matrix and Second-Order Optimizer Suite for Deep Learning},
year = {2026},
publisher = {PyPI and GitHub},
url = {https://github.com/AirBorneAI/airborne-muon}
}
Released under the Apache 2.0 Open Source License.
Metadata
Release files for omnimuon 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| omnimuon-0.1.2.tar.gz | 17.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| omnimuon-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.8 kB
Release files / omnimuon-0.1.2.tar.gz
| Download URL | omnimuon-0.1.2.tar.gz |
|---|---|
| Size | 17.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f0aa4c1bb758b6a97fc466507e96e4f424c3a0235c1470ffc311f76d7d3a66df
|
|
BLAKE2b-256 checksum How to use checksums |
36094bfcab197693d8fbc11855084a84ada68acdfe9f81b882573d2576de7e88
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / omnimuon-0.1.2-py3-none-any.whl
| Download URL | omnimuon-0.1.2-py3-none-any.whl |
|---|---|
| Size | 16.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
718ca30305d64807addbc4711d76631e56b502235b840cdcaf668da98aaedc46
|
|
BLAKE2b-256 checksum How to use checksums |
65df97d2ccc858ef3c39b629e791609b8a4807af11553c0ec6c42968d1838d14
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|