Skip to main content

⚡ OMNIMUON

Production-Grade Spectral & Second-Order Optimizer Suite for Deep Learning & LLMs

PyPI version Python Versions License PyTorch Downloads

OmniMuon delivers high-velocity, mathematically principled optimization for modern deep neural networks. By moving beyond coordinate-wise diagonal heuristics (AdamW), OmniMuon provides drop-in optimizers that cut training steps by up to 64% while preserving 100% compute efficiency (zero extra backward passes).

Installation • Quickstart • Optimizer Architecture Guide • Empirical Benchmarks • Production Runbooks • Citation


📌 Executive Summary

Modern large-scale models are bottlenecked by standard first-order adaptive gradient descent ($g / \sqrt{v}$). OmniMuon provides a unified toolkit implementing both Spectral Matrix Momentum and Finite-Difference Second-Order Curvature Probing:

  • Muon: Specialized for Autoregressive LLMs & Transformers. Replaces coordinate-wise scaling with Quintic Newton-Schulz matrix sign orthogonalization on 2D weight matrices, paired with decoupled AdamW on 1D vectors and embedding tables.
  • OmniMuon: The Universal Multi-Modal formulation. Adds Canonical Tensor Matricization (unfolding 3D/4D/5D convolution tensors) and Dynamic RMS Energy Matching for Vision Transformers (ViT) and ConvNets.
  • SophiaFD: Second-Order Stochastic optimization utilizing 2-pass Hutchinson diagonal Hessian probing with scale-aware numerical stability.
  • Polaris: First-principles Riemannian Soft-Polar geodesic optimizer with parameter-free self-consistent regularization $\tau = \text{RMS}(M)$, running on a single unified learning rate.

🚀 Installation

Install the production package directly from PyPI:

pip install omnimuon

Or install the bleeding-edge source from GitHub:

pip install git+https://github.com/AirBorneAI/airborne-muon.git

System Requirements

  • Python $\ge 3.9$
  • PyTorch $\ge 2.0.0$
  • Recommended: CUDA $\ge 11.8$ with native torch.bfloat16 or torch.float16 support.

⚡ Quickstart

1. Training Large Language Models with Muon (Drop-in for AdamW)

import torch
import torch.nn as nn
from omnimuon import Muon

model = nn.TransformerEncoder(
    nn.TransformerEncoderLayer(d_model=768, nhead=12, batch_first=True),
    num_layers=12
).cuda()

# Zero boilerplate! Defaults automatically use matrix_lr=0.02, lr=1e-3, momentum=0.95, weight_decay=0.01:
optimizer = Muon(model.parameters())

# Or customize any parameter whenever needed:
# optimizer = Muon(model.parameters(), lr=1e-3, matrix_lr=0.02, momentum=0.95, weight_decay=0.01)

# Standard training loop (no closures, zero extra backward passes!)
for tokens, targets in dataloader:
    tokens, targets = tokens.cuda(), targets.cuda()
    optimizer.zero_grad(set_to_none=True)
    
    with torch.autocast('cuda', dtype=torch.bfloat16):
        loss = criterion(model(tokens), targets)
        
    loss.backward()
    optimizer.step()

2. Multi-Modal Vision & ConvNet Training with OmniMuon

from omnimuon import OmniMuon

# Works seamlessly across Conv2D, Vision Transformers, and Multi-Modal Models
optimizer = OmniMuon(
    model.parameters(),
    lr=3e-4,              # Vector learning rate
    matrix_lr=0.02,       # Spectral tensor learning rate
    rms_scaling=True,     # Scales updates by parameter RMS (prevents patch-embed overshoot)
    spectral_blend=0.8,   # 80% Orthogonal Matrix Sign + 20% AdamW residual
    weight_decay=0.01
)

3. Second-Order Curvature Probing with SophiaFD

from omnimuon import SophiaFD

optimizer = SophiaFD(
    model.parameters(),
    lr=3e-3,
    betas=(0.96, 0.99),
    gamma=5.0,            # Curvature clipping threshold
    rho=1.0,              # Maximum parameter step bound
    delta=1e-3,           # Finite-difference perturbation magnitude
    k=10                  # Re-compute diagonal Hessian every 10 steps
)

for step, (inputs, targets) in enumerate(dataloader):
    optimizer.zero_grad(set_to_none=True)
    outputs = model(inputs)
    loss = criterion(outputs, targets)
    loss.backward()
    
    # Sophia-FD updates diagonal Hessian via 2-pass Hutchinson probing every k steps
    if step % optimizer.param_groups[0]['k'] == 0:
        base_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]
        probes = optimizer.sample_probe_vectors()
        delta = optimizer.param_groups[0]['delta']

        with torch.no_grad():
            for p, u in zip(model.parameters(), probes):
                if u is not None:
                    p.add_(u, alpha=delta)

        model.zero_grad(set_to_none=True)
        criterion(model(inputs), targets).backward()
        pert_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]

        with torch.no_grad():
            for p, u in zip(model.parameters(), probes):
                if u is not None:
                    p.sub_(u, alpha=delta)

        for p, bg in zip(model.parameters(), base_grads):
            if bg is not None:
                p.grad = bg

        optimizer.update_hessian(pert_grads, probes)
        
    optimizer.step()

🏛️ Optimizer Architecture Guide

Optimizer Primary Domain Core Mathematical Mechanism Compute Overhead Best For
Muon LLMs / Causal Transformers Quintic Newton-Schulz Spectral Orthogonalization ($O = M(M^TM)^{-1/2}$) + AdamW hybrid routing 0% (Exact match to AdamW) Pretraining GPT, Llama, Mistral, BERT architectures.
OmniMuon Universal (Vision, Conv, Diffusion) Tensor Matricization ($C_{\text{out}} \times C_{\text{in}} \cdot K_h \cdot K_w$) + Parameter RMS Energy Matching 0% ViT, ResNet, ConvNeXt, DiT, Multi-modal models.
SophiaFD Non-Convex Surface Navigation Finite-Difference Hutchinson Diagonal Hessian Preconditioning ($m / \max(\gamma h, \epsilon)$) ~10% (1 extra pass every $k=10$ steps) Highly ill-conditioned non-convex landscapes.
Polaris First-Principles Research Continuous Soft Polar Geodesic ($\Phi_\tau(M) = M(M^TM + \tau^2 I)^{-1/2}$) with $\tau = \text{RMS}(M)$ 0% Single-learning-rate parameter-free optimization.

📊 Empirical Benchmarks

1. The 6-Way LLM Pre-Training Tournament (MiniGPT on TinyShakespeare)

All optimizers evaluated under identical deterministic parameter initializations, token streams, and cosine decay learning rate schedules:

Rank Optimizer Class Final Val Loss Val Perplexity (PPL) Steps to Match AdamW Final Score Relative Efficiency
🥇 Airborne Muon Spectral Matrix Orthogonal 2.0141 7.49 PPL 250 steps 2.0x Faster (50% Fewer Steps)
🥈 Sophia-FD Second-Order Hessian Probe 2.0794 8.00 PPL 300 steps 1.6x Faster
🥉 AdamW Industry Standard Baseline 2.1864 8.90 PPL 500 steps (Baseline) 1.0x Baseline
4 Adan Adaptive Nesterov Momentum 2.2643 9.62 PPL Did not match -
5 diffGrad Friction Gradient Difference 2.3228 10.20 PPL Did not match -
6 Polaris Single-LR Soft Polar 2.7632 15.85 PPL Under-converged Requires warm-up tuning
LLM Optimizer Tournament

2. Full-Scale 124M GPT-2 on NVIDIA A100-80GB (FineWeb-Edu)

Tested under native torch.bfloat16 FlashAttention on real streaming educational web text:

Optimizer Final Val Loss (500 steps) Final Val Perplexity Steps to Match AdamW Final Loss Total Step Reduction
AdamW Baseline 2.8104 16.62 PPL 500 steps Baseline
Airborne Muon 2.2677 9.66 PPL $\mathbf{\le 180}$ steps 64% Compute Reduction
A100 Benchmark Curves

🛠️ Production Runbooks & Best Practices

  1. Cosine Decay with Warmup: Always use a short linear warmup ($2-5%$ of total training steps) followed by half-period cosine decay down to $10%$ of peak learning rate.

    Model Family Muon (matrix_lr) Muon (lr - AdamW) Weight Decay
    Small LLMs (< 500M params) 0.02 1e-3 0.01
    Medium LLMs (1B - 7B params) 0.015 6e-4 0.01
    Large LLMs (7B - 70B params) 0.01 3e-4 0.05
    Vision Transformers (ViT) 0.01 (OmniMuon) 3e-4 0.05
  2. Distributed Training (DDP / FSDP / DeepSpeed): Muon and OmniMuon operate entirely locally on the parameter gradients computed during backpropagation. They require no global cross-GPU communication during the Newton-Schulz polynomial step. You can use standard DistributedDataParallel or PyTorch FSDP without modifying your distributed training harness.

  3. Mixed Precision: Newton-Schulz iterations automatically run in native torch.bfloat16 on Ampere, Hopper, and Blackwell architectures for maximum throughput.


🔬 Scientific Foundations & Prior Art

  • Matrix Orthogonalization: Utilizes the degree-5 quintic Newton-Schulz iteration with optimal convergence coefficients $(a=3.4445, b=-4.7750, c=2.0315)$ discovered by Keller Jordan (2024).
  • Curvature Optimization: Builds upon finite-difference Hutchinson stochastic Hessian probing formulated by Liu et al. (Stanford, 2023).
  • The $1/\eta$ Minibatch Secant Collapse: Proved by AirBorne AI Research, demonstrating why consecutive-batch secant estimations collapse into learning rate artifacts.

📜 Citation

If you use OmniMuon in your academic research or production infrastructure, please cite:

@software{omnimuon2026,
  author = {Singh, Suryaansh Prithvijit and AirBorne AI Research},
  title = {OmniMuon: Universal Spectral Matrix and Second-Order Optimizer Suite for Deep Learning},
  year = {2026},
  publisher = {PyPI and GitHub},
  url = {https://github.com/AirBorneAI/airborne-muon}
}

Built with high rigor by AirBorne AI Research.
Released under the Apache 2.0 Open Source License.

Metadata

Release files for omnimuon 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for omnimuon 0.1.2
File Size Uploaded
omnimuon-0.1.2.tar.gz 17.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for omnimuon 0.1.2
File Interpreter ABI Platform
omnimuon-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 33.8 kB

Release files / omnimuon-0.1.2.tar.gz

Download URL omnimuon-0.1.2.tar.gz
Size 17.2 kB
Tags Source
SHA-256 checksum
How to use checksums
f0aa4c1bb758b6a97fc466507e96e4f424c3a0235c1470ffc311f76d7d3a66df
BLAKE2b-256 checksum
How to use checksums
36094bfcab197693d8fbc11855084a84ada68acdfe9f81b882573d2576de7e88
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / omnimuon-0.1.2-py3-none-any.whl

Download URL omnimuon-0.1.2-py3-none-any.whl
Size 16.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
718ca30305d64807addbc4711d76631e56b502235b840cdcaf668da98aaedc46
BLAKE2b-256 checksum
How to use checksums
65df97d2ccc858ef3c39b629e791609b8a4807af11553c0ec6c42968d1838d14
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page