Research
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
Overview Research area: Machine learning theory — specifically the implicit bias (implicit regularization) of gradient-based optimizers on homogeneous neural networks. Technical level: Advanced. The p

- arXiv
- 2602.16340
- Published
- 2026-02-18
- Authors
- Eitan Gronich, Gal Vardi
AI summary
Overview
- Research area: Machine learning theory — specifically the implicit bias (implicit regularization) of gradient-based optimizers on homogeneous neural networks.
- Technical level: Advanced. The paper is a theoretical optimization/dynamical-systems analysis built on flow (infinitesimal step size) ODEs, Clarke subdifferentials, KKT conditions, and stratified geometry. A reader needs comfort with convex analysis, norms and dual norms, and convergence analysis.
- Scope in one sentence: The paper proves that Adam, Muon, and a broad family of momentum-based steepest descent algorithms converge in direction to KKT points of a margin-maximization problem on smooth homogeneous models, with the identity of the maximized norm determined by the optimizer.
What This Paper Is About
Deep networks generalize well even when they are heavily overparameterized and trained with no explicit regularization. One leading explanation is that the optimizer itself has an "implicit bias" that steers training toward solutions with large margin. Prior work established this mostly for plain gradient descent (and some for Adam in linear models). This paper asks what happens for the modern momentum-based optimizers — Muon, Adam (analyzed without its stability constant), Signum, and composites like Muon-Signum and Muon-Adam — when they train smooth homogeneous models rather than just linear predictors.
Key Contributions
- A general theory of "approximate steepest descent." The authors define a notion (Definition 5.1) of a trajectory that only approximately follows the steepest descent direction, and show that such trajectories still converge in direction to KKT points of the norm-specific margin-maximization problem. This is the technical core that lets momentum-based optimizers be handled.
- Implicit bias of momentum steepest descent and Muon. Theorem 3.2 covers normalized and unnormalized momentum steepest descent; Corollary 3.3 shows Muon on a collection of weight matrices is a special case with the max-spectral norm, and Corollary 3.4 shows Muon-Signum corresponds to a hybrid max-of-max-spectral-and-ℓ∞ norm.
- Implicit bias of Adam (without the stability constant). Theorem 3.5 shows Adam with c₁ ≥ c₂ converges in direction to a KKT point of the ℓ∞ margin-maximization problem under a non-increasing, decaying learning rate. Theorem 3.6 extends this to Muon-Adam with distinct momentum and base learning-rate parameters.
- Extension beyond smooth models. Section 4 shows the smoothness assumption can be weakened to a Whitney C¹-stratifiable, locally Lipschitz model (M1-Weak) if the normalized model subgradients converge (condition T3), and reports that this condition appears to be violated in their two-layer ReLU experiments on MNIST.
Main Findings
- Optimizer choice determines which margin is maximized. Muon (spectral norm), MomentumGD (ℓ₂ norm), and Signum (ℓ∞ norm) each have a bias toward KKT points of the margin-maximization problem defined by their corresponding norm, given a decaying learning rate schedule.
- Muon's bias is a max-spectral norm bias. Running Muon simultaneously on several weight matrices yields normalized momentum steepest descent with ‖(W₁,…,W_K)‖_msp := max_k ‖W_k‖_sp.
- Composite algorithms inherit a hybrid norm. Muon-Signum corresponds to the norm max{‖(W₁,…,W_K)‖_msp, ‖u‖∞}; Muon-Adam corresponds to max{(η₀ᴬ/η₀ᴹ)‖(W₁,…,W_K)‖_msp, ‖u‖∞}, so the ratio of base learning rates scales the matrix part of the norm.
- Adam maximizes the ℓ∞ margin, but only under additional structure. Theorem 3.5 requires the momentum parameters to satisfy c₁ ≥ c₂ (analogous to β₁ ≤ β₂), a non-increasing learning rate (Assumption LR-Adam), and a technical initialization condition (A1) ensuring each coordinate of the second-moment estimate stays positive.
- Adam is not a normalized momentum steepest descent algorithm. The paper states this explicitly: Adam's update is a ratio of two momentum estimates with different rates, so it needs separate treatment rather than falling out of the momentum steepest descent framework.
- The results hold for a family of losses. The loss is of the form ℒ(θ) = Σ_{i=1}^m e^{−φ(y_i f(x_i;θ))}, with φ twice continuously differentiable, strictly monotone increasing, convex, and with bounded first and second derivatives — this includes the exponential loss ℓ(u) = e^{−u} and the logistic loss ℓ(u) = log(1 + e^{−u}).
- Model class covered. Assumption (M2) requires f to be L-homogeneous for some L ≥ 1. This includes deep linear networks, the smooth nonlinear ReLU^q activation max{0, z}^q for q > 1, and the quadratic activation z ↦ z².
- The learning rate must decay but not too fast. Assumption (LR-MSD) requires ∫₀^∞ η(t)dt = ∞ and η(t) ≤ o(t^{1/L − 1}). The harmonic schedule η(t) = 1/t satisfies the assumptions for any L > 1.
- The proof reduces to gradient–parameter alignment. KKT stationarity of the limiting direction is established through the alignment quantity ⟨θ_t/‖θ_t‖, −g_t/‖g_t‖_⋆⟩ tending to 1 (Section 5, based on insights of Tsilivis et al. (2025) and Ji and Telgarsky (2020)).
- Experiments corroborate the theory. The paper reports that experiments confirm the theory and show the identity of the maximized margin depends on the optimizer, and Figure 1 shows empirical evidence of directional convergence and strictly positive margins. Specific numerical results are not reported in the provided content.
- Non-smooth setting caveat. The paper notes condition (T3) appears to be violated in their experiments on two-layer ReLU networks with the MNIST dataset, leaving open whether some settings comply.
Methodology in Plain English
The authors work in the continuous-time (flow) limit of optimization, where a training run becomes a differential equation rather than a sequence of discrete steps. In this limit, normalization of the update and the learning rate schedule do not change the path taken in parameter space; they only change how fast the path is traversed. This lets the authors separate "which direction do we move?" from "how quickly?"
They then express each optimizer in this language. Plain gradient descent, coordinate descent, and sign gradient descent are all instances of steepest descent with respect to ℓ₂, ℓ₁, and ℓ∞ norms. Muon, when its orthogonalization is treated exactly (the U Σ Vᵀ ↦ U Vᵀ map rather than the Newton–Schulz approximation), is steepest descent with respect to the spectral norm; running it on several matrices at once uses the max-spectral norm. Momentum is modeled as an exponentially-weighted average of past gradients with a smoothing parameter c₁, and Adam additionally keeps a running average of squared gradients with a second parameter c₂.
The key technical move is to define a class of trajectories that only approximately follow the steepest descent direction, and then prove that this approximation is enough. The authors show that if a trajectory's normalized direction converges, its loss decays, its norm grows, and the alignment between the parameter direction and the negative gradient direction approaches 1, then the limiting direction must satisfy the KKT conditions of the norm-specific hard-margin problem. They verify these properties separately for momentum steepest descent (choosing ν(t) = ‖dθ_t/dt‖) and for Adam (choosing ν(t) = η(t), the learning rate). For non-smooth models such as ReLU networks, they show the same conclusion follows if the normalized model subgradients themselves converge — which holds whenever the parameters eventually stay inside a single smooth stratum. Formal proofs are deferred to the appendix.
Why This Matters
Impact on research. Previous implicit-bias results for Adam and Muon were confined to linear predictors. This paper pushes them to the much broader smooth homogeneous class, and — more importantly — supplies a reusable abstraction (approximate steepest descent) that decouples the proof technique from any single optimizer. The finding that the maximized norm is a tunable consequence of optimizer choice reframes "which optimizer" as "which geometry."
Real-world applications (as implied by the paper):
- Training large language models, where Adam and Muon are used near-universally and where generalization behavior without explicit regularization is a central engineering concern.
- Training vision transformers, which the paper names as another major deployment area for these optimizers.
- Large-scale matrix-parameterized models, where Muon's per-weight-matrix orthogonalization and its max-spectral norm bias apply directly.
- Mixed-parameter architectures that use Muon (or Scion-style combinations with sign gradient descent) on weight matrices alongside Adam on non-matrix parameters, since the paper characterizes exactly which hybrid norm such composites maximize.
Industry relevance. Practitioners routinely combine multiple optimizers within a single model and tune momentum and base learning rates independently. Theorem 3.6 states that the ratio η₀ᴬ/η₀ᴹ of base learning rates literally reweights the matrix part of the effective norm, which turns a common hyperparameter choice into a statement about the geometry of the solution found. The paper also notes that Pytorch defaults are β₁ = 0.9 and β₂ = 0.999, consistent with the c₁ ≥ c₂ regime it analyzes.
Future Directions
- Multiclass and general loss settings. The authors explicitly generalize Zhang et al. (2024) "albeit in the binary classification setting," while Fan et al. (2025) handled the multiclass case in linear models — leaving the multiclass extension of this work open.
- When does condition (T3) hold in practice? The paper proves the non-smooth extension under (T3) but states it is "unclear whether this is satisfied in practice or under what conditions," and reports it appears violated in their two-layer ReLU experiments on MNIST. Characterizing which trajectories satisfy it is a direct open question.
- Closing the gap between assumptions and provable behavior. Assumptions (T1) and (T2) — nontrivial trajectory and directional convergence with positive margin — are taken as assumptions rather than derived for momentum steepest descent and Adam, unlike the case for (normalized) steepest descent. Proving decay and directional convergence for these optimizers would strengthen the results.
- Faithfulness of the Adam model. The analysis drops the stability constant and focuses on c₁ ≥ c₂. The paper notes prior works arguing the stability constant makes the analysis less faithful to practice; whether the results survive the c₁ < c₂ regime or other practical configurations remains open. It also notes that Shampoo with accumulation disabled is spectral descent, but that Shampoo's accumulation differs from momentum — another unexplored case.
Target Audience
This paper is written for optimization and learning-theory researchers — graduate students and faculty in machine learning theory, applied mathematics, and dynamical systems — who work on implicit bias, margin maximization, or convergence analysis of optimizers. It is also relevant to theoretically-minded practitioners who want principled guidance on how optimizer choice and learning-rate ratios shape the geometry of the solutions their models find. Readers without a background in normed steepest descent, Clarke subdifferentials, and KKT conditions will need to work through the preliminaries section before the results are accessible.
Authors’ abstract
We study the implicit bias of momentum-based optimizers on homogeneous models. We first extend existing results on the implicit bias of steepest descent in homogeneous models to normalized steepest descent with an optional learning rate schedule. We then show that for smooth homogeneous models, momentum steepest descent algorithms like Muon (spectral norm), MomentumGD ($\ell_2$ norm), and Signum ($\ell_\infty$ norm) are approximate steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms too have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the $\ell_\infty$ margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.