Skip to content
AI.info

Research

The Affine Divergence: Aligning Activation Updates Beyond Normalisation

The Affine Divergence: Aligning Activation Updates Beyond Normalisation Overview Research area: Deep learning optimisation theory and architecture design, specifically the mathematical relationship be

The Affine Divergence: Aligning Activation Updates Beyond Normalisation
arXiv
2512.22247
Published
2025-12-24
Authors
George Bird

AI summary

The Affine Divergence: Aligning Activation Updates Beyond Normalisation

Overview

Research area: Deep learning optimisation theory and architecture design, specifically the mathematical relationship between gradient descent updates on parameters versus activations, and its connection to normalisation layers.

Technical level: Advanced. The paper is built on tensor-calculus derivations, steepest-descent arguments and manifold/geometry language, although the underlying intuition is explainable in plain terms.

Scope: A single-author theoretical paper (George Bird, University of Manchester; arXiv:2512.22247v2 [cs.LG], 10 Mar 2026) that derives a systematic mismatch between the "ideal" and "effective" update of activations in affine layers, shows that correcting this mismatch incidentally produces normalisation-like functions, and validates two derived correction families on fully connected CIFAR10 networks.

What This Paper Is About

Backpropagation computes the steepest-descent direction for every intermediate quantity in a network, including the activations, but only the parameter gradients are actually used to update the model. The paper argues this creates a structural mismatch: the parameters take their mathematically ideal step, but the activations that are propagated into them do not, because propagating a parameter change forward multiplies it by a sample-dependent factor. The goal is to derive simple corrections that make both parameters and activations take the optimal step, and to test what those corrections look like and how well they work.

Key Contributions

  1. Derivation of the "affine divergence". For an affine layer z = Wx + b (with the activation function applied later), the paper shows the effective finite-difference gradient of the representation is ΔL/Δz_i = (∂L/∂z_i)(||x||² + 1), which is not equal to the ideal gradient ∂L/∂z_i. This amounts to a sample-wise quadratic bias, equivalent to an effective learning rate of η(||x||² + 1).

  2. Normalisation derived from first principles. Solving for equality between ideal and effective updates yields two structural corrections. One, z = W(x/||x||) + b, is effectively a classical parameterless L₂-normalisation, similar to parameterless RMSNorm without the √n width factor — arriving at normalisation as a consequence of an alignment requirement rather than as a design assumption.

  3. A second, non-normalising solution. The other correction, z = (Wx + b)/√(||x||² + 1), remedies the same divergence but is explicitly not a normaliser and does not have scale invariance. The paper argues its empirical success is a discriminative test of the divergence theory, since conventional explanations (statistical or scale-invariance-based) do not apply to it.

  4. Extension to convolution and to architecture theory. Appendix C derives a related patchwise divergence for convolution and introduces "PatchNorm", described as a compositionally inseparable normaliser built into the convolution. Appendix B argues for dissolving the distinction between normalisers and activation functions via an algebraic decomposition.

Main Findings

  • The mismatch is structural and simple. The divergence term is (||x||² + 1). The paper notes that prior normalisers only partially address it: BatchNorm typically reduces but does not identically cancel Var(||x||²[+1]), while LayerNorm and RMSNorm cancel the term sample-wise but introduce caveats such as rebalancing weight/bias responsibility, altering the effective learning rate, and losing geometric degrees of freedom.

  • Two correction families with opposite character. The norm-like solution (Eqn. 18) projects the representation onto the unit hypersphere (ℝⁿ → Sⁿ⁻¹ ↪ ℝⁿ), irreversibly removing a radial degree of freedom, and features a genuine singularity as ||x|| → 0; the ε-positivity trick mitigates this but quietly reintroduces divergence with ε-distortion. The affine-like solution (Eqn. 19) instead acts as a soft bound on representations and preserves all degrees of freedom, and is non-singular at ||x|| → 0 without an ε term.

  • Backward-pass difference. For the norm-like map, ∂L/∂W = g x̂ᵀ (scale-invariant). For the affine-like map, ∂L/∂W = √(||x||²/(||x||² + 1)) · g x̂ᵀ, which smoothly limits to the same expression as ||x|| → ∞ and becomes closely comparable when 1 << ||x||. The paper argues this indicates a theoretical preference for the affine-like form.

  • The norm-like correction roughly doubles the effective learning rate. Its propagated update reduces to z_i - 2η g_i = z_i - η' g_i, so the paper halves the learning rate in experiments (using both η and η/2) for comparability.

  • Tanh results. Structural corrections — affine-like, or norm-like at η and η/2 — consistently and significantly outperform all other non-divergence-corrective normalisers by a wide margin, more so for deeper networks. Affine-like outperforms all other normalisers except in the 3-layer, 16-width case, where the norm-like structural correction marginally outperforms. RMSNorm underperforms, which the paper attributes to its parameterless form.

  • Leaky-ReLU results. Early learning is faster across all normalisers, followed by a performance drop for BatchNorm and LayerNorm only before stagnation. Affine-like typically performs optimally, particularly at larger widths and depths. For very narrow networks (n=16), no normaliser performs well, which the paper links to the larger fractional loss of representation degrees of freedom at small width. The clearest separation is at a single hidden layer with n=128.

  • Width trends. Performance is plotted against log-layer-width from 1 through 256. For Tanh, affine-like and norm-like structural solutions show significantly higher performance that grows with width, and affine-like performs notably better even at a width of 1 neuron. Both no-normaliser and RMSNorm peak at n = 52 and then decline. For Leaky-ReLU differences are slighter, with affine-like better at larger widths and no normaliser outperforming at shallower widths.

  • Scale of the experiments. Each result reports the mean over 5 repeats with standard error, totalling 560 networks and 112 distinct model permutations.

  • A validated auxiliary prediction. Appendix C.1 predicts a counterintuitive negative correlation between batch size and performance, because cross-sample interference produces further ideal-effective misalignment. The paper states this prediction is empirically validated, which it treats as independent support for the mechanism.

  • ADAM robustness. The paper reports that results using ADAM show the proposed considerations still yield improvement, while noting that the implications of statistical rescalings for representation deflection remain unclear.

  • Relation to natural gradient descent. The paper acknowledges conceptual alignment with natural gradient methods (Amari, 1998; Amari et al., 2019; Martens, 2020) and the contravariant gradient ∇L̃ = G⁻¹∇L with the Fisher information metric G, but states the approaches differ substantially in variable prioritisation, solution method and computational tractability (discussed in Appendix D).

Methodology in Plain English

The argument proceeds in three stages.

First, take the simplest possible layer, a linear map plus a bias, and ask: if I apply the standard gradient descent update to the weights and bias, what change actually arrives at the layer's output? Doing the algebra shows the output changes by -η g_i (||x||² + 1) rather than by the ideal -η g_i. The extra factor depends on the input sample, so samples with larger magnitude get disproportionately large effective steps.

Second, ask what would need to change about the layer's functional form for the two to match. The paper derives candidate corrections by inserting scaling factors into the affine map and solving for the scale that makes the effective update equal the ideal one. This produces the two structural solutions, plus a separate "gradient-only" class that edits the gradient rather than the layer's forward form.

Third, test empirically. The corrections are compared against BatchNorm, LayerNorm and RMSNorm on fully connected CIFAR10 networks, varying activation function (Tanh and Leaky-ReLU), width and depth. The paper explains that fully connected models are used because the affine correction applies to affine layers, and because single-layer approximations are more likely to remain valid when there are not many prior layers compounding propagated corrections. The content provided is truncated, so the full experimental protocol in Appendix E and the PatchNorm empirical results are not visible here.

Several simplifying assumptions are stated explicitly and flagged for future relaxation: single-step corrections that are first-order in the learning rate, and a single-layer approximation in which propagation starts at each affine layer's parameters and terminates at that same layer's output activations, thereby avoiding the activation function. A single-sample assumption is made in the main text and relaxed in Appendix C.1.

Why This Matters

Impact on research. The paper reframes normalisation as an approximate correction of a mathematical misalignment rather than primarily as a statistical or conditioning device. If the argument holds, it gives a priori derivational support for normalisers (alongside existing post-hoc explanations), predicts that a non-normalising, non-scale-invariant alternative should work — and reports that it does — and yields new functional forms such as PatchNorm. It also argues that normalisers are better analysed as geometric operators decomposed into activation-function-like maps with parameterised scaling, rather than as statistical operations. A concrete example given is that LayerNorm's mean subtraction can be reweighted to any weighted mean, including a one-hot weighting, which the paper says has no geometric consequence beyond reorienting the hypersphere.

Potential application areas. The paper itself reports no applied or real-world deployment experiments; its empirical work is fully connected CIFAR10 classification. The following are areas where the findings could plausibly matter, not applications demonstrated in the paper:

  • Layer design for networks that currently rely on normalisation, particularly fully connected and convolutional stacks.
  • Convolutional architectures, via the derived PatchNorm form, which the paper describes as inseparable from the convolution operation (the truncated content does not report its quantitative results here).
  • Optimisation pipelines where normalisation interacts with optimiser choice, an area the paper explicitly flags as unresolved.
  • Settings where batch size is constrained, given the paper's prediction and reported observation of a negative correlation between batch size and performance.

Industry relevance. Normalisation layers are standard components in production deep learning systems, and the paper's central claim is that a widely deployed design pattern may be an approximate fix for an underlying misalignment that has a better, or at least complementary, alternative. The reported result that the affine-like correction tends to match or exceed conventional normalisers on the tested setting, and that its success cannot be explained by scale invariance because it does not have it, is the kind of claim that, if it generalised, would affect architecture choices. The paper cautions against over-reading this: it states the affine-like solution should not be taken as widely practical until it is mechanistically generalised.

Future Directions

  • Relax the approximations. The paper explicitly encourages relaxing the first-order, single-step and single-layer approximations, noting that propagating through multiple nonlinear layers becomes analytically non-trivial and potentially intractable.
  • Generalise to other architectures. Divergences for residual and attention layers are outlined, and the paper speculates that attention's lack of normalisation may offer weak, circumstantial evidence for the theory. PatchNorm's scope is derived and tested but the truncated content does not report how widely it applies.
  • Investigate the optimiser interaction. The paper suggests future work on the relation to optimiser choice and on the ramifications of rescaling by running statistics on representation updates, including ADAM's element-wise rescaling, noting that cumulative representation deflection remains unclear.
  • Test the auxiliary prediction further. The predicted negative correlation between batch size and performance provides an independent, falsifiable hypothesis; expanding such tests would probe the mechanism beyond the reported CIFAR10 setting.

Target Audience

Researchers working on optimisation theory, normalisation layers and deep learning architecture design will get the most from this paper, as will readers interested in geometric rather than statistical interpretations of network components. It is written for a mathematically comfortable audience: the derivations use tensor notation and manifold language, and the appendices carry a substantial share of the argument — including the convolution extension, the activation/normaliser decomposition, the comparison with natural gradient methods, and the supplementary experiments. The truncated content does not include the full experimental details, so readers wanting to reproduce or scrutinise the numbers should go to Appendix E of the original.

Authors’ abstract

A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, as they are closer to the loss in the computational graph and carry sample-dependent information through the network. Yet their propagated updates do not take the optimal steepest-descent step. These quantities exhibit non-ideal sample-wise scaling across affine, convolutional, and attention layers.Solutions to correct for this are trivial and, incidentally, derive normalisation from first principles despite motivational independence. Consequently, such considerations offer a fresh, conceptual reframe of normalisation's action, with auxiliary experiments bolstering this mechanistic interpretation. Moreover, this analysis makes clear a second possibility: a solution that is functionally distinct from modern normalisations, without scale invariance, yet remains empirically successful -- an alternative to the affine map. This outperforms conventional normalisers across several tests. This generalises to convolution via a new functional form, ``PatchNorm'', a compositionally inseparable normaliser. Together, these provide an alternative mechanistic framework that both adds to and counters some of the discussion of normalisation. Further, it is argued that normalisers are better decomposed into activation-function-like maps with parameterised scaling. Overall, this constitutes a theoretically principled approach that yields new functions with empirical validation and raises questions about the affine + nonlinear approach.

Read the original paper