Research
Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
Overview Research area: Statistical physics of learning / theoretical deep learning (stat.ML), combining spin-glass replica theory with random matrix models. Technical level: Advanced. The paper assum

- arXiv
- 2510.24616
- Published
- 2025-10-28
- Authors
- Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
AI summary
Overview
Research area: Statistical physics of learning / theoretical deep learning (stat.ML), combining spin-glass replica theory with random matrix models.
Technical level: Advanced. The paper assumes familiarity with statistical mechanics of disordered systems, the replica method, random matrix theory, and Bayesian learning theory.
One-sentence scope: The paper derives the Bayes-optimal generalisation limits and the sufficient statistics (order parameters) describing what an optimally trained multi-layer perceptron learns, in the matched teacher-student setting, when the network width scales linearly with the input dimension and the number of samples is quadratic in it (the interpolation regime).
What This Paper Is About
Existing quantitative theories of neural networks either analyse very narrow networks, or kernel-like models whose inner layers are effectively frozen, or networks trained far from interpolation, so none captures the combined effect of depth, non-linear activations and finite (proportional) width when the model must genuinely adapt to the task. The authors study supervised learning of a multi-layer perceptron (MLP) where the data are generated by another MLP of the same architecture, in a scaling limit where the width of each hidden layer grows proportionally to the input dimension and the number of training samples grows as the square of the input dimension. The goal is to identify the fundamental limits of learning such targets and the sufficient statistics that describe what an optimally trained network has actually learnt as the data budget grows.
Key Contributions
-
A tractable theory of proportional-width, non-linear, deep networks near interpolation. The authors provide replica-symmetric formulas for the free entropy and order parameters of MLPs satisfying four properties simultaneously: width proportional to input dimension (P1), broad classes of non-linear activations (P2), possibly multiple hidden layers (P3), and learning in the interpolation regime (P4).
-
A unified formalism for "hybrid" inference problems. They combine the replica method with the Harish Chandra–Itzykson–Zuber (HCIZ) spherical integral in a new way, producing a single formalism able to handle both the correlations induced by the matrix nature of the problem and the mean-field symmetry-breaking (specialisation) transitions. The paper argues that prior work required rotational invariance, or treated the two phases separately with different formalisms.
-
Analytical results by depth. Result 1 gives the replica-symmetric free entropy and order parameters for shallow networks (L = 1) with generic activations; Result 3 covers two hidden layers (L = 2); Result 4 covers arbitrary depth L. Result 2 yields the Bayes generalisation error in all cases, providing the Bayes-optimal limits of learning an MLP target.
-
A learning phase diagram and algorithmic study for L ≤ 2. The experimental section validates the theory numerically, identifies the phases and transitions as a function of the data budget, and examines whether practical training algorithms reach optimal performance or are blocked by statistical-computational gaps.
Main Findings
-
A universal phase precedes specialisation. As the sample rate α increases, the network first operates in what the paper calls a "universal" phase, before a critical sample rate α_sp. There, the network predicts using specific non-linear combinations of the teacher's features without disentangling them, effectively learning the best "quadratic network approximation" of the target, and performance is asymptotically independent of the detailed law of the target hidden weights.
-
Specialisation transitions appear, one per layer. With enough data, optimal performance is attained through the model's specialisation towards the target — in the teacher-student setting, recovery of the target weights. In the deep case there is one specialisation transition per layer, so the model exhibits a rich phenomenology of learning transitions.
-
Specialisation is inhomogeneous. It occurs not only across layers, propagating from shallow towards deep ones, but also across neurons within each layer.
-
Deeper targets are harder to learn. The paper states that deeper targets are harder to learn, and that for many targets the time for the network to specialise grows as exp(c·d), where c < 1 is a positive, activation-dependent constant.
-
An optimally trained network beats kernels and random features. Figure 2 compares the Bayes-optimal mean-square generalisation error of a two-layer MLP against the random feature model and a kernel-related baseline. The random feature model is given width r = 3kd, roughly three times larger than the total number k(d+1) of parameters of the NN, and is trained by exact empirical risk minimisation with L2 regularisation strength chosen by cross-validation; the GAMP-RIE estimator (extended to generic activations) reaches the performance of an optimally regularised kernel. The optimally trained NN still outperforms them, and the reason is that it can specialise whereas the random feature model cannot.
-
Numerical validation of the theory. In Figure 2 the theoretical curves follow from the results of Sec. II, with Gaussian inputs and noisy responses from a target two-layer NN with Gaussian random weights, activation ReLU or tanh(2x), d = 150 and k = 75, in the scaling limit where d, n and the width k diverge with n/d² and k/d → 0.5 held fixed. Empty circles are Bayesian NNs trained by Hamiltonian Monte Carlo initialised close to the target (yielding the best achievable error), evaluated on 10^5 test data, with error bars as the standard deviation over 10 instances of the training set and target.
-
Not all tractable architectures can specialise. Shallow quadratic networks reduce to a matrix sensing problem and, since the target depends only on the Gaussian weight matrix through a product, rotational invariance prevents recovery of the target weights, so no specialisation transitions can occur. The paper also defines "weak feature learning" as regimes where feature learning occurs but the model's inner weights have zero overlap with the teacher's.
-
Depth imposes activation restrictions for tractability. For L = 1, any activation admitting a Hermite expansion with vanishing zeroth coefficient μ0 = 0 is allowed (relaxed in App. B.1.7). For L = 2, μ0 = μ2 = 0 is required, as for odd activations such as tanh; a product term W^(2)W^(1) then appears and requires its own order parameter. For L ≥ 3, μ0 = μ1 = μ2 = 0 is required. The authors state these restrictions are practical, not fundamental: relaxing them yields a combinatorial explosion in L of the number of order parameters to track.
-
Algorithmic reachability depends on the target. For L ≤ 2 the answer to whether practical training algorithms reach optimal performance depends on the target, in particular on whether its readout weights are discrete. The theory predicts sub-optimal solutions that can attract training algorithms.
Methodology in Plain English
The authors set up a "teacher-student" experiment as a theoretical model. A teacher network with randomly drawn weights generates the labels; a student network with the same architecture is assumed to be trained in a Bayesian way, and the analysis asks how well it can do on average. Because the teacher and student have matching architectures, priors and output noise, the student is Bayes-optimal, which means no other method trained on the same data can do better — so the results are fundamental limits rather than properties of one particular algorithm.
The technical engine is the replica method from spin-glass physics, which converts the average over random data and random target weights into a problem about a small set of summary quantities (order parameters) that describe what the student has learnt. The new ingredient is a combined use of the HCIZ spherical integral with the replica method, which lets the authors handle matrix-valued degrees of freedom that are correlated with each other — something earlier work avoided by assuming rotational invariance or by treating different phases with separate formalisms.
The analysis then takes a large-system limit: the input dimension, the layer widths and the number of samples all diverge together, with each width proportional to the input dimension and the number of samples proportional to its square. In this "interpolation" regime the number of trainable parameters and the number of data points are comparable, which forces the model to adapt to the task rather than relying on a fixed embedding.
The theoretical predictions are checked numerically: the authors solve the order-parameter equations, compare the predicted generalisation error to Hamiltonian Monte Carlo sampling of the Bayesian posterior initialised near the target, and compare against random feature and optimally regularised kernel baselines. They also extend the setting to structured inputs (Gaussian with a covariance) and to a real dataset, MNIST, and describe an algorithm called GAMP-RIE (generalised approximate message-passing with rotational invariant estimator) that reaches optimally regularised kernel performance.
Why This Matters
The paper argues that statistical physics is not limited to narrow networks or kernel methods, and can describe deep non-linear networks in the regime where feature learning genuinely occurs. It supplies the Bayes-optimal benchmark against which any training procedure on the same data can be measured, and shows which features a perfectly trained network extracts at each data budget. The authors note that modern architectures, including generative diffusion models and large language models, train near interpolation: compute-optimal training scales parameters and tokens in equal proportion, with typical sizes in the range 10^10 to 10^12, and these models are both highly expressive and show signs of feature learning. MLPs are also a basic building block of LLMs, so the authors hope qualitative insights from this idealised setting carry over.
Real-world applications this connects to:
- Large language models and foundation models, which are trained in a regime with parameters and tokens scaled in equal proportion and reach sizes of 10^10 to 10^12.
- Generative diffusion models, cited as another class of highly expressive architectures trained near interpolation.
- Small-data image classification, since the paper generalises its results to structured data and tests on MNIST.
- Kernel and random feature pipelines, where the paper quantifies how much performance is lost by freezing the embedding layer, using a random feature model with width r = 3kd as a concrete baseline.
Industry relevance: the results identify when a network's extra trainable parameters translate into genuine task adaptation rather than behaviour equivalent to a fixed-feature kernel, which matters for deciding model size, data budget and whether to invest in feature-learning architectures versus cheaper kernel-approximation methods.
Future Directions
- Relaxing the activation hypotheses. The paper's restrictions on μ0, μ1 and μ2 as depth grows are attributed to a combinatorial explosion in L of the number of order parameters and to cumbersome formulas. Analysing the most general case is explicitly left for future work.
- Extending beyond static analysis. The paper states it makes no theoretical claims about how learning occurs during training; a dynamics-level theory of how specialisation is reached (or missed) at depth is open.
- Bridging theory and practical algorithms. The observed gap between achievable optimal performance and what training algorithms find — especially when the target's readout weights are discrete and algorithms get attracted to sub-optimal solutions — calls for algorithms and analyses that can close it.
- Generalising the architecture and data assumptions. The setting assumes matched teacher and student and a Bayes-optimal posterior; the authors note the approach can be extended to account for prior mismatch, and they begin generalising the inputs to structured data and MNIST, leaving broader architectures and data types open.
Target Audience
Theoretical machine learning researchers and statistical physicists working on the statistical mechanics of learning, particularly those interested in replica methods, matrix models and Bayesian asymptotics. It is also relevant to mathematically inclined deep learning researchers seeking precise statements about feature learning and interpolation, and to researchers studying optimisers who want to know how far practical training is from Bayes-optimal performance. The paper is not beginner-friendly: it assumes comfort with spin-glass theory, order parameters and random matrix integrals.
Authors’ abstract
For four decades statistical physics has been providing a framework to analyse neural networks. A long-standing question remained on its capacity to tackle deep learning models capturing rich feature learning effects, thus going beyond the narrow networks or kernel methods analysed until now. We positively answer through the study of the supervised learning of a multi-layer perceptron. Importantly, (i) its width scales as the input dimension, making it more prone to feature learning than ultra wide networks, and more expressive than narrow ones or ones with fixed embedding layers; and (ii) we focus on the challenging interpolation regime where the number of trainable parameters and data are comparable, which forces the model to adapt to the task. We consider the matched teacher-student setting. Therefore, we provide the fundamental limits of learning random deep neural network targets and identify the sufficient statistics describing what is learnt by an optimally trained network as the data budget increases. A rich phenomenology emerges with various learning transitions. With enough data, optimal performance is attained through the model's "specialisation" towards the target, but it can be hard to reach for training algorithms which get attracted by sub-optimal solutions predicted by the theory. Specialisation occurs inhomogeneously across layers, propagating from shallow towards deep ones, but also across neurons in each layer. Furthermore, deeper targets are harder to learn. Despite its simplicity, the Bayes-optimal setting provides insights on how the depth, non-linearity and finite (proportional) width influence neural networks in the feature learning regime that are potentially relevant in much more general settings.