Research
Deriving Neural Scaling Laws from the statistics of natural language
Overview Research area: Deep learning theory and the physics of learning, specifically the theory of neural scaling laws for large language models trained on natural text. Technical level: Advanced. T

- arXiv
- 2602.07488
- Published
- 2026-02-07
- Authors
- Francesco Cagnetta, Allan Raventós, Surya Ganguli, Matthieu Wyart
AI summary
Overview
- Research area: Deep learning theory and the physics of learning, specifically the theory of neural scaling laws for large language models trained on natural text.
- Technical level: Advanced. The paper combines information theory (conditional entropy, entropy rate), statistical physics ideas (power-law decays, scaling collapse, universality classes), and large-scale transformer training experiments.
- Scope: The paper derives a parameter-free formula predicting the data-limited neural scaling exponent of LLMs directly from two measurable statistics of a text corpus, and verifies it against GPT-2 and LLaMA-style models trained on TinyStories and WikiText.
What This Paper Is About
Large language models improve predictably as they are given more training data, following approximate power laws known as neural scaling laws, but no existing theory could quantitatively predict the exponent of those power laws for a real LLM on a real language dataset. This paper claims the first such theory in the data-limited regime, arguing that the exponent is determined entirely by two statistical properties of the text: how quickly next-token conditional entropy decays as conditioning context grows, and how quickly token-token correlations decay with temporal separation. The authors measure these two properties in two corpora and check whether the resulting prediction matches what they actually observe when training transformers from scratch.
Key Contributions
- A first-principles prediction of the data-limited scaling exponent. The paper proposes the formula α_D = γ / (2β), where γ is the exponent of the power-law decay of the next-token conditional entropy with context length and β is the exponent of the power-law decay of token-token correlation strength with temporal lag. The formula requires no free parameters and no synthetic data model.
- A scaling-collapse prediction for individual n-gram learning curves. The theory predicts that the family of n-gram losses ℒ_n(P) should collapse onto a single master curve under the rescalings ℒ_n → n^γ ℒ_n and P → P / n^{2β}, expressed as ℒ_n(P) ≡ n^{−γ} ℓ(P / n^{2β}).
- Empirical estimation of two dataset-level exponents across architectures. The authors measure γ and β for TinyStories and WikiText-103, and confirm that γ is consistent across GPT-2-style transformers with absolute positional embeddings, GPT-2-style transformers with rotary positional embeddings (RoPE), LLaMA-style transformers, and also non-transformer models (infini-gram and Mamba).
- Verification against trained models on two qualitatively different corpora. The predicted exponents match empirically measured learning curves for GPT-2 and LLaMA-style transformers trained from scratch on both datasets, with reproducible code released at https://github.com/fracagnetta/small-language-modelling.
Main Findings
- Two statistics suffice. The paper isolates the decay of pairwise token correlations with time separation, and the decay of next-token conditional entropy with conditioning context length, as the two properties of language that alone predict neural scaling exponents.
- Measured entropy exponents. For TinyStories, γ = 0.325 ± 0.003; for WikiText, γ = 0.265 ± 0.016. These values are slightly larger than the 0.23 reported by Takahira et al. (2016), who used different data (mostly news articles), character-level tokens, and estimated entropies via compression.
- Measured correlation exponents. For TinyStories, β = 0.88 ± 0.06; for WikiText, β = 0.94 ± 0.16.
- Predicted versus observed scaling exponents. Plugging in the TinyStories values gives α_D ≈ 0.325 / (2 × 0.88) ≈ 0.185, which matches the experimental learning curves. For WikiText, γ = 0.27 and β = 0.94 yield α_D ≈ 0.265 / (2 × 0.94) ≈ 0.141, also consistent with the empirical scaling.
- Uncertainty on the predictions. Propagating uncertainty gives α_D = 0.185 ± 0.013 for TinyStories and α_D = 0.141 ± 0.025 for WikiText.
- Collapse holds. In the rescaled variables, all n-gram learning curves collapse onto a single curve across the whole range of n and P probed, for both TinyStories and WikiText, under a fixed GPT-2-style APE transformer and across varying maximal context lengths T.
- Context-size dependence. The prediction matches empirical curves especially at larger context sizes T, as the theory expects; the theory is expected to break down when n*(P) approaches T.
- The prediction time horizon grows as a power law. The maximal context window the model can beneficially leverage scales as n*(P) ≍ P^{1/(2β)}, so more data mainly extends the usable context rather than refining performance within it.
- Fast learning within the horizon. For n-gram losses with n ≤ 12, the fitted excess-loss exponents δ_n are all greater than γ/(2β), supporting the assumption that within-horizon errors decay faster than the P^{−γ/(2β)} prediction.
- Broken power law in WikiText. WikiText correlations are better described as a broken power law with two stages; the authors use the short-lag stage n ≲ 32, the range where n*(P) falls for the P values considered. A localized peak near n ≈ 10 also appears, more prominently in WikiText than TinyStories, and is not taken into account.
- Confidence intervals. Using bootstrap resampling to obtain 95% confidence intervals, the predicted range overlaps empirical exponents for up to 10 of the first 12 points on TinyStories and for all 8 points on WikiText.
- Sensitivity of the collapse. For TinyStories, collapse deteriorates when β is taken outside its standard-error interval and when γ falls outside [0.31, 0.34]; for WikiText, outside [0.23, 0.30]. An exception is the LLaMA model trained on TinyStories, whose collapse appears visually sharper for β ∈ [0.6, 0.7], though the corresponding scaling law remains compatible with the predicted α_D.
Methodology in Plain English
The authors start from a simple decomposition of the loss: part of the error comes from how far back in the text the model can usefully look (the prediction time horizon), and part from how well it uses the tokens inside that horizon. They argue that the first effect dominates. To look back n tokens profitably, a model needs enough data to distinguish a real token-token correlation at lag n from random noise in the finite sample; because sampling noise shrinks like P^{−1/2} while the signal decays like n^{−β}, the usable horizon grows as P^{1/(2β)}. Since the best achievable loss at horizon n*(P) is the conditional entropy H_{n*(P)}, and that entropy decays as n^{−γ}, substituting one into the other yields α_D = γ/(2β).
Measuring γ directly from raw counts is infeasible because the number of distinct contexts grows exponentially with n, so the authors instead train increasingly large models and treat their n-gram losses as progressively tighter upper bounds on the true conditional entropy, fitting a power law to the small-n portion of the curve from the largest model. They check this limit is architecture-independent by repeating it across several model families. β is measured more directly, from empirical token co-occurrence counts over the full training set, using the operator norm (largest singular value) of the token-token covariance matrix as a scalar summary, and reported alongside the closely tracking Frobenius norm. Finally, they train transformers from scratch on slices of P tokens from each corpus and compare the measured learning-curve slopes to the parameter-free prediction.
Why This Matters
Impact on research. The paper offers a route from measurable statistics of a corpus to the exponents of the learning curves a model will exhibit on it, replacing a purely empirical law with a quantitative, testable one. It also connects scaling behavior to the hidden hierarchical structure of language, and frames architectures in terms of universality classes, where many systems share the same critical exponents. The authors note that existing kernel-based theories assume fixed features and offer limited guidance for modern feature-learning LLMs, since LLMs learn syntactic and semantic representations during training.
Real-world applications (as implications of the result):
- Data-budget planning. Because the loss-versus-data exponent governs the expected returns from collecting more data, a formula that predicts the exponent from corpus statistics directly informs how much additional text is worth acquiring.
- Corpus selection and comparison. Measuring γ and β on a candidate corpus gives an early quantitative indication of how quickly loss will fall with data, without first running a full scaling-law sweep.
- Benchmark design for small-scale studies. The scaling-collapse diagnostic lets practitioners compare n-gram learning curves in rescaled units rather than needing to run every configuration to full scale.
- Model/architecture comparisons. Because γ is reported as consistent across architectures and datasets, it can serve as a dataset property against which architecture-specific deviations are measured.
Industry relevance. The paper notes that scaling relations guide large-scale training decisions across the AI industry, and that the data-limited exponent is particularly important because it governs returns on data collection. The authors also lay out the limits of applicability: within-horizon learning is expected to be fast for a class of deep networks including the LLMs studied, but the assumption is expected to fail for shallow networks, kernel methods, and n-gram models, where the curse of dimensionality with growing context would prevent fast learning. The prediction is also expected to break down at fixed T once n*(P) approaches T.
Future Directions
- Determining which model architectures and training dynamics give rise to the fast within-horizon learning regime the theory assumes. The paper's conclusion raises this question as the text ends; the supplied content does not report an answer.
- Extending the theory beyond the data-limited regime, since the foundational work cited studies model-size-limited, data-limited, and compute-limited regimes, and this paper explicitly focuses only on the data-limited case.
- Interpreting the LLaMA-on-TinyStories discrepancy, where the collapse appears visually sharper for β ∈ [0.6, 0.7], below the standard-error interval, even though the resulting scaling law stays compatible with the predicted α_D. The authors leave this for future work.
- Testing at token scales and context lengths beyond those probed. The authors note that prior hierarchical-structure evidence came from small-scale, character-level settings with very limited context lengths, and that their own comparison holds "in the range of dataset sizes and context lengths probed."
Target Audience
Machine learning theorists and statistical physicists working on scaling laws and learning dynamics; LLM practitioners making data-acquisition and training-budget decisions; and researchers in computational linguistics or cognitive science interested in which statistics of a corpus drive learnability. Familiarity with conditional entropy, power laws, and transformer training is assumed, given the paper's derivations and architecture-specific experiments.
Authors’ abstract
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.