Research
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
Overview Research area: Machine learning — transformer architecture design, optimization theory, and dynamical-systems views of deep networks. Technical level: Advanced (the paper combines Riemannian

- arXiv
- 2601.23236
- Published
- 2026-01-30
- Authors
- Aleksandr Zimin, Yury Polyanskiy, Philippe Rigollet
AI summary
Overview
- Research area: Machine learning — transformer architecture design, optimization theory, and dynamical-systems views of deep networks.
- Technical level: Advanced (the paper combines Riemannian gradient-flow arguments, operator splitting, and Nesterov acceleration with large-scale language-model training).
- Scope: The paper proposes a variational framework in which transformer blocks are discrete optimization steps on token configurations, and uses it to build Nesterov-accelerated transformers ("YuriiFormer") that are compared against nanoGPT baselines on TinyStories and OpenWebText.
What This Paper Is About
Transformer architectures are still largely an empirical design: attention, MLPs, residual connections, and normalization are known to be essential, but their combination is rarely described as a single coherent algorithm. This paper argues that attention and MLP layers are learned first-order gradient oracles for two different energy functionals over token embeddings, so a transformer block is really one iteration of a first-order optimization method on a composite objective. The goal is to show that this viewpoint makes architectural design principled — specifically, by swapping vanilla gradient descent for Nesterov's accelerated method while keeping the same attention and MLP structure.
Key Contributions
-
A variational interpretation of transformer blocks. Self-attention is shown to be a (modulated) gradient step on an interaction energy $\mathcal{E}(X) := \sum_{i,j=1}^{n} e^{\langle x_i, x_j \rangle}$, where the learned value matrix acts as a preconditioner and the query/key matrices act as a change of coordinates. MLP layers are shown to be a gradient step on a potential energy $\mathcal{F}(X) := \sum_{i=1}^{n} V(x_i)$ with $V(x) := \sum_{\ell=1}^{d} v(x^{(\ell)})$ and $v' = \sigma$, preconditioned by the learned matrices and an affine change of coordinates.
-
Standard GPT-style blocks as a splitting scheme. Alternating attention and MLP layers corresponds to vanilla gradient descent on the composite objective $\mathcal{E} + \mathcal{F}$ implemented via Lie–Trotter splitting. A forward Euler discretization instead yields the parallel update used in architectures such as PaLM.
-
The YuriiFormer suite. Replacing the gradient-descent template with Nesterov's accelerated gradient (lookahead, velocity update, state update) yields a two-stream (state–velocity) architecture family that keeps the same attention and MLP oracles and does not increase the number of attention or MLP evaluations per block. Two instantiations are given: Nesterov with Euler discretization and Nesterov with Lie–Trotter splitting.
-
Empirical demonstration plus a modularity check. The accelerated variants are trained against a nanoGPT baseline on TinyStories and OpenWebText, and the same template is instantiated with Polyak's heavy-ball method (Appendix-scale variants such as Verlet and IMEX are also evaluated).
Main Findings
- Nesterov + Lie–Trotter is the strongest variant. On TinyStories (small; 10k steps) it achieves the lowest validation loss: best 1.078 and final 1.090 nats/token, versus the GD+Lie–Trotter baseline at best 1.106 and final 1.114. Its training loss at 10k steps is 0.896 versus 0.985 for GD+Lie–Trotter.
- Consistent ordering on OpenWebText. After the 3k-step warmup, the ordering is stable for both small and medium models: GD+Euler is highest, Nesterov+Lie–Trotter is lowest, and GD+Lie–Trotter, Nesterov+Euler, and the Polyak variants lie in between.
- OpenWebText small (30k steps) best validation losses: GD+Euler 3.014, Polyak+Euler 2.965, Nesterov+Euler 2.959, GD+Lie–Trotter 2.990, Polyak+Lie–Trotter 2.925, Nesterov+Lie–Trotter 2.920.
- OpenWebText medium (30k steps) best validation losses: GD+Euler 2.770, Polyak+Euler 2.730, Nesterov+Euler 2.726, GD+Lie–Trotter 2.758, Polyak+Lie–Trotter 2.705, Nesterov+Lie–Trotter 2.702.
- Lie–Trotter splitting beats Euler discretization. The authors report that this superiority is less pronounced for OpenWebText than for TinyStories, and that the combination of Nesterov with Lie–Trotter yields the maximal improvement over the baseline.
- Downstream accuracy follows the validation-loss ordering. On OpenWebText small, 10-shot HellaSwag (length-normalized accuracy) moves from 30.0 to 31.8 between GD+Lie–Trotter and Nesterov+Lie–Trotter; on medium it moves from 35.5 to 36.8. On ARC-Easy small, the 25-shot gain is 40.5 to 42.6, and the 0-shot gain is 39.4 to 41.1.
- Comparison to GPT-2 references. Under the nanoGPT OpenWebText preprocessing and validation split, GPT-2 (124M) attains a validation loss of 3.12 and GPT-2 medium (350M) attains 2.84. The paper's small model reaches 2.92 and its medium model reaches 2.70, both below the respective GPT-2 checkpoints. nanoGPT can reach 2.85 at the small scale, but with a much larger budget of 600k steps / 294.9B tokens versus 30k steps / 14.75B tokens.
- Polyak variants also help but trail Nesterov. Polyak's heavy-ball method achieves validation losses comparable to the Nesterov variants and improves over the GD baselines; Nesterov's lookahead yields a small additional loss improvement at the same compute and parameter budget, with no extra attention/MLP oracle calls.
- TinyStories overfitting is visible. With about 10 epochs, validation loss reaches its minimum around step 7.6k for all methods and changes little thereafter, while training loss continues to decrease.
- Preliminary oracle-count result. On TinyStories small, under a matched budget of one attention and one MLP oracle call per block, Nesterov+Lie–Trotter is tied for best among the one-attention/one-MLP variants; variants that use additional attention/MLP oracle calls per block can achieve lower losses at higher compute.
Methodology in Plain English
The authors start from a mathematical observation rather than a new module. They write down the formula for a self-attention layer with a residual connection and compare it to the formula for the gradient of a pairwise "interaction energy" defined over token embeddings. The two match once you account for the query, key, and value matrices — which the authors describe as a preconditioner plus a change of coordinates. They then do the same for an MLP layer, matching it to the gradient of a per-token "potential energy" whose derivative is the activation function. Adding the two energies gives a composite objective that a standard transformer block is implicitly optimizing, one layer per step.
From there, the modeling decision is purely about which numerical scheme to use. The usual sequential attention-then-MLP block is Lie–Trotter splitting; a parallel sum of the two updates is forward Euler. The authors replace plain gradient descent with Nesterov's three-step accelerated method — compute a lookahead point, update a velocity from the gradient at that lookahead point, then update the state — applied to the same attention and MLP oracles. This creates a second "velocity" stream of token embeddings alongside the usual state stream, with a LayerNorm applied to the velocity after each velocity update.
Experiments are controlled: a nanoGPT baseline and all YuriiFormer variants are decoder-only language models trained autoregressively with a causal mask, context length 1024, learned position embeddings, pre-normalization with two LayerNorms per layer, GELU FFN with 4× expansion, weight-tied output projection, dropout 0, and no bias terms. All methods use identical batch size, sequence length, number of optimizer steps, and the same deterministic sequence of training batches, so differences reflect the update rule. Models are 12L/12H/768d (small) on TinyStories and OpenWebText and 24L/16H/1024d (medium) on OpenWebText, at 123.6M and 353.6M non-positional parameters respectively. Training uses a mixed Muon + AdamW optimizer.
Why This Matters
Impact on research. The paper reframes architectural design as a choice of optimization template and splitting scheme over a fixed set of learned oracles. If that framing holds, a large body of classical numerical analysis — acceleration schemes, higher-order integrators, alternative splittings — becomes directly importable into transformer design rather than being a source of loose analogies. The authors position this against earlier work that derived transformer variants from ODE splitting schemes and against energy-based reconstructions of attention, arguing that their specific contribution is embedding Nesterov acceleration as a two-stream state–velocity architecture at the representation level while preserving the standard GPT block structure.
Real-world applications implied by the work:
- Language model pretraining under fixed compute budgets, where the paper's own setting shows a small loss improvement at equal parameter count and equal optimizer steps.
- Latency- and memory-sensitive deployment of transformer-like models, since the accelerated variants add no extra attention or MLP evaluations per block and preserve the familiar sequential block structure.
- Architecture search and design tooling, where the optimization view gives a structured menu of update rules (gradient descent, Polyak heavy ball, Nesterov, and further alternatives) instead of ad hoc block modifications.
- Efficient small-model training on constrained datasets, as demonstrated on TinyStories with roughly 10 epochs of training.
Industry relevance. The modifications are drop-in in the sense that attention and MLP sublayers are untouched; only the depth-update rule and its learned scalars change. The paper reports the same non-positional parameter counts as the gradient-descent baselines up to negligible overhead from the update rule, plus a velocity LayerNorm of order $O(Ld)$ and $O(L)$ learned scalars, and momentum variants add separate $v_0$ embedding tables (adding 38.6M parameters for $d = 768$ and 51.5M for $d = 1024$). The evaluation focuses on 12-layer and 24-layer models under controlled budgets rather than production-scale training.
Future Directions
- Scaling beyond small and medium models. The authors explicitly state that evaluating these architectures at larger scales and in longer-context regimes remains an important direction for future work.
- Theory for the implicit objectives. The induced objectives are nonconvex, layer-dependent, and strongly shaped by preconditioning, placing them outside classical optimization theory; the paper treats the algorithms as design analogies rather than solvers with transferable guarantees, which is a stated limitation.
- Exploring more splitting schemes and integrators. The framework admits alternative discretizations; the paper examines Euler and Lie–Trotter in the main text, Polyak's heavy-ball variants, and Verlet and IMEX in the appendix, and the authors note that Strang–Marchuk splitting had previously produced moderate but consistent improvements.
- Understanding when extra oracle calls are worth the compute. The preliminary TinyStories comparison notes that variants using additional attention/MLP oracle calls per block can reach lower losses at higher compute, leaving the compute-accuracy tradeoff among these variants open.
Target Audience
This paper is best suited to researchers and advanced practitioners working on transformer architecture and optimization — particularly those interested in dynamical-systems or variational views of deep networks, in energy-based interpretations of attention, and in numerical schemes such as operator splitting. It will also be useful to engineers who need the precise training protocol and controlled-comparison methodology, and to readers already familiar with Nesterov acceleration and Polyak momentum who want to see those methods instantiated as concrete language-model architectures. Readers looking for formal convergence guarantees or large-scale production results will find those explicitly outside the paper's stated scope.
Authors’ abstract
We propose a variational framework that interprets transformer layers as iterations of an optimization algorithm acting on token embeddings. In this view, self-attention implements a gradient step of an interaction energy, while MLP layers correspond to gradient updates of a potential energy. Standard GPT-style transformers emerge as vanilla gradient descent on the resulting composite objective, implemented via Lie--Trotter splitting between these two energy functionals. This perspective enables principled architectural design using classical optimization ideas. As a proof of concept, we introduce a Nesterov-style accelerated transformer that preserves the same attention and MLP oracles. The resulting architecture consistently outperforms a nanoGPT baseline on TinyStories and OpenWebText, demonstrating that optimization-theoretic insights can translate into practical gains.