Research
On the Role of Hidden States of Modern Hopfield Network in Transformer
Overview Research area: Deep learning architecture — the theoretical connection between modern Hopfield networks (MHN) and Transformer self-attention, and a new attention variant derived from that con

- arXiv
- 2511.20698
- Published
- 2025-11-24
- Authors
- Tsubasa Masumura, Masato Taki
AI summary
Overview
Research area: Deep learning architecture — the theoretical connection between modern Hopfield networks (MHN) and Transformer self-attention, and a new attention variant derived from that connection.
Technical level: Advanced. The paper combines dynamical-systems derivations (adiabatic limit, discretization of coupled state equations) with a rank-collapse convergence proof and large-scale training experiments.
Scope: The paper derives a hidden-state version of self-attention from the non-adiabatic modern continuous Hopfield network, proposes "modern Hopfield attention" (MHA), and tests it on GPT-2, LLaMA, and Vision Transformer, with theory and experiments on rank collapse.
What This Paper Is About
Earlier work showed that the state update rule of the modern Hopfield network in the adiabatic approximation matches the self-attention layer of the Transformer. The authors go beyond that approximation: they keep the hidden-state dynamics that the adiabatic limit discards and ask what it implies for Transformer architecture. The result is a new attention mechanism, modern Hopfield attention (MHA), in which attention scores are accumulated across layers in a hidden state, and which the authors argue improves attention weights and mitigates rank collapse without adding training parameters.
Key Contributions
- Extends the Hopfield–Transformer correspondence beyond the adiabatic approximation by carefully discretizing the coupled visible/hidden state equations of the modern continuous Hopfield network (MCHN), producing a generalized attention rule with two parameters, α and α′.
- Proposes modern Hopfield attention (MHA), which carries a hidden state H that accumulates each layer's attention score Q Kᵀ as an exponential moving average, so that attention weights inherit score information from other layers. MHA adds no training parameters and adds only about O(T²) computation on top of the O(dT²) cost of self-attention.
- Demonstrates systematic perplexity and accuracy improvements on GPT-2 Small/Medium (WikiText103), a LLaMA architecture (WikiText-103, CNN DailyMail, BookCorpus), and ViT-Tiny/Small/Base/Large (CIFAR10/CIFAR100, ImageNet-1k) plus four downstream transfer datasets.
- Proves a theorem showing that a non-zero α′ relaxes the double-exponential rank decay of attention-only networks (Theorem 5.1 of Dong et al.) to linear decay, and supports it with skip-free depth experiments and cosine-similarity measurements.
Main Findings
- GPT-2 perplexity improves: On WikiText103, GPT-2 Small (124M) went from 22.87 (self-attention) to 20.70 with MHA (α = 0.5); GPT-2 Medium (350M) went from 20.85 to 19.61. α was simply set to 0.5 based on a rough hyperparameter search.
- LLaMA also improves: Tested with a miniLLaMA implementation at α = 0.5, perplexity dropped from 14.49 to 14.29 on WikiText-103, from 19.36 to 18.97 on CNN DailyMail, and from 23.76 to 23.50 on BookCorpus.
- Vision Transformer gains grow with model size on CIFAR100: For ViT-Tiny (5.5M), self-attention scored 73.080 versus 72.030 for MHA(α = 0.5) and 72.570 for MHA(α = 0.7); for ViT-Small (22M), 74.485 versus 75.420 and 75.590; for ViT-Base (86M), 75.360 versus 76.215 and 75.590; for ViT-Large (303M), 72.910 versus 75.775 and 75.365. The paper states MHA's advantage becomes clear for models larger than Small. CIFAR10 results were near saturation with no clear effect, though effects began appearing in Base and Large.
- ImageNet-1k improvement is under 1% but non-trivial: ViT-B (86M) trained for 300 epochs reached 76.074 (self-attention), 76.434 (MHA, α = 0.5) and 77.058 (MHA, α = 0.7).
- Downstream transfer results are mixed: With α = 0.7, MHA beat self-attention on Flower102 (93.85 vs 81.15), Food101 (87.99 vs 74.51) and Stanford cars (87.54 vs 51.54), but on Stanford dogs self-attention scored 95.00 versus 83.64 for MHA — the paper frames the MHA variants as achieving "consistently good transfer performance."
- Both α and α′ matter: In a ViT-T CIFAR100 experiment, varying α′ with α fixed at 0.5 gave scores from 66.10 (α′ = 1.0) up to 72.29 (α′ = 0.2); varying α with α′ fixed at 0.5 gave 1.00 at α = 1.0, 69.89 at α = 0.0, and 72.66 at α = 0.6. Setting either parameter to 0 degrades performance relative to the joint setting.
- Depth degradation is much milder with MHA: In skip-free networks based on ViT-T, self-attention peaked at depth 2 and fell sharply beyond depth 4 (depth 4: 57.38 on CIFAR10, 32.25 on CIFAR100; depth 8: 48.59 and 17.19; depth 12: 10.00 and 1.00), whereas MHA kept improving to depth 4 (85.74 on CIFAR10, 64.39 on CIFAR100) and was still far above the baseline at depth 8 (80.34 and 49.90). Both configurations collapsed to 10.00 and 1.00 at depth 12.
- Token uniformity is reduced: In violin plots of token cosine similarity for GPT-2 Medium (WikiText103) and ViT-B (CIFAR100), ordinary models have a similarity mode of 1.0 in all layers, while MHA eliminates tokens with perfect similarity of 1, though MHA layers with high average token similarity still exist.
- Theory of rank collapse: The authors prove that with non-zero α′, the residual upper bound becomes a max over m = 0 to L rather than the double-exponentially decaying term of Dong et al., so the m = 0 term dominates and rank decay becomes linear; setting α′ = 0 reproduces the original double-exponential decay result. In MHA this effect comes from the hidden state rather than from a skip connection.
Methodology in Plain English
The authors start from the modern continuous Hopfield network, a model with two coupled variables: a visible state x (the feature neurons) and a hidden state h (the memory neurons), each following a differential equation involving time constants and the network's weight matrices. Writing those equations in discrete time introduces a step-size-to-time-constant ratio, which they parameterize as α and α′. In the adiabatic limit used in prior work, the hidden state is assumed to be instantaneously equilibrated, and the update rule collapses exactly to ordinary self-attention. Keeping a finite step size instead leaves the hidden state alive, and translating the resulting equations into Transformer notation yields a rule where the softmax is applied to a hidden state h that is an exponential moving average of previous layers' attention scores, plus a skip connection weighted by (1 − α).
They then implement this as MHA in existing architectures and train from scratch under matched settings: GPT-2 Small (124M) and Medium (350M) on WikiText103, a LLaMA implementation on three datasets, and ViT-Tiny through ViT-Large on CIFAR10/CIFAR100 with ImageNet-1k for ViT-B, plus linear-probing transfer to four downstream datasets. The experiments were run on up to eight A100 GPUs. For theory, they analyze an attention-only network with no skip connections, following the setup of Dong et al., and derive a new upper bound on how the residual shrinks with depth. They complement this with skip-free networks trained at depths 1, 2, 4, 8, and 12, and with cosine-similarity measurements of token representations.
Why This Matters
Impact on research: The paper argues that the Hopfield-network view of self-attention is useful not only as a reinterpretation but as a source of concrete architectural improvements, and gives a mechanism — hidden-state accumulation of attention scores — that is parameter-free yet affects rank collapse. It also extends attention-score reuse work from encoder-only, technically motivated designs to a principled derivation covering decoder Transformers.
Real-world applications:
- Large language model pretraining, where GPT-2 and LLaMA experiments show perplexity reductions on WikiText103, CNN DailyMail and BookCorpus.
- Vision backbones for image classification, given gains on CIFAR100 and ImageNet-1k and transfer to Flower102, Food101, Stanford Dogs and Stanford Cars.
- Deep Transformer training, where reduced rank collapse and milder depth degradation may make very deep stacks more trainable.
- Deployment-constrained settings where an accuracy gain with no new parameters and only a small complexity increase is attractive.
Industry relevance: Because MHA adds no training parameters and only about O(T²) extra computation relative to the O(dT²) attention cost, it is a drop-in replacement in existing Transformer pipelines rather than a new architecture requiring retraining infrastructure.
Future Directions
- The authors note that they do not tie W₁ and W₂, which breaks the symmetry assumption underlying the monotonically decreasing energy function of Krotov and Hopfield; interpreting the energy function of MHA in this asymmetric setting is described as an interesting theoretical challenge.
- MHA has two hyperparameters, α and α′; the paper mostly studies α = α′, and a systematic study of independently tuned values, and of how to choose them, remains open.
- The theory is proved for attention-only, skip-free networks while the experiments use standard architectures; the authors explicitly note it is unclear how much additional effect MHA has in architectures that already have skip connections.
- Extending the analysis to the full Transformer architecture with FFN layers is stated to be straightforward in principle but is not carried out in the truncated content.
Target Audience
Researchers and graduate students working on Transformer architecture, attention mechanisms, and associative memory models, particularly those familiar with the Hopfield-network/self-attention correspondence and with rank-collapse or oversmoothing literature. It will also interest practitioners looking for a parameter-free attention modification with reported gains on language and vision benchmarks, though the theoretical sections and the rank-collapse proof require comfort with dynamical systems and norm-based convergence analysis.
Authors’ abstract
Associative memory models based on Hopfield networks and self-attention based on key-value mechanisms have been popular approaches in the study of memory mechanisms in deep learning. It has been pointed out that the state update rule of the modern Hopfield network (MHN) in the adiabatic approximation is in agreement with the self-attention layer of Transformer. In this paper, we go beyond this approximation and investigate the relationship between MHN and self-attention. Our results show that the correspondence between Hopfield networks and Transformers can be established in a more generalized form by adding a new variable, the hidden state derived from the MHN, to self-attention. This new attention mechanism, modern Hopfield attention (MHA), allows the inheritance of attention scores from the input layer of the Transformer to the output layer, which greatly improves the nature of attention weights. In particular, we show both theoretically and empirically that MHA hidden states significantly improve serious problem of deep Transformers known as rank collapse and token uniformity. We also confirm that MHA can systematically improve accuracy without adding training parameters to the Vision Transformer or GPT. Our results provide a new case in which Hopfield networks can be a useful perspective for improving the Transformer architecture.