Skip to content
AI.info

Research

How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs

How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs Overview Research area: Theoretical machine learning — specifically the theory of in-context learning (ICL

How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
arXiv
2510.25753
Published
2025-10-29
Authors
Samet Demir, Zafer Dogan

AI summary

How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs

Overview

Research area: Theoretical machine learning — specifically the theory of in-context learning (ICL) in Transformer architectures, combining random matrix theory / Gaussian universality with the study of multi-source data mixtures.

Technical level: Advanced. The paper is written for readers comfortable with high-dimensional asymptotics, random feature models, Hermite polynomial expansions, and Transformer internals (attention, MLPs, ridge regression). The prose is accessible but the machinery is mathematical.

One-sentence scope: The paper proves that a Transformer with a linear attention block and a trained two-layer nonlinear MLP head is asymptotically equivalent — in terms of ICL error — to a finite-degree polynomial predictor, and uses that equivalence to analyze how mixing multiple data sources of differing quality shapes ICL performance and feature learning.

Note: the provided paper content is truncated mid-way through the experimental results section (the text cuts off in the caption of Figure 2). Findings from the real-world multilingual sentiment analysis experiment are described only at a high level in the abstract and introduction; the specific numerical results of that experiment are not present in the available content.

What This Paper Is About

Most theoretical work on in-context learning simplifies the problem heavily: it studies linear regression tasks, drops the MLP layers entirely, and assumes training data comes from a single homogeneous source. Real Transformers, however, have nonlinear MLPs, face nonlinear tasks, and are pretrained on mixtures of many heterogeneous datasets of varying quality. This paper asks two questions in response: how do nonlinear MLPs change ICL behavior under realistic data conditions, and what role does the mixture of training sources play in shaping a Transformer's ICL ability and its internal feature learning?

Key Contributions

  1. An asymptotic equivalence result for Transformers with MLPs. The authors prove that a Transformer with linear attention and a two-layer nonlinear MLP head (first layer updated by a single gradient step, second layer fully trained by ridge regression) achieves the same ICL error as a structured polynomial predictor of finite degree, built from a truncated Hermite expansion of the activation function. This extends Gaussian universality results from two-layer networks to Transformer-based ICL.

  2. A characterization of high-quality data sources. Through the equivalent polynomial model, the authors identify which properties make a training source valuable for ICL: low target noise and structured (non-isotropic) covariance in the input and task distributions. They analyze how the mixing ratio across sources affects the resulting ICL error.

  3. A link between data mixing and feature learning. The paper shows that feature learning emerges only when the task covariance exhibits sufficient structure — meaning that some mixtures of data sources can enhance, and others can suppress, the model's ability to learn useful nonlinear features.

  4. Empirical validation plus a real-world demonstration. The equivalence and the data-mixing insights are tested across different activation functions (including ReLU and tanh), model sizes, and data distributions, and are further illustrated with a multilingual sentiment analysis experiment in which each language is treated as a separate data source.

Main Findings

  • Nonlinear MLPs beat linear baselines on nonlinear ICL tasks. The Transformer with a nonlinear MLP head (the paper's model) achieves lower ICL error than the linear Transformer without MLPs. The authors note this is a non-trivial baseline: without training the first layer with a gradient step, there are cases where the nonlinear MLP head fails to outperform the linear model. Outperforming the linear model therefore serves as an indicator of effective feature learning.

  • The Transformer matches its polynomial surrogate. Across the experiments in Figure 1, the nonlinear MLP Transformer closely aligns with its equivalent polynomial model even at moderate dimensionality, supporting the asymptotic equivalence claim.

  • Double descent appears in ICL error. The experiments show a characteristic double descent phenomenon in ICL error with respect to sample size n (Figure 1a) and hidden dimension k (Figure 1c).

  • Longer contexts help. Increasing context length ℓ consistently reduces ICL error, reflecting improved task estimation with more in-context examples.

  • Data source quality is defined by noise and covariance structure. Sources with low target noise and structured covariances (in the input and/or task vector distributions) are the high-quality sources; an isotropic Gaussian source in both input and task space represents a low-quality, noise-like dataset.

  • Feature learning requires task structure. The ability of the model to learn useful features depends crucially on the structure of the task distribution — feature learning only emerges when the task covariance exhibits sufficient structure.

  • Experimentally configured regimes. The simulation studies use two data sources (S = 2) with equal probability, d = 80, ℓ = d, n = k = 0.5d², step size η ≍ d², regularization λ = 5×10⁻⁵, target noise Δ_s = 0.01 (and Δ_0 = 0.2 in the noise-mixing comparison), ReLU target functions, and averages over 20 Monte Carlo runs. The equivalent polynomial degree was set to p = 4 in Figure 1 and p = 5 in Figure 2.

Methodology in Plain English

The authors set up an in-context learning problem where a model sees a context of input–output pairs followed by a query input, and must predict the query's label. Data comes from several sources, each with its own distribution over input vectors, task vectors, noise level, and nonlinear target function. The task vector is fixed within a context but resampled across contexts, so the model must infer it from the demonstrations.

The model is a single block: linear attention followed by a two-layer MLP head. Training is split into two stages — one gradient descent step on the first layer (which merges the attention parameters and the first MLP layer into one matrix F), then ridge regression on the second layer w, using fresh samples so the two stages stay statistically independent. This two-stage scheme is drawn from prior theoretical work on MLPs and keeps the analysis tractable while preserving feature learning.

The analysis works in a high-dimensional limit where the input dimension, context length, number of training samples, and hidden dimension all go to infinity together at fixed ratios. Under these assumptions, the authors show that the pre-activation outputs of the first layer become jointly Gaussian with the task-relevant projection of the query input, and that the gradient matrix decomposes into a dominant rank-one term plus a negligible residual. Combining these with the orthogonality properties of Hermite polynomials under the Gaussian measure yields the main result: the Transformer's ICL error equals that of a polynomial model whose activation is a truncated Hermite expansion of the true activation. Because this surrogate is far simpler, the authors can then derive precise statements about data mixing and feature learning, and they check all of it with simulations and a multilingual sentiment analysis experiment.

Why This Matters

Impact on research. Prior ICL theory largely assumed linear tasks, attention-only architectures, or single homogeneous data sources. This paper connects Gaussian universality theory to Transformers with MLPs for nonlinear ICL, and it is among the first to explicitly model heterogeneous multi-source training data and the effect of mixture ratios. It also gives a concrete characterization of what the model learns — a low-degree polynomial approximation of the task function — which opens the door to designing or optimizing MLP nonlinearities through the polynomial surrogate.

Real-world applications:

  • Pretraining data curation and mixing decisions — deciding which datasets to include and in what proportions, guided by the finding that low-noise, structured sources drive ICL quality.
  • Multilingual model development — the paper's own sentiment-analysis example treats each language as a separate source, which is directly relevant to balancing language data in multilingual models.
  • Model architecture choices — evidence that nonlinear MLP heads matter for ICL on nonlinear tasks, informing trade-offs between architecture complexity and capability.
  • Benchmark and evaluation design — the per-source ICL error decomposition (reported in the paper's appendix) suggests evaluating models per data source rather than only on an average.

Industry relevance. Anyone making pretraining data-mixture decisions — a central practical lever in frontier model training — can use the qualitative criteria identified here: prefer low-noise sources and sources with structured covariances, and be aware that certain mixtures can suppress feature learning rather than help it.

Future Directions

  • Relaxing the analysis assumptions. The authors explicitly flag Assumptions 4.3–4.5 (step-size scaling, low-rank covariance of attention outputs, joint scaling) as limitations, and note that their synthetic and real-world experiments suggest these can be partially relaxed in practice. Making the theory hold more generally is a natural next step.
  • Extending beyond a single attention-plus-MLP block. The analysis covers one block; generalizing to stacked attention and MLP layers remains open.
  • Replacing linear attention with softmax attention. The paper uses linear attention for tractability, leaving the question of how nonlinear attention mechanisms interact with the MLP-based equivalence.
  • Turning the equivalence into a design tool. Since the Transformer behaves like a polynomial predictor, the surrogate could be used to choose or optimize activation functions and mixture ratios rather than only to analyze them.

Target Audience

This paper is most valuable to theory-oriented machine learning researchers working on in-context learning, Gaussian universality, random matrix theory, or the training dynamics of two-layer networks. It is also relevant to practitioners and engineers who make pretraining data-mixture decisions for Transformers, and to anyone interested in why MLP layers matter for in-context learning. A reader needs a solid mathematical background — high-dimensional asymptotics, Hermite polynomials, and ridge regression — to follow the derivations, though the experimental sections and the qualitative conclusions about data quality are readable without the full proofs. Beginners will likely find the theoretical core inaccessible.

Authors’ abstract

Pretrained Transformers demonstrate remarkable in-context learning (ICL) capabilities, enabling them to adapt to new tasks from demonstrations without parameter updates. However, theoretical studies often rely on simplified architectures (e.g., omitting MLPs), plain data models (e.g., linear regression with isotropic inputs), and single-source training, limiting their relevance to realistic settings. In this work, we study ICL in pretrained Transformers with nonlinear MLP heads on nonlinear tasks drawn from multiple data sources with heterogeneous input, task, and noise distributions. We analyze a model where the MLP comprises two layers, with the first layer trained via a single gradient step and the second layer fully optimized. Under high-dimensional asymptotics, we prove that such models are equivalent in ICL error to structured polynomial predictors, leveraging results from the theory of Gaussian universality and orthogonal polynomials. This equivalence reveals that nonlinear MLPs meaningfully enhance ICL performance, particularly on nonlinear tasks, compared to linear baselines. It also enables a precise analysis of data mixing effects: we identify key properties of high-quality data sources (low noise, structured covariances) and show that feature learning emerges only when the task covariance exhibits sufficient structure. These results are validated empirically across various activation functions, model sizes, and data distributions. Finally, we experiment with a real-world scenario involving multilingual sentiment analysis where each language is treated as a different source. Our experimental results for this case exemplify how our findings extend to real-world cases. Overall, our work advances the theoretical foundations of ICL in Transformers and provides actionable insight into the role of architecture and data in ICL.

Read the original paper