Research
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
Overview Research area: AI safety and alignment, specifically the scaling behavior of AI failure modes; also touches on evaluation methodology, bias–variance decomposition, and reasoning-model inferen

- arXiv
- 2601.23045
- Published
- 2026-01-30
- Authors
- Alexander Hägele, Aryo Pradipta Gema, Henry Sleight, Ethan Perez, Jascha Sohl-Dickstein
AI summary
Overview
Research area: AI safety and alignment, specifically the scaling behavior of AI failure modes; also touches on evaluation methodology, bias–variance decomposition, and reasoning-model inference scaling.
Technical level: Advanced. The paper builds on statistical bias–variance decomposition (Kullback–Leibler, Brier, and 0/1 error formulations), scaling laws, and reinforcement-learning-trained reasoning models.
Scope: A single-sentence summary: the paper measures whether AI errors become more systematic (bias) or more inconsistent (variance) as models reason longer, act longer, and grow larger.
What This Paper Is About
When an AI system does something other than what we intended, that failure could come in two very different flavors: it could be consistently pursuing the wrong goal (bias), or it could be behaving erratically in ways that do not further any goal at all (variance). The authors ask which of these two dominates as models become more capable and take on longer, more complex tasks. To answer this, they decompose model error into bias and variance and track the ratio between them — a quantity they call error-incoherence — across benchmarks, agentic coding tasks, safety evaluations, and a synthetic optimizer setting.
Key Contributions
-
A bias–variance framework for AI failures. The authors adapt the classical bias–variance decomposition (formulated for regression) to classification settings using a KL-divergence decomposition, and also run experiments with Brier and 0/1 formulations. They report that all three decompositions produce qualitatively similar results.
-
The error-incoherence metric. They define error-incoherence as the proportion of a model's total error attributable to variance rather than bias, which lands in the range [0, 1]. A value of 0 means every error is consistent; a value of 1 means every error is inconsistent. Because it is a relative measure, models with different overall error rates can be compared.
-
Empirical scaling results across five task families. They measure incoherence on the GPQA and MMLU multiple-choice benchmarks, SWE-Bench agentic coding, the advanced AI risk subset of Model-Written Evals (in both multiple-choice and open-ended formats), a synthetic optimizer task, and a human survey.
-
Interventions that change incoherence. They test reasoning budgets and ensembling, finding that ensembling reduces variance at a rate of 1/E (where E is ensemble size) without affecting bias, and that reasoning budgets reduce incoherence but far less than the effect of natural variation in reasoning length.
Main Findings
-
Longer reasoning and action sequences mean more incoherent failures. Across every setup the authors measure — GPQA, SWE-Bench, Model-Written Evals, and the synthetic optimizer — grouping questions by average reasoning length or number of actions and bucketing them shows error-incoherence increasing significantly with length. The slopes differ by model and model family. Notably, for Qwen3, incoherence levels and slopes are nearly identical across all model sizes even though larger models perform better.
-
Natural variation in reasoning length predicts incoherence better than task difficulty alone. When samples are split into above-median and below-median reasoning length per question (GPQA) or above/below-median actions per question (SWE-Bench), the longer group is substantially more incoherent. Average accuracy and SWE-Bench score are similar between the two groups, and this effect is much stronger than the effect of increasing the reasoning budget.
-
Scaling effects depend on task difficulty. On MMLU with the Qwen3 family (same architecture, up to 32B parameters, thinking enabled), performance improves consistently with model size, with the fastest improvement on the hardest questions. But error-incoherence drops with scale on easy questions while remaining constant or increasing on the hardest ones. The authors attribute this to different scaling slopes: bias slopes are similar across difficulty groups, but variance slopes decrease sharply for harder groups, and in the hardest group variance slopes fall below bias slopes, leaving variance as the limiting factor.
-
Human survey ranking agrees. Using results from the survey of Sohl-Dickstein (2023) — previously released in blog form — disjoint groups of human subjects independently ranked the intelligence and the coherence of AI models, humans, non-human beings, and organizations. Entities judged more intelligent by one group were judged more incoherent by another group, across all categories.
-
A synthetic optimizer task shows variance dominating asymptotically. Training transformers of increasing size to emulate an optimizer descending a quadratic loss, the trained models produce rollouts with lower loss as they scale, but the final loss becomes more variance dominated. Larger models reduce bias much faster than variance. Smaller models reach a lower error-incoherence plateau after a tipping point where they can no longer follow the correct trajectory and stagnate, which reduces variance.
-
Reasoning budgets help, but modestly. Increasing reasoning budgets improves performance and slightly reduces error-incoherence for all models tested except Sonnet 4. The paper notes this effect is overshadowed by incoherence arising from natural variation, i.e., from models thinking longer than the median for a question. The authors state that the implementation details of reasoning budgets for frontier models are not public, so the mechanism is unclear; they speculate it relates to better backtracking and error correction.
-
Ensembling reduces incoherence as theory predicts. For GPQA with o4-mini, using 320 samples per question, the authors average probabilities over ensembles of size E. Variance drops like the inverse of ensemble size and error-incoherence therefore drops, while bias is unaffected.
-
Bias can itself be split. The authors decompose Bias into Bias_mesa (deviation of model behavior from the training objective) and Bias_spec (deviation of the training objective from the intended objective). They believe there was no meaningful reward misspecification in their tasks, but worry that Bias_spec would dominate in settings with poorly specified training objectives.
Methodology in Plain English
The authors start from a standard statistical idea: the expected error of a predictor equals Bias² + Variance. They adapt this to language models by treating a single fixed model as the source of randomness (rather than retraining many models, as classical bias–variance analysis does), taking the expectation over sampling randomness and few-shot context randomness for the same task. For multiple-choice tasks they use a KL-divergence version of the decomposition; for a model to have a "bias," there must be a well-defined target, so they select tasks with objective ground truth — correct answers on GPQA and MMLU, passing unit tests on SWE-Bench, and the optimum on a quadratic function.
To estimate bias and variance reliably, they collect at least 30 samples per question, each with a different seed for autoregressive generation; GPQA and MMLU samples additionally use a different random few-shot context. For open-ended safety questions from Model-Written Evals, where no single correct answer exists, they embed only the answers (not the reasoning chains) using text-embedding-3-large and report the variance of the embedding vectors in Euclidean norm. For SWE-Bench, each sample becomes a binary vector marking which of the task's unit tests the generated code passes; coverage error is the mean squared difference from a vector of all 1s, decomposed into bias and variance.
The frontier models evaluated are Sonnet 4 with reasoning enabled, o3-mini, and o4-mini. For scaling analysis with respect to model size as an imperfect proxy for intelligence, they use the Qwen3 family with thinking enabled. In the synthetic setting, they train autoregressive transformers of varying sizes using decoding-based regression and teacher forcing, with a vocabulary of digits and signs, on trajectories produced by steepest descent with a fixed step norm on an ill-conditioned quadratic with condition number 50. Questions are grouped by reasoning length — using Qwen3 32B as a reference model — as a proxy for task complexity.
Why This Matters
The paper reframes the debate about how advanced AI will fail. Rather than assuming that a capable model will consistently pursue a misaligned goal, the results suggest that failures on long, complex tasks will increasingly look like unpredictable, inconsistent behavior — closer to an industrial accident than to a coherent power grab. This shifts the relative priority toward research on reward hacking and goal misspecification, and toward understanding the mechanistic origins of incoherence.
Real-world applications discussed in the paper:
- Critical software development — the authors note that AI is already relied on for writing critical software, and SWE-Bench measures exactly this kind of agentic coding work.
- Bail determination — cited as a consequential high-stakes deployment where consistent, predictable behavior matters.
- News feed curation — cited as a domain where AI decides what stories to present.
- Industrial and physical systems — the authors predict a future with industrial accidents caused by unpredictable AI misbehavior, but less consistent pursuit of a misaligned goal.
Industry relevance: The findings matter for anyone deploying reasoning models with extended thinking budgets, agentic tool-use loops, or multi-step pipelines. The authors emphasize that actions in the real world are often irreversible, so noise introduced by model actions frequently cannot be corrected — unlike in a benchmark, where resampling is cheap. Ensembling works as an error-correction technique in benchmark settings, but the authors explicitly note it is impractical for action loops in the world, since state typically cannot be reset.
Future Directions
-
Explain the mechanism. The authors state plainly that they do not experimentally or theoretically explore why error-incoherence rises with trajectory length and model size. They motivate it with an argument that optimizers occupy a measure-zero set of dynamical systems, and that broader capabilities expand the effective state and action space — but this is not tested.
-
Extend to open-ended goals. The current framework requires a measurable target. The authors call the extraction of hidden goals and complex incoherent behaviors important, and present the embedding-variance analysis of Model-Written Evals as only an initial exploration of a setting where bias is not easily defined.
-
Study reward misspecification directly. In settings with poorly specified training objectives, the authors expect Bias_spec to dominate once variance and Bias_mesa shrink with capability. They call for characterizing and mitigating goal misspecification during training.
-
Find broader classes of error correction. The small incoherence reduction attributed to reasoning budgets may operate through a mechanism similar to ensembling. The authors expect other error-correction techniques to reduce incoherence as well and frame this as an open direction.
Target Audience
AI safety and alignment researchers will benefit most, particularly those working on misalignment scenarios, reward hacking, and goal misspecification. The paper is also useful for evaluation researchers interested in output variance and benchmark reliability, for practitioners deploying reasoning models and agents on long-horizon tasks, and for policy-oriented readers trying to weigh different AI risk scenarios. Some statistical background — bias–variance decomposition, KL divergence, and scaling laws — is assumed, so beginners will find the appendices and conceptual figures more accessible than the decomposition math.
Authors’ abstract
As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will they fail by systematically pursuing goals we do not intend? Or will they fail by being a hot mess, and taking nonsensical actions that do not further any goal? We operationalize this question using a bias-variance decomposition of the errors made by AI models: An AI's \emph{error-incoherence} on a task is measured over test-time randomness as the fraction of its error that stems from variance rather than bias in task outcome. Across all tasks and frontier models we measure, the longer models spend reasoning and taking actions, \emph{the more incoherent} their failures become. Error-incoherence changes with model scale in a way that is experiment dependent. However, in several settings, larger, more capable models are more incoherent than smaller models. Consequently, scale alone seems unlikely to eliminate error-incoherence. Instead, as more capable AIs pursue harder tasks, requiring more sequential action and thought, our results predict failures to be accompanied by more incoherent behavior. This suggests a future where AIs sometimes cause industrial accidents (due to unpredictable misbehavior), but are less likely to exhibit consistent pursuit of a misaligned goal. This increases the relative importance of alignment research targeting reward hacking or goal misspecification.