Research
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Overview Research area: Natural Language Processing / LLM evaluation and robustness (submitted to the TAE — Trust-AI-Eval: "Can

- arXiv
- 2610.01428
- Published
- 2026-10-01
- Authors
- Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi
AI summary
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMsOverview
Research area: Natural Language Processing / LLM evaluation and robustness (submitted to the TAE — Trust-AI-Eval: "Can We Trust AI Evaluation?" workshop).
Technical level: Intermediate — the paper assumes familiarity with standard LLM evaluation tooling (BLEU, BERTScore, token log-probabilities, hidden states, predictive entropy) but explains its framework from first principles.
Scope: The paper proposes SAGO (Stability-Aware Generalization Objective), a framework that redefines LLM generalization as per-instance behavioral stability under semantically equivalent prompt variations, and applies it to eleven instruction-tuned models across six datasets and four behavioral axes.
What This Paper Is About
Most LLM evaluations reduce generalization to an aggregate accuracy score on one prompt format, one task, or one kind of perturbation. That conflates "the model is good on this benchmark" with "the model behaves the same way when the same request is phrased differently." The paper argues this lets models look robust while their responses, confidence, internal representations, or stylistic mirroring shift substantially on individual examples — failures that cancel out in the average. SAGO instead measures, per prompt, how much model behavior changes across families of meaning-preserving rewrites, across multiple datasets and multiple behavioral dimensions.
Key Contributions
-
A per-instance stability framework. SAGO evaluates LLM generalization as the change in behavior for the same input under semantically equivalent variations, across multiple datasets and multiple families of variation — rather than as aggregate benchmark performance.
-
A multi-axis metric. The framework scores behavioral change along four axes instantiated as six deltas: activation geometry (Δ-Cos), generation consistency (Δ-BLEU, Δ-BERT), confidence and uncertainty (Δ-Prob, Δ-Ent), and response mirroring (Δ-MR). Mirroring is treated as a behavioral sensitivity that may reflect either intended alignment or unintended instability.
-
A stability score with a statistical test. The Stability Generalization Score (SGS) normalizes each axis to a common scale, aggregates per-prompt root-mean-square deltas, and tests against a non-zero tolerance threshold ε using a one-sided t-test at ε ∈ {0.01, 0.05, 0.10} (primary results at ε = 0.05).
-
An empirical study of eleven LLMs across six datasets, showing that no model generalizes uniformly, that generation consistency is the most sensitive axis, and that content stability and response mirroring are independent dimensions of behavioral sensitivity.
Main Findings
-
No model generalizes uniformly. Every open-source model exceeds ε = 0.05 on at least one axis, and Δ-BLEU and Δ-MR are statistically significant at p < 0.001 for every model. This held across three architecturally distinct families (Llama, Gemma, Qwen), which the authors read as a structural property of current instruction tuning rather than a fixable gap in any single model.
-
Scale does not resolve the problem. L-8B, G-7B, and Q-7B are equally or more unstable than their smaller counterparts on most axes, ruling out the hypothesis that the failure mode disappears with parameter count.
-
Token-level entropy is decoupled from output stability. Δ-Ent stays within the acceptable threshold across all models (maximum 0.012), meaning distributional sharpness does not track output-level consistency.
-
Axes capture independent failure modes. G4-E4B records the highest generation instability in the open-source pool (Δ-BERT 0.178; Δ-BLEU 0.837) while its predecessors score substantially lower (G-2B 0.097; G-7B 0.084) — a newer model is less generation-stable than its predecessors. Q3.5-9B shows the opposite inversion: lowest Δ-BERT (0.080) and Δ-BLEU (0.520) in the open-source pool, but the highest confidence instability (Δ-Prob 0.088) and substantial mirroring instability (Δ-MR 0.441). Aggregating across axes would give these two models similar robustness scores while masking opposite failure profiles.
-
Cross-dataset variation can reverse rankings. For L-8B, Δ-Cos spans a 3.7× range and Δ-MR more than doubles across the six datasets. Because this variation is large enough to flip model rankings depending on dataset choice, an evaluation anchored to one benchmark characterizes only that domain.
-
Closed-source models do not eliminate instability, they shift it. Total SGS scores: GPT-5.4 (Δ-BLEU 0.414 ± 0.13; Δ-BERT 0.056 ± 0.02; Δ-MR 0.020 ± 0.08), Gem-2.5F (0.312 ± 0.28; 0.105 ± 0.02; 0.005 ± 0.03), Cl-S-4.6 (0.505 ± 0.13; 0.120 ± 0.09; 0.061 ± 0.11). Activation geometry (Δ-Cos) and sequence log-probability (Δ-Prob) were not available for closed-source models due to API limitations, and Δ-Ent was computed only for GPT-5.4.
-
Stability is a structured profile, not a scalar. A model stable on one axis may be unstable on another; stability under one dataset or variation family does not imply stability elsewhere.
Methodology in Plain English
The researchers start from a simple definition of generalization: if two prompts mean the same thing, the model should behave the same way on both. They build a benchmark by taking factual QA prompts and generating meaning-preserving rewrites in three families:
- Social register — politeness, directness, greetings, hedging, gratitude, and negative framing, varied across eleven levels.
- Surface noise — irregular spacing, inserted punctuation, and letter-casing changes at multiple strengths.
- Structural rewriting — compressing or expanding prompt length and converting between interrogative and imperative forms.
Non-structural variants are applied at prefix, suffix, or global positions; structural variants are global by construction. To keep the rewrites truly meaning-preserving, the authors retain only prompts where all variants exceed BERTScore > 0.85 (surface-noise variants are meaning-preserving by construction).
For each original prompt and each variant, the model generates a response. The authors then compute a delta on each behavioral axis — how much activation geometry, generation similarity, confidence, entropy, or mirroring changed relative to the baseline. Deltas are normalized by a fixed per-axis scale factor (R_Act = 2, R_BLEU = 100, R_BERT = 1, R_Ent = 10, R_Prob = 250, R_MR = 1) so axes with different units can be compared while zero still means "no change." Per-prompt instability is the root mean square of these normalized deltas across the whole variation set; the SGS is the average of that across prompts, and the total SGS pools all six datasets with equal per-prompt weight. Since RMS is zero only when every delta is zero, a model that shifts by the same amount on every variation is still flagged as unstable.
Each axis score is then tested against a tolerance threshold ε with a one-sided t-test to check whether instability exceeds a level considered negligible for practical deployment.
Evaluation setup: eight open-source instruction-tuned models — Llama-3.2-3B-Instruct and Llama-3.1-8B-Instruct, Gemma-2B-IT, Gemma-7B-IT and Gemma-4-E4B-IT, Qwen2.5-1.5B-Instruct, Qwen2.5-7B-Instruct and Qwen3.5-9B — plus three closed-source models via API: GPT-5.4, Gemini-2.5-Flash, and Claude-Sonnet-4.6. Decoding is deterministic (T = 0). Datasets: TruthfulQA, Natural Questions, Alpaca, SimpleQA Verified, TriviaQA, and HotpotQA, with N = 16 prompts sampled per dataset. Experiments ran on four NVIDIA RTX 2080 Ti and four GTX 1080 Ti GPUs (CUDA 12.8). Social-register mirroring uses GPT-4o-mini as a judge (10% of labels manually verified), surface-noise mirroring uses rule-based heuristics, and length mirroring checks whether response length matches the target multiplier within ±0.25. BERTScore-F1 uses the RoBERTa-large backbone.
Why This Matters
Impact on research. The paper argues that claims about LLM generalization should be scoped to the axis, domain, and variation family evaluated, because robustness on one benchmark is not evidence of a broadly stable representation of user intent. It also recommends profile-based comparison over rank-based comparison — one model may be preferable when content stability matters most, another when confidence stability or reduced surface mirroring does. This challenges the practice of summarizing robustness as a single number.
Real-world applications:
- Search and retrieval systems — the same query phrased politely, tersely, or with typos should retrieve the same results; the paper cites search among core LLM use cases.
- Coding assistants — a request to fix a bug should yield equivalent behavior whether stated as a question or a command.
- Agentic workflows — agents that shift behavior under surface noise or prompt register may take different actions toward the same user goal.
- Safety and policy compliance — the paper stresses that an aligned model should still refuse harmful requests and follow its policies regardless of how the context is worded, toned, or formatted.
Industry relevance. Because instability appears in every model tested and does not shrink with parameter count, teams deploying or fine-tuning LLMs cannot assume that higher benchmark accuracy implies reliability for the same intent expressed differently. The independence of the axes matters practically: a single robustness score can hide the specific failure mode a deployment actually cares about, and content-stable models can still drift in confidence or imitate input noise.
Future Directions
- Multilingual and cross-cultural extension. The datasets are English-only; whether the same axis dissociations hold across languages, scripts, or culturally specific prompt conventions is untested.
- Causal mechanism analysis. SAGO measures where and how much behavior changes, not why. The authors propose mediation analysis or activation patching to ask whether generation inconsistency, confidence drift, and response mirroring share internal components.
- Larger and more adversarial sampling. The N = 16 operating point is validated for typical per-prompt instability, but rare or worst-case failures may require larger samples.
- Evaluating interventions by transfer. Future training interventions should be judged by whether gains transfer across axes, datasets, and semantically equivalent input forms — not by reduction on one axis or variation family.
- Reducing judge dependence. Structural variants and social-register mirroring rely on GPT-4o-mini, which may introduce stylistic bias, and limited API access restricts closed-source conclusions to the observable axes.
Target Audience
Researchers in LLM evaluation, robustness, and alignment; practitioners who benchmark or deploy instruction-tuned models and need to know what a benchmark score does and does not guarantee; and evaluation designers looking for a per-instance, multi-axis alternative to aggregate accuracy. Readers without prior exposure to prompt-sensitivity research will find the framework accessible, but familiarity with metrics such as BLEU, BERTScore, and token-level log-probabilities helps in interpreting the axis definitions.
Authors’ abstract
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.