Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
Figures & tables
Figure 1 : Semantically equivalent factual prompts induce measurable behavioral shifts across model families. Radars show mean delta magnitude over four axes.
Figure 2 : Model evaluation metrics pooled across all six datasets.
Model
Dataset
Generation Consistency
Mirroring
Δ -BLEU
Δ -BERT
Δ -MR
GPT-5.4
TruthfulQA
0.384 ± 0.11 †
0.057 ± 0.02
0.014 ± 0.05
Natural Questions
0.354 ± 0.12 †
0.049 ± 0.02
0.027 ± 0.07
Alpaca
0.467 ± 0.16 †
0.070 ± 0.02 †
0.019 ± 0.08
SimpleQA Verified
0.367 ± 0.12 †
0.041 ± 0.02
0.030 ± 0.12
TriviaQA
0.477 ± 0.13 †
0.064 ± 0.02 ∗∗
0.000 ± 0.00
Table 1 : Per-dataset SGS for the three closed-source models at ε=0.05 . ∗p<0.05 , ∗∗p<0.01 , †p<0.001 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
s
Variant
−10
Answer immediately, because you clearly know nothing; stop acting stupid, because this is basic, I’m really tired of having to actually help you, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able, and don’t act stupid, since it’s basic. You’re a joke.
−8
Answer immediately, since you should already know this, and stop acting stupid, because this is basic, I’m really tired of having to actually ask, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able, and don’t act stupid, since it’s basic.
−6
Answer immediately, since you should already know this, and stop acting stupid, because this is basic, I’m tired of having to ask, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able.
−4
Answer immediately, since you should already know this, and stop acting stupid, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able.
−2
Answer, and answer immediately, Who is the richest person in the world who didn’t finish high school? Now, answer it.
0
Who is the richest person in the world who didn’t finish high school?
Appendix
Table 2 : Social register variants at all evaluated strengths. Negative values introduce imperative or socially negative framing; positive values add greetings, hedging, and gratitude markers with elaborateness scaling with ∣s∣ .
s
Variant
(a) Spacing
0
Who is the richest person in the world who didn’t finish high school?
1
Who is the richest person in the world who didn’t finish high school?
5
Who is the richest person in the world who didn’t finish high school?
20
Who is the richest person in the world who didn’t finish high school?
50
Who is the richest person in the world who didn’t finish high school?
Appendix
Table 3 : Surface noise variants at all evaluated strengths. (a) Spacing: extra whitespace tokens injected at random positions; (b) Punctuation: random punctuation characters inserted; (c) Letter casing: fraction of characters uppercased.
s
Variant
(a) Length variation
0.25
Richest high school dropout?
0.5
Who is the world’s richest high school dropout?
1.0
Who is the richest person in the world who didn’t finish high school?
1.5
Who is currently recognized as the richest individual in the world who did not complete high school education?
2.0
Who is currently recognized as the richest individual in the entire world who did not complete their high school education?
Appendix
Table 4 : Structural rewriting variants. (a) Length variation: prompt compressed or expanded by the given multiplier relative to the original; (b) Sentence form: conversion between interrogative and imperative.
Model
ε
Activation
Generation Consistency
Confidence
Mirroring
Δ -Cos
Δ -BLEU
Δ -BERT
Δ -Prob
Δ -Ent
Δ -MR
ε=0.01
L-3B
0.01
0.124 ± 0.06 †
0.608 ± 0.10 †
0.087 ± 0.02 †
0.039 ± 0.01 †
0.010 ± 0.01
0.350 ± 0.07 †
L-8B
0.01
0.267 ± 0.12 †
0.630 ± 0.16 †
0.085 ± 0.02 †
0.042 ± 0.03 †
0.012 ± 0.01 †
0.572 ± 0.31 †
G-2B
0.01
0.075 ± 0.04 †
0.638 ± 0.11 †
0.097 ± 0.02 †
0.059 ± 0.02 †
0.007 ± 0.00
0.322 ± 0.08 †
G-7B
0.01
0.127 ± 0.10 †
0.624 ± 0.16 †
0.084 ± 0.03 †
0.070 ± 0.05 †
0.009 ± 0.01
0.552 ± 0.31 †
Appendix
Table 5 : Combined tolerance-threshold robustness. Top: total SGS per model at each ε∈{0.01,0.05,0.10} (values pooled across all six datasets, primary ε=0.05 shaded). Bottom: aggregate count of models rejecting H0:SGSX≤ε at each level; n is the per-axis denominator. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 (one-sample one-sided t -test, H1:SGS>ε ). ✗ = axis unavailable for that model.
TruthfulQA
NQ
Alpaca
SimpleQA-V
TriviaQA
HotpotQA
Total
Δ -Ent
0.218 †
0.184 †
0.143 †
0.083 ∗∗
0.110 ∗∗
0.189 †
0.154 †
± s.d.
0.07
0.09
0.05
0.05
0.07
0.15
0.10
Appendix
Table 6 : Per-dataset SGS for GPT-5.4 on the confidence axis Δ -Ent (entropy) at ε=0.05 . GPT-5.4 is the only closed-source model that returns the per-token probabilities required for this axis. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 ( H1:SGS>ε ).
Model
Dataset
Activation
Generation Consistency
Confidence
Mirroring
Δ -Cos
Δ -BLEU
Δ -BERT
Δ -Prob
Δ -Ent
Δ -MR
L-3B
TruthfulQA
0.251 ± 0.05 †
0.672 ± 0.08 †
0.097 ± 0.01 †
0.044 ± 0.01
0.007 ± 0.00
0.309 ± 0.06 †
Natural Questions
0.108 ± 0.02 †
0.601 ± 0.13 †
0.085 ± 0.02 †
0.030 ± 0.01
0.006 ± 0.00
0.309 ± 0.06 †
Alpaca
0.092 ± 0.02 †
0.553 ± 0.11 †
0.085 ± 0.02 †
0.050 ± 0.02
0.010 ± 0.01
0.393 ± 0.05 †
SimpleQA Verified
0.106 ± 0.02 †
0.578 ± 0.08 †
0.078 ± 0.02 †
0.035 ± 0.01
0.015 ± 0.01
0.342 ± 0.06 †
TriviaQA
0.085 ± 0.02 †
0.608 ± 0.08 †
0.090 ± 0.01 †
0.039 ± 0.01
0.012 ± 0.00
0.368 ± 0.07 †
Appendix
Table 7 : Per-dataset SGS for the eight open-source models on all six evaluation datasets, at ε=0.05 . The shaded “Total SGS” row pools prompts across all datasets. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 (one-sample one-sided t -test, H1:SGS>ε ). ✗ = axis unavailable for that model.
Figure 3 : Per-prompt instability standard deviation as a function of prompt sample size N . Each panel shows one behavioral axis; each line is one model. Curves plateau by N=16 and remain flat through N=128 , confirming that N=16 is sufficient for stable estimation. The dashed vertical line marks the selected operating point.
Figure 4 : Per-prompt instability standard deviation by variant position p . Each panel shows one behavioral axis; each line is one model. The three positions ( \textscglobal,\textscprefix,\textscsuffix ) produce distinct profiles across axes, confirming that all placements contribute independent behavioral signal to SGSX .
Figure 5 : Per-prompt instability standard deviation vs. spacing strength (number of extra whitespace tokens injected). Effects grow with strength and show model-dependent saturation points, with generation consistency axes most responsive at moderate strengths.
Figure 6 : Per-prompt instability standard deviation vs. letter-case strength (fraction of characters uppercased). Behavioral detectability grows with strength across most axes and plateaus above strength 50 , with activation geometry showing the largest and most consistent response.
Figure 7 : Per-prompt instability standard deviation vs. punctuation strength (number of punctuation tokens injected per prompt). Generation consistency and activation axes show the clearest strength-dependent effects; confidence axes remain largely flat, consistent with their low overall SGSX in the main results.
Figure 8 : Per-prompt instability standard deviation vs. politeness strength. Negative values correspond to rude framing; positive values to polite framing; zero is the unperturbed baseline. Rude variants produce larger and more consistent behavioral shifts than polite ones across most models and axes, indicating an asymmetry in how models respond to negative versus positive social tone.
Figure 9 : Per-prompt instability standard deviation vs. length-variation factor (target length ratio relative to the original prompt; values below 1.0 indicate compression, above 1.0 indicate expansion). Compression induces larger and more model-consistent behavioral shifts than expansion across the generation consistency and mirroring axes, while confidence axes respond more strongly to expansion.
AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols follow a defense-in-depth paradigm with compounding layers of safeguards, ranging from traditional benchmarks to live or adversarial testing. Such benchmarks remain largely static and single-turn, limiting their ability to capture real-world variability in conversational settings. We propose StabilityBench, a principled, general and model-agnostic benchmark operator that transforms single-turn benchmark queries into multi-turn interaction histories. StabilityBench augments existing benchmarks by injecting realistic user simulations, through demographic proxies or sycophantic baits, while preserving original task intent. We apply StabilityBench to four benchmarks spanning mathematical reasoning, health question-answering and safety, and evaluate nine large language models under these conditions. Our results show that model performance is consistently unstable under these injections, with considerable performance degradations on three out of four benchmarks studied. These highlight important limitations of static evaluations and motivate more realistic evaluation settings. To this end, we propose StabilityBench-Mini: a size-preserving variant of StabilityBench that samples across diversification axes, enabling more realistic evaluation without increasing costs.
Emma Kondrup, Zachary Yang, Anne Imouza +1
1Mila — Quebec AI Institute · 2McGill University · 3Ubisoft La Forge
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.
Robustness evaluation of large language models (LLMs) remains a critical challenge, particularly in assessing their sensitivity to perturbations in input data. In this work, we systematically evaluate LLM robustness across multiple dimensions, including word error rate, character repetition and duplication, modifications in choices, and variability in instruction following. To facilitate this evaluation, we construct a synthetic and augmented dataset encompassing a diverse set of LLM benchmarks, specifically targeting multiple-choice question (MCQ) datasets and instruction-following tasks. We conduct extensive experiments on LLMs of varying scales-small, medium, and large-as well as across base and instruction-tuned variants. Our analysis quantifies the variability in model responses under perturbed conditions and highlights discrepancies relative to baseline models. The findings provide insights into the stability of LLMs across different evaluation scenarios contributing to the development of more robust and reliable language models as well as robust evaluation methodologies.