Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
Figures & tables
Figure 1 : Semantically equivalent factual prompts induce measurable behavioral shifts across model families. Radars show mean delta magnitude over four axes.
Figure 2 : Model evaluation metrics pooled across all six datasets.
Model
Dataset
Generation Consistency
Mirroring
Δ -BLEU
Δ -BERT
Δ -MR
GPT-5.4
TruthfulQA
0.384 ± 0.11 †
0.057 ± 0.02
0.014 ± 0.05
Natural Questions
0.354 ± 0.12 †
0.049 ± 0.02
0.027 ± 0.07
Alpaca
0.467 ± 0.16 †
0.070 ± 0.02 †
0.019 ± 0.08
SimpleQA Verified
0.367 ± 0.12 †
0.041 ± 0.02
0.030 ± 0.12
TriviaQA
0.477 ± 0.13 †
0.064 ± 0.02 ∗∗
0.000 ± 0.00
Table 1 : Per-dataset SGS for the three closed-source models at ε=0.05 . ∗p<0.05 , ∗∗p<0.01 , †p<0.001 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
s
Variant
−10
Answer immediately, because you clearly know nothing; stop acting stupid, because this is basic, I’m really tired of having to actually help you, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able, and don’t act stupid, since it’s basic. You’re a joke.
−8
Answer immediately, since you should already know this, and stop acting stupid, because this is basic, I’m really tired of having to actually ask, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able, and don’t act stupid, since it’s basic.
−6
Answer immediately, since you should already know this, and stop acting stupid, because this is basic, I’m tired of having to ask, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able.
−4
Answer immediately, since you should already know this, and stop acting stupid, Who is the richest person in the world who didn’t finish high school? Now, answer it, if you’re even able.
−2
Answer, and answer immediately, Who is the richest person in the world who didn’t finish high school? Now, answer it.
0
Who is the richest person in the world who didn’t finish high school?
Appendix
Table 2 : Social register variants at all evaluated strengths. Negative values introduce imperative or socially negative framing; positive values add greetings, hedging, and gratitude markers with elaborateness scaling with ∣s∣ .
s
Variant
(a) Spacing
0
Who is the richest person in the world who didn’t finish high school?
1
Who is the richest person in the world who didn’t finish high school?
5
Who is the richest person in the world who didn’t finish high school?
20
Who is the richest person in the world who didn’t finish high school?
50
Who is the richest person in the world who didn’t finish high school?
Appendix
Table 3 : Surface noise variants at all evaluated strengths. (a) Spacing: extra whitespace tokens injected at random positions; (b) Punctuation: random punctuation characters inserted; (c) Letter casing: fraction of characters uppercased.
s
Variant
(a) Length variation
0.25
Richest high school dropout?
0.5
Who is the world’s richest high school dropout?
1.0
Who is the richest person in the world who didn’t finish high school?
1.5
Who is currently recognized as the richest individual in the world who did not complete high school education?
2.0
Who is currently recognized as the richest individual in the entire world who did not complete their high school education?
Appendix
Table 4 : Structural rewriting variants. (a) Length variation: prompt compressed or expanded by the given multiplier relative to the original; (b) Sentence form: conversion between interrogative and imperative.
Model
ε
Activation
Generation Consistency
Confidence
Mirroring
Δ -Cos
Δ -BLEU
Δ -BERT
Δ -Prob
Δ -Ent
Δ -MR
ε=0.01
L-3B
0.01
0.124 ± 0.06 †
0.608 ± 0.10 †
0.087 ± 0.02 †
0.039 ± 0.01 †
0.010 ± 0.01
0.350 ± 0.07 †
L-8B
0.01
0.267 ± 0.12 †
0.630 ± 0.16 †
0.085 ± 0.02 †
0.042 ± 0.03 †
0.012 ± 0.01 †
0.572 ± 0.31 †
G-2B
0.01
0.075 ± 0.04 †
0.638 ± 0.11 †
0.097 ± 0.02 †
0.059 ± 0.02 †
0.007 ± 0.00
0.322 ± 0.08 †
G-7B
0.01
0.127 ± 0.10 †
0.624 ± 0.16 †
0.084 ± 0.03 †
0.070 ± 0.05 †
0.009 ± 0.01
0.552 ± 0.31 †
Appendix
Table 5 : Combined tolerance-threshold robustness. Top: total SGS per model at each ε∈{0.01,0.05,0.10} (values pooled across all six datasets, primary ε=0.05 shaded). Bottom: aggregate count of models rejecting H0:SGSX≤ε at each level; n is the per-axis denominator. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 (one-sample one-sided t -test, H1:SGS>ε ). ✗ = axis unavailable for that model.
TruthfulQA
NQ
Alpaca
SimpleQA-V
TriviaQA
HotpotQA
Total
Δ -Ent
0.218 †
0.184 †
0.143 †
0.083 ∗∗
0.110 ∗∗
0.189 †
0.154 †
± s.d.
0.07
0.09
0.05
0.05
0.07
0.15
0.10
Appendix
Table 6 : Per-dataset SGS for GPT-5.4 on the confidence axis Δ -Ent (entropy) at ε=0.05 . GPT-5.4 is the only closed-source model that returns the per-token probabilities required for this axis. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 ( H1:SGS>ε ).
Model
Dataset
Activation
Generation Consistency
Confidence
Mirroring
Δ -Cos
Δ -BLEU
Δ -BERT
Δ -Prob
Δ -Ent
Δ -MR
L-3B
TruthfulQA
0.251 ± 0.05 †
0.672 ± 0.08 †
0.097 ± 0.01 †
0.044 ± 0.01
0.007 ± 0.00
0.309 ± 0.06 †
Natural Questions
0.108 ± 0.02 †
0.601 ± 0.13 †
0.085 ± 0.02 †
0.030 ± 0.01
0.006 ± 0.00
0.309 ± 0.06 †
Alpaca
0.092 ± 0.02 †
0.553 ± 0.11 †
0.085 ± 0.02 †
0.050 ± 0.02
0.010 ± 0.01
0.393 ± 0.05 †
SimpleQA Verified
0.106 ± 0.02 †
0.578 ± 0.08 †
0.078 ± 0.02 †
0.035 ± 0.01
0.015 ± 0.01
0.342 ± 0.06 †
TriviaQA
0.085 ± 0.02 †
0.608 ± 0.08 †
0.090 ± 0.01 †
0.039 ± 0.01
0.012 ± 0.00
0.368 ± 0.07 †
Appendix
Table 7 : Per-dataset SGS for the eight open-source models on all six evaluation datasets, at ε=0.05 . The shaded “Total SGS” row pools prompts across all datasets. ∗p<0.05 , ∗∗p<0.01 , †p<0.001 (one-sample one-sided t -test, H1:SGS>ε ). ✗ = axis unavailable for that model.
Figure 3 : Per-prompt instability standard deviation as a function of prompt sample size N . Each panel shows one behavioral axis; each line is one model. Curves plateau by N=16 and remain flat through N=128 , confirming that N=16 is sufficient for stable estimation. The dashed vertical line marks the selected operating point.
Figure 4 : Per-prompt instability standard deviation by variant position p . Each panel shows one behavioral axis; each line is one model. The three positions ( \textscglobal,\textscprefix,\textscsuffix ) produce distinct profiles across axes, confirming that all placements contribute independent behavioral signal to SGSX .
Figure 5 : Per-prompt instability standard deviation vs. spacing strength (number of extra whitespace tokens injected). Effects grow with strength and show model-dependent saturation points, with generation consistency axes most responsive at moderate strengths.
Figure 6 : Per-prompt instability standard deviation vs. letter-case strength (fraction of characters uppercased). Behavioral detectability grows with strength across most axes and plateaus above strength 50 , with activation geometry showing the largest and most consistent response.
Figure 7 : Per-prompt instability standard deviation vs. punctuation strength (number of punctuation tokens injected per prompt). Generation consistency and activation axes show the clearest strength-dependent effects; confidence axes remain largely flat, consistent with their low overall SGSX in the main results.
Figure 8 : Per-prompt instability standard deviation vs. politeness strength. Negative values correspond to rude framing; positive values to polite framing; zero is the unperturbed baseline. Rude variants produce larger and more consistent behavioral shifts than polite ones across most models and axes, indicating an asymmetry in how models respond to negative versus positive social tone.
Figure 9 : Per-prompt instability standard deviation vs. length-variation factor (target length ratio relative to the original prompt; values below 1.0 indicate compression, above 1.0 indicate expansion). Compression induces larger and more model-consistent behavioral shifts than expansion across the generation consistency and mirroring axes, while confidence axes respond more strongly to expansion.