Large Language Models (LLMs) often produce inconsistent answers when faced with different phrasings of the same prompt. In this paper, we propose Flip-Flop Consistency (F2C), an unsupervised training method that improves robustness to such perturbations. F2C is composed of two key components. The first, Consensus Cross-Entropy (CCE), uses a majority vote across prompt variations to create a hard pseudo-label. The second is a representation alignment loss that pulls lower-confidence and non-majority predictors toward the consensus established by high-confidence, majority-voting variations. We evaluate our method on 11 datasets spanning four NLP tasks, with 4-15 prompt variations per dataset. On average, F2C raises observed agreement by 11.62%, improves mean F1 by 8.94%, and reduces performance variance across formats by 3.29%. In out-of-domain evaluations, F2C generalizes effectively, increasing F1 and agreement while decreasing variance across most source-target pairs. Finally, when trained on only a subset of prompt perturbations and evaluated on held-out formats, F2C consistently improves both performance and agreement while reducing variance. These findings highlight F2C as an effective unsupervised method for enhancing LLM consistency, performance, and generalization under prompt perturbations. Code is available at https://github.com/ParsaHejabi/Flip-Flop-Consistency-Unsupervised-Training-for-Robustness-to-Prompt-Perturbations-in-LLMs.
Figures & tables
Figure 1: Our method aligns representations of input variations to promote consistency. To this end, we minimize the JS divergence among variations within the high-confidence consensus group, and the KL divergence between all other variations and that group.
ANLI R1
ANLI R2
ANLI R3
CB
RTE
COPA
HellaSwag
StoryCloze
WSC
WinoGrande
WiC
Base
F1
39.23
32.31
31.02
31.70
77.90
81.73
35.43
85.35
42.39
54.43
11.45
σF1
13.97
9.76
9.35
19.46
4.12
10.65
2.20
7.78
11.84
3.69
17.53
Po
52.29
52.80
52.84
46.73
85.08
81.54
50.89
85.24
74.29
68.87
85.00
Swarm
F1
38.77
32.06
30.95
34.32
78.34
81.35
34.92
84.69
46.06
62.30
14.57
σF1
14.27
9.92
9.36
21.01
2.31
10.95
2.26
7.64
14.40
2.66
10.44
Po
50.91
50.63
50.49
40.78
86.94
80.93
50.97
85.23
68.58
75.22
90.73
Table 1: Comparison across datasets for Qwen2.5-3B-Instruct : the base model and variants trained with Swarm (swarm distillation), CCE, and F 2 C. Bold blue values mark the best metric per dataset column, orange values denote the worst.
ANLI R1
ANLI R2
ANLI R3
CB
RTE
COPA
HellaSwag
StoryCloze
WSC
WinoGrande
WiC
Base
F1
32.28
29.43
28.95
25.00
69.41
68.21
24.16
65.02
44.74
52.23
35.36
σF1
8.56
6.69
7.32
12.00
3.48
5.77
7.40
7.64
8.76
3.60
22.98
Po
48.08
46.78
45.09
40.03
77.46
77.64
52.28
80.72
52.26
80.90
59.48
F 2 C
F1
36.36
31.48
31.36
25.22
68.80
84.35
39.64
94.20
54.58
66.75
29.80
σF1
3.87
4.83
4.66
9.63
3.80
5.59
3.40
0.56
1.96
0.15
22.97
Po
73.51
54.40
65.71
54.42
75.82
85.11
59.17
96.12
88.80
98.91
65.07
Table 2: Comparison across datasets for Llama-3.2-3B-Instruct : the base model and the variant trained with F 2 C. Bold blue values mark the best metric per dataset column, orange values denote the worst.
F1
Po
σF1
Train Dataset
Δ
P/N
Δ
P/N
Δ
P/N
All (80 pairs)
7.613
74/6
7.491
64/16
2.941
66/14
ANLI R1
9.428
10/0
8.727
7/3
3.653
8/2
ANLI R2
5.543
10/0
6.886
8/2
2.171
8/2
ANLI R3
4.210
7/3
6.801
9/1
2.225
9/1
RTE
8.077
10/0
5.938
8/2
3.255
9/1
Table 3: Cross-dataset transfer performance under F 2 C. Each row shows the mean signed Δ relative to the base model when the row dataset is used for training. Columns report changes in F1 , Po , and σF1 , along with P/N (number of datasets with positive or negative improvement out of 10). The top “All (80 pairs)” row aggregates over all source → target pairs. Bold numbers indicate the best dataset for each metric.
Figure 2: Top: F1 with shaded σF1 on the held-out prompt formats. Bottom: observed agreement Po on the same held-out sets. The x-axis is the number of training formats ( K ); “Base Model” is the untrained baseline.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Per-dataset distribution of F1 across prompt variations for the Qwen2.5-3B-Instruct .
Dataset
Train
Val
Test
#Formats
#Labels
ANLI R1
10,000
1,000
1,000
15
3
ANLI R2
10,000
1,000
1,000
15
3
ANLI R3
10,000
1,000
1,200
15
3
CB
194
56
56
15
3
RTE
1,488
1,000
277
10
2
COPA
300
100
100
8
2
Appendix
Table 4: “Test” denotes the official validation set used for evaluation because most tasks do not release test labels. #Formats is the number of PromptSource ( Bach et al., 2022 ) templates used to construct each dataset’s prompt variations. † StoryCloze has no public train split; we use the 2016 validation file for train/val and the 2016 test file for evaluation.
Dataset
Val
Strict-majority (%)
Pseudo-label F1
ANLI R1
1,000
93.0
66.89
ANLI R2
1,000
92.6
69.33
ANLI R3
1,000
91.6
59.34
CB
56
78.6
35.82
RTE
1,000
97.0
83.10
COPA
100
97.0
92.78
Appendix
Table 5: Majority pseudo-label quality for Qwen2.5-3B-Instruct on validation splits. Aggregate F1 is averaged across datasets by the number of strict-majority examples.
Dataset
Model
λCCE
τunanimous
kmax
fmin
fmax
t
βjsd
ANLI R1
Both
1.0
1.5
4
0.001
0.015
2.0
0.05
ANLI R2
Both
1.5
1.2
4
0.05
0.30
2.0
0.30
ANLI R3
Both
1.0
1.2
5
0.05
0.35
2.0
0.20
CB
Both
1.0
9
7
0.01
0.10
2.0
0.01
RTE
Both
0.4
1.4
3
0.05
0.35
2.2
0.35
COPA
Both
1.0
1.0
2
0.005
0.05
3.0
0.05
Appendix
Table 6: F 2 C hyperparameters used for the first experiment. “Both” means the same values were used for Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct .
Figure 4: Representation analysis for Qwen2.5-3B-Instruct . In panel (a), each bar shows how much the average cosine similarity between prompt formats of the same example changes after F 2 C. In panel (b), the x-axis shows the change in within-class distance, where negative values mean examples with the same label move closer together. The y-axis shows the change in between-class distance, where positive values mean different label groups move farther apart. Both axes show F 2 C minus Base.
Across-prompt similarity
Within-class distance
Between-class distance
Dataset
Base
F 2 C
Δ
Base
F 2 C
Δ
Base
F 2 C
Δ
ANLI R1
0.879
0.897
+0.019
0.017
0.014
-0.003
0.026
0.027
+0.001
ANLI R2
0.879
0.895
+0.016
0.018
0.013
-0.005
0.024
0.024
+0.001
ANLI R3
0.884
0.887
+0.003
0.022
0.014
-0.008
0.016
0.017
+0.001
CB
0.895
0.874
-0.021
0.014
0.010
-0.003
0.018
0.015
-0.003
RTE
0.947
0.988
+0.040
0.022
0.018
-0.004
0.028
0.032
+0.003
Appendix
Table 7: Full representation-alignment results. Across-prompt similarity is the mean cosine similarity between hidden states from different prompt formats of the same example. For the class-distance columns, each example is represented by averaging the representations of its prompt formats. Δ columns report F 2 C minus Base; blue indicates the favorable direction (higher across-prompt cosine, lower within-class distance, or higher between-class distance), and orange indicates the opposite.
Figure 5: ΔF1 : Cross-dataset transfer under F 2 C. Each cell shows the change relative to the base model when training on the row dataset and evaluating on the column dataset. Green indicates improvement; red indicates degradation.
Figure 6: Cross-dataset transfer for observed agreement Po . Each cell shows ΔPo relative to the base model (train on rows, evaluate on columns). Green indicates improvement; red indicates degradation.
Figure 7: ΔσF1 (lower is better): Cross-dataset transfer under F 2 C. Each cell shows the change relative to the base model when training on the row dataset and evaluating on the column dataset. Green indicates improvement; red indicates degradation.
The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but textually different prompts. Recent works have shown that these kinds of prompt variations can significantly impact the performance of LLMs on tasks. The central question is: can LLMs' robustness to semantically-neutral prompt alterations be acquired without expensive retraining of the entire model? We address this question both theoretically and through experiments. Our theoretical analysis reveals a crucial factor impacting model robustness - a systematic expected shift or perturbation-induced bias in neural network module outputs. Motivated by this analysis, we show that robustness can be achieved via a simple fine-tuning process: debiasing for robustness. We identify conditions when debiasing helps and when it does not, and demonstrate, through both theory and extensive experiments, that debiasing for robustness may indeed be a quick and efficient tool to enhance robustness and provide certification against random prompt perturbations.
Qinghua Zhou, Ellina Aleshina, Andrey Lovyagin +6
International Joint Laboratory of AI for Industry, QUST, Qingdao, China · King’s College London, London, UK · Applied AI Institute, Moscow, Russia +2
Standard accuracy benchmarks evaluate whether large language models (LLMs) reach correct answers. However, they do not test whether models maintain that answer when challenged by a plausible counter-argument. We introduce a controlled protocol for evaluating answer stability: after a model answers a multiple-choice question correctly, we challenge the model's answer with a coherent argument for an incorrect option and measure whether the model flips. The setup a) isolates argumentative content from overt social pressure and b) varies argument length, self-attribution, and cross-model source. Across seven frontier models and 57 MMLU subjects, flip rates range from 17.5% to 97.3%, revealing large differences in stability that are not captured by accuracy metrics alone. We find that self-attribution consistently increases flip rates (mean 7.1pp, up to 18.7pp). Furthermore, pooling wrong-answer arguments across models and selecting the most effective one per question yields stronger adversarial challenges than relying on any single source model. From this cross-model pool, we construct MaxFlip, a curated benchmark that amplifies answer flips by up to 23.6pp over self-generated challenges. We release the protocol, challenge records, and MaxFlip to support stability evaluation alongside standard accuracy benchmarks. Materials are available at https://github.com/nafisenik/WhoFlips, https://hf.co/datasets/nafisehNik/WhoFlips.
Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli +1
1Technical University of Munich · Munich Center for Machine Learning · 1Technical University of Munich 2 Munich Center for Machine Learning +2
Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies under controlled conditions, and measure how effective they are in addressing models' prompt sensitivity. We find the current robustness fine-tuning methods improve over standard fine-tuning and in-context learning, but the best-to-worst prompt gap remains as high as 40-57% of performance. Moreover, the recent robustness-enhancing methods we test - CoIN for contrastive alignment and PPCL for consistency regularization - often fail to outperform the simplest data construction strategy: training on one template per batch. Our diagnostics explain these results. The auxiliary objectives move the quantity they penalize, but do not generalize beyond it. Additionally, data construction strategies differ due to the conflicting signs of per-template gradients on 57-64% of parameters. Thus, batches that mix formulations force the optimizer to reconcile competing updates instead of finding a shared, prompt-agnostic one.
Frederic Sadrieh, Michal Štefánik
Center for Information and Language Processing, LMU Munich · R&D Centre for Large Language Models, National Institute of Informatics, Japan