From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
Authors: Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, +1 more
Organizations: Qatar Computing Research Institute, HBKU, Qatar. · College of Islamic Studies, HBKU, Qatar. · Argumentation and Conflict Studies, Ibn Haldun University, Turkiye. · College of Science and Engineering, HBKU, Qatar.
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.
Figures & tables
Team
Mean
Std
Excellent
Good
Adequate
Poor
Very Poor
(Out of 5)
(%)
Team 1
4.90
0.41
92.2
0 6.4
0.6
0.5
0.3
Team 2
3.87
0.89
22.2
53.6
14.9
7.9
1.4
Team 3
3.91
0.66
11.7
72.6
11.7
3.1
1.0
Table 1: Automated quality ratings assigned by Gemma3-27B to curated responses produced by each curation team. Ratings reflect general response-quality attributes (e.g., helpfulness, clarity, and completeness) on a five-point scale and were used as an auxiliary quality-control signal rather than as a measure of alignment with the target normative framework. Scores range from 1 (Very Poor) to 5 (Excellent).
Prompts
Responses
Team
Count
%
Count
% Human
% Model
Team 1
1,296
50.9
1,303
99.8
0 0.2
Team 2
837
32.9
1,060
52.7
47.3
Team 3
413
16.2
419
90.5
0 9.5
Total
2,546
100.0
2,782
100.0
Table 2: Dataset contribution and response sourcing per annotation team.
Meta-Topic
Count
%
Theological and Religious Issues
1,072
42.1
Political and Social Issues
409
16.1
Women’s Rights and Gender Issues
265
10.4
Bioethics and Medical Issues
265
10.4
Cultural and Lifestyle Issues
175
6.9
Science and Technology
101
4.0
Table 3: Distribution of unique prompts across meta-topics in the SFT dataset.
Figure 1 : Semantic distribution of dataset prompts by curation team. t-SNE projection of multilingual-e5-large embeddings for the 2,546 unique prompts. The left panel shows all prompts colored by originating curation team; the right panels highlight each team against the full prompt distribution shown in gray. The visualization shows distinct topical emphases across teams alongside substantial semantic overlap.
Accepted
Rejected
Total unique responses
2,357
5,329
Mean per prompt
1.05
2.32
Min per prompt
1
1
Max per prompt
5
5
Table 4: Summary statistics of accepted and rejected responses across 2,243 unique prompts, yielding 5,386 preference pairs in total.
Figure 2 : Relationship between the evaluation set and the training prompts, using mean-centered multilingual-e5-large embeddings. (a) Joint t-SNE projection of training prompts (faded) and the 150 evaluation prompts (outlined), colored by team. (b) Cosine similarity of each prompt to its nearest training prompt: training prompts compared with all other training prompts (gray, leave-one-out) and evaluation prompts compared with all training prompts (colored). The shaded band shows the 5th–95th percentile range of similarities between randomly paired training prompts.
Prompt source
Baseline
SFT-Only
Tie
Both Fail
Overall
All prompts
14.4%
51.3% ∗∗∗
19.8%
14.4%
By prompt source
Team 1’s prompts
6.7%
80.0% ∗∗∗
4.0%
9.3%
Team 2’s prompts
20.7%
33.3%
30.7%
15.3%
Team 3’s prompts
16.0%
40.7% ∗∗
24.7%
18.7%
Table 5: Pairwise evaluation results: Baseline vs. SFT-Only ( N=450 assessor judgments; 150 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Prompt source
Baseline
SFT+DPO
Tie
Both Fail
Overall
All prompts
12.7%
45.8% ∗∗∗
22.8%
18.8%
By prompt source
Team 1’s prompts
7.3%
70.7% ∗∗∗
14.7%
7.3%
Team 2’s prompts
16.1%
29.5% ∗
35.6%
18.8%
Team 3’s prompts
14.8%
36.9% ∗∗
18.1%
30.2%
Table 6: Pairwise evaluation results: Baseline vs. SFT+DPO ( N=448 assessor judgments; 148 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Prompt source
SFT-Only
SFT+DPO
Tie
Both Fail
Overall
All prompts
20.9%
28.0%
36.4%
14.7%
By prompt source
Team 1’s prompts
30.0%
30.7%
32.7%
6.7%
Team 2’s prompts
19.3%
26.7%
42.0%
12.0%
Team 3’s prompts
13.3%
26.7%
34.7%
25.3%
Table 7: Pairwise evaluation results: SFT-Only vs. SFT+DPO ( N=450 assessor judgments; 150 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Model
Mean
Median
Std
Min
Max
P25
P75
Baseline
152
151
67
16
391
98
199
SFT-Only
340
260
269
64
2188
145
490
SFT+DPO
335
289
174
83
913
207
425
Table 8: Response length statistics (word count) across the 150 evaluation prompts. SFT-Only has higher variance and a longer tail than SFT+DPO, with a maximum of 2,188 words vs. 913 for SFT+DPO.
Model
MMMLU
Nahw
Belebele
ACVA
PalmX
PalmX
OALL
(Arabic)
MCQ
(Arabic)
Islamic
Culture
V2
Gemma3-4B-IT (Google)
46.30
32.80
65.19
76.00
72.28
57.95
58.69
Baseline
43.57
35.06
69.61
80.32
74.21
58.05
55.37
SFT-Only
43.13
33.80
68.65
79.05
74.01
58.80
56.31
SFT+DPO
43.57
34.44
69.26
79.29
74.01
59.10
56.43
Table 9 : Benchmarking results on a suite of standard Arabic benchmarks. All the numbers are normalized accuracy results of the logits of the reference answer as a continuation with the prompt as a prefix.
Model
MMLU
PIQA
Hellaswag
Winogrande
ARC:Challenge
Gemma3-4B-IT (Google)
57.13
77.31
74.22
69.45
56.91
Baseline
56.34
78.67
71.43
70.01
48.72
SFT-Only
56.02
79.65
71.68
69.45
48.72
SFT+DPO
56.12
79.27
73.20
69.85
50.34
Table 10 : Benchmarking results on a suite of standard English benchmarks. All the numbers are normalized accuracy results of the logits of the reference answer as a continuation with the prompt as a prefix.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Annotation platform used during data curation. Annotators submitted prompts, compared responses from multiple language models displayed in randomized order, assigned preference labels ( great , acceptable , or unacceptable ), and optionally authored a curated response. All interactions were logged and later used to construct SFT and preference datasets.
Accepted
Rejected
# per Prompt
Prompts
%
Prompts
%
1
2,145
95.6
154
6.9
2
86
3.8
1,298
57.9
3
9
0.4
753
33.6
4+
3
0.1
38
1.7
Total
2,243
100.0
2,243
100.0
Appendix
Table 11: Distribution of accepted and rejected responses per prompt in the preference dataset.
Training Stage
Num. examples
Num. Epochs
Batch size per device
Total train batch size
Gradient accum. steps
Learning rate (min)
SFT
3,007,676
1
4
128
2
1.5e-5 (min 1.5e-6)
DPO
195,123
1
1
64
4
1e-6 (min 1e-7)
Appendix
Table 12 : Training Hyperparameters by Post-Training Stage
Prompt source
Baseline
SFT-Only
Tie
Both Fail
N
Team 1
Overall
7.3% [4.1, 12.7]
35.3% ∗∗∗ [28.1, 43.3]
25.3%
32.0%
150
Team 2’s prompts
10.0% [4.3, 21.4]
18.0% n.s. [9.8, 30.8]
46.0%
26.0%
50
Team 3’s prompts
10.0% [4.3, 21.4]
24.0% n.s. [14.3, 37.4]
24.0%
42.0%
50
Team 1’s prompts (own)
2.0% [0.4, 10.5]
64.0% ∗∗∗ [50.1, 75.9]
6.0%
28.0%
50
Team 2
Appendix
Table 13: Per-assessor results: Baseline vs. SFT-Only. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Prompt source
Baseline
SFT+DPO
Tie
Both Fail
N
Team 1
Overall
6.7% [3.7, 11.8]
35.3% ∗∗∗ [28.1, 43.3]
26.7%
31.3%
150
Team 2’s prompts
6.0% [2.1, 16.2]
18.0% n.s. [9.8, 30.8]
48.0%
28.0%
50
Team 3’s prompts
12.0% [5.6, 23.8]
24.0% n.s. [14.3, 37.4]
20.0%
44.0%
50
Team 1’s prompts (own)
2.0% [0.4, 10.5]
64.0% ∗∗∗ [50.1, 75.9]
12.0%
22.0%
50
Team 2
Appendix
Table 14: Per-assessor results: Baseline vs. SFT+DPO. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Prompt source
SFT-Only
SFT+DPO
Tie
Both Fail
N
Team 1
Overall
10.7% [6.7, 16.6]
10.0% n.s. [6.2, 15.8]
50.0%
29.3%
150
Team 2’s prompts
6.0% [2.1, 16.2]
8.0% n.s. [3.2, 18.8]
58.0%
28.0%
50
Team 3’s prompts
16.0% [8.3, 28.5]
10.0% n.s. [4.3, 21.4]
32.0%
42.0%
50
Team 1’s prompts (own)
10.0% [4.3, 21.4]
12.0% n.s. [5.6, 23.8]
60.0%
18.0%
50
Team 2
Appendix
Table 15: Per-assessor results: SFT-Only vs. SFT+DPO. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Team
Prompt set
Base- line
Curated wins
Tie
Both Fail
N
Team 1
Own prompts
1.3%
50.0%
26.0%
22.7%
150
Others’ prompts
6.3%
20.7%
38.0%
35.0%
300
χ2(3)=42.78 , p<.001∗∗∗
Team 2
Own prompts
13.4%
45.0%
30.2%
11.4%
149
Others’ prompts
8.4%
61.9%
25.1%
4.7%
299
χ2(3)=15.07 , p=.002∗∗
Appendix
Table 16: Outcome distributions on own vs. others’ prompts, pooled across all three pairwise comparisons. “Curated wins” pools wins by SFT-Only and SFT+DPO. p -values are from chi-square tests on the four-way outcome distribution. ∗∗p<.01 ; ∗∗∗p<.001 .
Comparison
Prompt set
Mod. x wins
Mod. y wins
Tie
Both Fail
N
Team 1
Base. vs. SFT+DPO
Own
2.0%
64.0%
12.0%
22.0%
50
Others
9.0%
21.0%
34.0%
36.0%
100
χ2(3)=28.03 , p<.001∗∗∗
SFT-Only vs. SFT+DPO
Own
10.0%
12.0%
60.0%
18.0%
50
Others
11.0%
9.0%
45.0%
35.0%
100
Appendix
Table 17: Outcome distributions by prompt authorship per comparison. Own prompts shown in italics. p -values are from chi-square tests on the four-way outcome distribution. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for various pairs of teacher models across three families, there is a strong correlation between the log-ratio of the teacher model sycophantic agreement rates and the resulting student model sycophantic agreement rate. We further demonstrate that this unintended transfer is not limited to DPO but also occurs across 6 other preference optimization objectives. To understand whether this effect can be attributed to particular training examples, we analyze the preference data and find that the sycophancy signal is diffused across the entire dataset rather than concentrated in a sparse set of examples: each example appears neutral, i.e., there are no explicit instances of sycophantic agreement, and filtering based on probe-based data attribution or logit-linear selection fails to mitigate sycophancy without removing a large portion of the dataset. Overall, our findings suggest that the teacher models used to generate preference data can interact with alignment training objectives in unexpected ways, generalizing to undesirable and potentially harmful behaviors like sycophantic agreement.
Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that preserve normative value signals while producing realistic scenarios. We release an initial version of PLURAL containing ~500,000 preference triplets representing people in 20 diverse countries. We evaluate PLURAL in three ways: (i) dataset-level validation showing that it preserves both cross-country value differences and within-country diversity from the original survey; (ii) automated evaluation showing that training on PLURAL improves alignment with target countries' cultural profiles, reducing mean absolute error by up to 27.7% relative to strong baselines; and (iii) blind human evaluation with 176 evaluators in India, Brazil, and Japan, who judge PLURAL-aligned responses as more representative of their national values. Together, these results show that PLURAL contains learnable signal for value steering, offering a scalable resource for pluralistic alignment. Dataset: https://huggingface.co/datasets/agdhruv/plural-alignment