From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
Authors: Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, +1 more
Organizations: Qatar Computing Research Institute, HBKU, Qatar. · College of Islamic Studies, HBKU, Qatar. · Argumentation and Conflict Studies, Ibn Haldun University, Turkiye. · College of Science and Engineering, HBKU, Qatar.
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.
Figures & tables
Team
Mean
Std
Excellent
Good
Adequate
Poor
Very Poor
(Out of 5)
(%)
Team 1
4.90
0.41
92.2
0 6.4
0.6
0.5
0.3
Team 2
3.87
0.89
22.2
53.6
14.9
7.9
1.4
Team 3
3.91
0.66
11.7
72.6
11.7
3.1
1.0
Table 1: Automated quality ratings assigned by Gemma3-27B to curated responses produced by each curation team. Ratings reflect general response-quality attributes (e.g., helpfulness, clarity, and completeness) on a five-point scale and were used as an auxiliary quality-control signal rather than as a measure of alignment with the target normative framework. Scores range from 1 (Very Poor) to 5 (Excellent).
Prompts
Responses
Team
Count
%
Count
% Human
% Model
Team 1
1,296
50.9
1,303
99.8
0 0.2
Team 2
837
32.9
1,060
52.7
47.3
Team 3
413
16.2
419
90.5
0 9.5
Total
2,546
100.0
2,782
100.0
Table 2: Dataset contribution and response sourcing per annotation team.
Meta-Topic
Count
%
Theological and Religious Issues
1,072
42.1
Political and Social Issues
409
16.1
Women’s Rights and Gender Issues
265
10.4
Bioethics and Medical Issues
265
10.4
Cultural and Lifestyle Issues
175
6.9
Science and Technology
101
4.0
Table 3: Distribution of unique prompts across meta-topics in the SFT dataset.
Figure 1 : Semantic distribution of dataset prompts by curation team. t-SNE projection of multilingual-e5-large embeddings for the 2,546 unique prompts. The left panel shows all prompts colored by originating curation team; the right panels highlight each team against the full prompt distribution shown in gray. The visualization shows distinct topical emphases across teams alongside substantial semantic overlap.
Accepted
Rejected
Total unique responses
2,357
5,329
Mean per prompt
1.05
2.32
Min per prompt
1
1
Max per prompt
5
5
Table 4: Summary statistics of accepted and rejected responses across 2,243 unique prompts, yielding 5,386 preference pairs in total.
Figure 2 : Relationship between the evaluation set and the training prompts, using mean-centered multilingual-e5-large embeddings. (a) Joint t-SNE projection of training prompts (faded) and the 150 evaluation prompts (outlined), colored by team. (b) Cosine similarity of each prompt to its nearest training prompt: training prompts compared with all other training prompts (gray, leave-one-out) and evaluation prompts compared with all training prompts (colored). The shaded band shows the 5th–95th percentile range of similarities between randomly paired training prompts.
Prompt source
Baseline
SFT-Only
Tie
Both Fail
Overall
All prompts
14.4%
51.3% ∗∗∗
19.8%
14.4%
By prompt source
Team 1’s prompts
6.7%
80.0% ∗∗∗
4.0%
9.3%
Team 2’s prompts
20.7%
33.3%
30.7%
15.3%
Team 3’s prompts
16.0%
40.7% ∗∗
24.7%
18.7%
Table 5: Pairwise evaluation results: Baseline vs. SFT-Only ( N=450 assessor judgments; 150 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Prompt source
Baseline
SFT+DPO
Tie
Both Fail
Overall
All prompts
12.7%
45.8% ∗∗∗
22.8%
18.8%
By prompt source
Team 1’s prompts
7.3%
70.7% ∗∗∗
14.7%
7.3%
Team 2’s prompts
16.1%
29.5% ∗
35.6%
18.8%
Team 3’s prompts
14.8%
36.9% ∗∗
18.1%
30.2%
Table 6: Pairwise evaluation results: Baseline vs. SFT+DPO ( N=448 assessor judgments; 148 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Prompt source
SFT-Only
SFT+DPO
Tie
Both Fail
Overall
All prompts
20.9%
28.0%
36.4%
14.7%
By prompt source
Team 1’s prompts
30.0%
30.7%
32.7%
6.7%
Team 2’s prompts
19.3%
26.7%
42.0%
12.0%
Team 3’s prompts
13.3%
26.7%
34.7%
25.3%
Table 7: Pairwise evaluation results: SFT-Only vs. SFT+DPO ( N=450 assessor judgments; 150 prompts). Win rates are computed over all assessor judgments (Tie and Both Fail retained in denominator). p -values are from one-sided binomial tests at the prompt level, aggregating three assessor judgments per prompt by majority vote; prompts with no majority are excluded. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Model
Mean
Median
Std
Min
Max
P25
P75
Baseline
152
151
67
16
391
98
199
SFT-Only
340
260
269
64
2188
145
490
SFT+DPO
335
289
174
83
913
207
425
Table 8: Response length statistics (word count) across the 150 evaluation prompts. SFT-Only has higher variance and a longer tail than SFT+DPO, with a maximum of 2,188 words vs. 913 for SFT+DPO.
Model
MMMLU
Nahw
Belebele
ACVA
PalmX
PalmX
OALL
(Arabic)
MCQ
(Arabic)
Islamic
Culture
V2
Gemma3-4B-IT (Google)
46.30
32.80
65.19
76.00
72.28
57.95
58.69
Baseline
43.57
35.06
69.61
80.32
74.21
58.05
55.37
SFT-Only
43.13
33.80
68.65
79.05
74.01
58.80
56.31
SFT+DPO
43.57
34.44
69.26
79.29
74.01
59.10
56.43
Table 9 : Benchmarking results on a suite of standard Arabic benchmarks. All the numbers are normalized accuracy results of the logits of the reference answer as a continuation with the prompt as a prefix.
Model
MMLU
PIQA
Hellaswag
Winogrande
ARC:Challenge
Gemma3-4B-IT (Google)
57.13
77.31
74.22
69.45
56.91
Baseline
56.34
78.67
71.43
70.01
48.72
SFT-Only
56.02
79.65
71.68
69.45
48.72
SFT+DPO
56.12
79.27
73.20
69.85
50.34
Table 10 : Benchmarking results on a suite of standard English benchmarks. All the numbers are normalized accuracy results of the logits of the reference answer as a continuation with the prompt as a prefix.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Annotation platform used during data curation. Annotators submitted prompts, compared responses from multiple language models displayed in randomized order, assigned preference labels ( great , acceptable , or unacceptable ), and optionally authored a curated response. All interactions were logged and later used to construct SFT and preference datasets.
Accepted
Rejected
# per Prompt
Prompts
%
Prompts
%
1
2,145
95.6
154
6.9
2
86
3.8
1,298
57.9
3
9
0.4
753
33.6
4+
3
0.1
38
1.7
Total
2,243
100.0
2,243
100.0
Appendix
Table 11: Distribution of accepted and rejected responses per prompt in the preference dataset.
Training Stage
Num. examples
Num. Epochs
Batch size per device
Total train batch size
Gradient accum. steps
Learning rate (min)
SFT
3,007,676
1
4
128
2
1.5e-5 (min 1.5e-6)
DPO
195,123
1
1
64
4
1e-6 (min 1e-7)
Appendix
Table 12 : Training Hyperparameters by Post-Training Stage
Prompt source
Baseline
SFT-Only
Tie
Both Fail
N
Team 1
Overall
7.3% [4.1, 12.7]
35.3% ∗∗∗ [28.1, 43.3]
25.3%
32.0%
150
Team 2’s prompts
10.0% [4.3, 21.4]
18.0% n.s. [9.8, 30.8]
46.0%
26.0%
50
Team 3’s prompts
10.0% [4.3, 21.4]
24.0% n.s. [14.3, 37.4]
24.0%
42.0%
50
Team 1’s prompts (own)
2.0% [0.4, 10.5]
64.0% ∗∗∗ [50.1, 75.9]
6.0%
28.0%
50
Team 2
Appendix
Table 13: Per-assessor results: Baseline vs. SFT-Only. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Prompt source
Baseline
SFT+DPO
Tie
Both Fail
N
Team 1
Overall
6.7% [3.7, 11.8]
35.3% ∗∗∗ [28.1, 43.3]
26.7%
31.3%
150
Team 2’s prompts
6.0% [2.1, 16.2]
18.0% n.s. [9.8, 30.8]
48.0%
28.0%
50
Team 3’s prompts
12.0% [5.6, 23.8]
24.0% n.s. [14.3, 37.4]
20.0%
44.0%
50
Team 1’s prompts (own)
2.0% [0.4, 10.5]
64.0% ∗∗∗ [50.1, 75.9]
12.0%
22.0%
50
Team 2
Appendix
Table 14: Per-assessor results: Baseline vs. SFT+DPO. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Prompt source
SFT-Only
SFT+DPO
Tie
Both Fail
N
Team 1
Overall
10.7% [6.7, 16.6]
10.0% n.s. [6.2, 15.8]
50.0%
29.3%
150
Team 2’s prompts
6.0% [2.1, 16.2]
8.0% n.s. [3.2, 18.8]
58.0%
28.0%
50
Team 3’s prompts
16.0% [8.3, 28.5]
10.0% n.s. [4.3, 21.4]
32.0%
42.0%
50
Team 1’s prompts (own)
10.0% [4.3, 21.4]
12.0% n.s. [5.6, 23.8]
60.0%
18.0%
50
Team 2
Appendix
Table 15: Per-assessor results: SFT-Only vs. SFT+DPO. Own prompts shown in italics. 95% Wilson score CIs shown in brackets. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .
Team
Prompt set
Base- line
Curated wins
Tie
Both Fail
N
Team 1
Own prompts
1.3%
50.0%
26.0%
22.7%
150
Others’ prompts
6.3%
20.7%
38.0%
35.0%
300
χ2(3)=42.78 , p<.001∗∗∗
Team 2
Own prompts
13.4%
45.0%
30.2%
11.4%
149
Others’ prompts
8.4%
61.9%
25.1%
4.7%
299
χ2(3)=15.07 , p=.002∗∗
Appendix
Table 16: Outcome distributions on own vs. others’ prompts, pooled across all three pairwise comparisons. “Curated wins” pools wins by SFT-Only and SFT+DPO. p -values are from chi-square tests on the four-way outcome distribution. ∗∗p<.01 ; ∗∗∗p<.001 .
Comparison
Prompt set
Mod. x wins
Mod. y wins
Tie
Both Fail
N
Team 1
Base. vs. SFT+DPO
Own
2.0%
64.0%
12.0%
22.0%
50
Others
9.0%
21.0%
34.0%
36.0%
100
χ2(3)=28.03 , p<.001∗∗∗
SFT-Only vs. SFT+DPO
Own
10.0%
12.0%
60.0%
18.0%
50
Others
11.0%
9.0%
45.0%
35.0%
100
Appendix
Table 17: Outcome distributions by prompt authorship per comparison. Own prompts shown in italics. p -values are from chi-square tests on the four-way outcome distribution. ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 ; n.s. p≥.05 .