cs.CLSep 28, 2026

From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

Authors: Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, +1 more

Organizations: Qatar Computing Research Institute, HBKU, Qatar. · College of Islamic Studies, HBKU, Qatar. · Argumentation and Conflict Studies, Ibn Haldun University, Turkiye. · College of Science and Engineering, HBKU, Qatar.

Abstract

Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Alignment Forecasting: Predicting Misalignment From Training Data

    Sep 19, 2026Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky +1Large Language Model AlignmentMisalignment Persona

  2. Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

    Aug 31, 2026Camila Blank, Zhuofan Ying, Christopher Potts +2SycophancyAgreement

  3. PLURAL: A Global Dataset for Value Alignment

    Jul 9, 2026Dhruv Agarwal, Anya Shukla, Tanya Goyal +1Preference DatasetsPluralistic Alignment