Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
Figures & tables
Domain
Dimension
Benchmark
Evaluated Behavior
Safety
Refusal
HarmBench ( Mazeika et al., 2024 )
Refusal of harmful, unethical, or out-of-policy requests.
Jailbreak Robustness
HarmBench ( Mazeika et al., 2024 )
Preservation of refusal behavior under adversarial prompt templates.
Factuality
Truthfulness
TruthfulQA ( Lin et al., 2022 )
Accurate answers rather than plausible but false claims.
Hallucination
HalluLens ( Bang et al., 2025 )
Fabrication about non-existent entities or misremembered facts.
Stance stability
Belief Consistency
VAL-Bench ( Gupta et al., 2025 )
Stable stances under paraphrase, role prompting, or irrelevant context.
Sycophancy
Anthropic Sycophancy ( Perez et al., 2023 ; Sharma et al., 2024 )
Agreement with user views even when they conflict with evidence or prior answers.
Table 1: Alignment dimensions and benchmarks used for evaluation.
Figure 1 : Average RLVR vs SFT drift across categories. The figure shows the mean absolute alignment drift, in percentage points relative to the baseline, for each alignment category, pooled across all four models and both training domains. SFT exceeds RLVR drift most strongly in safety, factuality and controllability.
Figure 2 : Alignment drift relative to the instruction-tuned baseline, averaged across training domains for SFT and RLVR. Each cell reports the improvement in the metric value in percentage points. Coloring shows the sign-aligned change, with blue indicating improvement and red indicating degradation. Colors are clipped at ±6 pp for readability.
Figure 3 : Example alignment metrics over training, evaluated at intermediate checkpoints. Dotted lines mark the corresponding instruction-tuned baseline. Arrows after metric names indicate the preferred direction.
Figure 4 : Effect of increasing the KL coefficient β for KL-SFT. Sub-graphs show representative alignment metrics for the two models. Dashed gray lines mark the instruction-tuned baseline.
Beta
Shifted
Equivalent
Indeterminate
Total
Mean abs. drift
SFT (β=0)
32
33
47
112
4.68pp
0.05
33
30
49
112
4.26pp
0.10
27
35
50
112
3.80pp
0.50
13
46
53
112
2.29pp
Table 2: Effect of increasing the KL coefficient under KL-SFT. Counts use the δ=5 pp TOST margin and are computed over Qwen2.5-3B and Llama3.2-3B MATH and TACO sub-metric cells for SFT and KL-SFT.
Figure 5 : Concept-direction shift across post-training regimes. The y-axis shows the concept-distance ratio ∥μpos−μneg∥trained/∥μpos−μneg∥baseline at the AUC-selected layer. The dotted line marks the baseline ratio of 1.0 ; values below 1.0 indicate compression.
Figure 6 : Concept-distance shift versus signed behavioral alignment change for each probed concept. Each point is a trained model checkpoint.
Concept
Pearson r
CI
Holm sig.
Corrigibility
+0.65
[+0.29,+0.85]
✓
Sycophancy
−0.78
[−0.91,−0.49]
✓
Coordination
+0.57
[+0.19,+0.82]
✓
Power-seeking
−0.95
[−0.98,−0.88]
✓
Self-awareness
+0.81
[+0.60,+0.93]
✓
Stereotype
+0.69
[+0.38,+0.88]
✓
Table 3: Correlation between concept-direction shifts and behavioral drift. All reported Pearson correlations remain statistically significant after Holm–Bonferroni correction, with 95% confidence intervals given.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Data
Steps
LR
βtrain
Epochs
RLVR
MATH
150
5×10−6
0.005
≈2.6
RLVR
TACO
200
5×10−6
0.005
≈8.4
SFT
MATH
351
2×10−5
—
3
SFT
TACO
420
1×10−5
—
10
Appendix
Table 4: Run-specific training configuration. “Steps” is the optimizer-step count of the evaluated checkpoint.
Setting
GRPO
SFT
KL-SFT
Optimizer
AdamW ( β1=0.9 , β2=0.95 , ε=10−8 )
Weight decay
0.01
LR schedule
constant
cosine, 10% warmup
cosine, 10% warmup
Mini-batch
128 prompts
64 examples
64 examples
Micro-batch/GPU
8
4
1
Reference model
frozen πref
—
frozen πref
Appendix
Table 5: Settings shared within each training method. KL-SFT uses βtrain∈{0.05,0.10,0.50} and computes the full-distribution per-position DKL(πθ∥πref) .
Setting
MATH GRPO
TACO GRPO
Group size G
16 rollouts/prompt
Temperature
1.0
0.7
Top-p
1.0
0.95
Max prompt length
1024
1536
Max response length
1024
Clip ratio ε
0.2
Appendix
Table 6: GRPO rollout and reward configuration.
Figure 7 : The AUC curves for each of the seven concepts considered for each layer of each model. Higher AUC indicates stronger linear separability between the positive and negative examples for a given concept. These curves provide the validation basis for the concept-specific layer choices used in the residual-stream analyses.
Benchmark
N
Prompt
Decoding
Metric
Dir.
Aggregation
MMLU
14,042
0-shot
greedy
4-way accuracy
↑
mean over 57 subjects
GSM8K-CoT
1,319
8-shot CoT
greedy
numeric exact match
↑
single metric
Minerva-MATH
5,000
0-shot
greedy
symbolic equivalence
↑
mean over 7 subjects
HumanEval-instruct
164
0-shot
greedy
pass@1
↑
mean over tasks
MBPP-instruct
500
0-shot
greedy
pass@1
↑
mean over tasks
TACO
200
0-shot
sampled
unit-test pass rate
↑
mean over problems
Appendix
Table 7 : Evaluation benchmarks and headline metrics used in this work. Direction indicates whether higher or lower scores are preferred.
Large language models require continuous adaptation to new tasks while preserving safety alignment. However, fine-tuning on even benign data often compromises safety behaviors, including refusal of harmful requests, truthfulness, and commonsense reasoning. We investigate which training samples cause alignment drift through a data-centric lens. Our empirical analysis shows samples contribute unequally: high-gradient samples cause greater safety degradation and drive models toward pretrained distributions, while moderate-gradient samples enable task learning with minimal alignment loss. We propose gradient-based sample selection that filters high-gradient samples during fine-tuning. Across multiple model families on continual domain tasks, our method substantially improves alignment preservation while maintaining competitive task performance, without requiring curated safe data or architectural modifications. Our method is robust across selection ratios, task orderings, and diverse attack benchmarks.
We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at https://github.com/NWULIST/DAPA.
Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under subsequent fine-tuning. Existing explanations either attribute alignment fragility to gradient geometry or characterize it as a distributional shift in model outputs, yet few provide a unified account that bridges parameter-space learning dynamics with function-space alignment behavior during fine-tuning. In this work, we introduce a tractable alignment score and derive its closed-form update during fine-tuning, yielding a unified framework for alignment dynamics. Our analysis decomposes alignment updates into two competing components: a \textbf{\color{red!60!black} Rebound Force}, governed jointly by the current alignment state and the narrowness of model distribution, and a \textbf{\color{green!60!black} Driving Force}, determined by how the training distribution aligns with outcome-conditioned posteriors over aligned and non-aligned completions. This decomposition explains why prior alignment can be reversed by later fine-tuning and why narrower posterior structure strengthens such reversal. Moreover, our framework predicts a \textbf{Rehearsal Priming Effect}: prior alignment leaves a latent posterior imprint that amplifies the effective Driving Force upon re-exposure, leading to faster re-alignment. We validate these predictions across safety alignment, emergent misalignment, and sentiment settings, demonstrating consistent alignment reversal and accelerated re-alignment under re-exposure. In addition, controlled experiments in safety alignment confirm the predicted dependence of rebound strength on posterior narrowness. Together, these results provide a unified dynamical perspective on how alignment is disrupted and reactivated during LLM fine-tuning.