Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
Figures & tables
Domain
Dimension
Benchmark
Evaluated Behavior
Safety
Refusal
HarmBench ( Mazeika et al., 2024 )
Refusal of harmful, unethical, or out-of-policy requests.
Jailbreak Robustness
HarmBench ( Mazeika et al., 2024 )
Preservation of refusal behavior under adversarial prompt templates.
Factuality
Truthfulness
TruthfulQA ( Lin et al., 2022 )
Accurate answers rather than plausible but false claims.
Hallucination
HalluLens ( Bang et al., 2025 )
Fabrication about non-existent entities or misremembered facts.
Stance stability
Belief Consistency
VAL-Bench ( Gupta et al., 2025 )
Stable stances under paraphrase, role prompting, or irrelevant context.
Sycophancy
Anthropic Sycophancy ( Perez et al., 2023 ; Sharma et al., 2024 )
Agreement with user views even when they conflict with evidence or prior answers.
Table 1: Alignment dimensions and benchmarks used for evaluation.
Figure 1 : Average RLVR vs SFT drift across categories. The figure shows the mean absolute alignment drift, in percentage points relative to the baseline, for each alignment category, pooled across all four models and both training domains. SFT exceeds RLVR drift most strongly in safety, factuality and controllability.
Figure 2 : Alignment drift relative to the instruction-tuned baseline, averaged across training domains for SFT and RLVR. Each cell reports the improvement in the metric value in percentage points. Coloring shows the sign-aligned change, with blue indicating improvement and red indicating degradation. Colors are clipped at ±6 pp for readability.
Figure 3 : Example alignment metrics over training, evaluated at intermediate checkpoints. Dotted lines mark the corresponding instruction-tuned baseline. Arrows after metric names indicate the preferred direction.
Figure 4 : Effect of increasing the KL coefficient β for KL-SFT. Sub-graphs show representative alignment metrics for the two models. Dashed gray lines mark the instruction-tuned baseline.
Beta
Shifted
Equivalent
Indeterminate
Total
Mean abs. drift
SFT (β=0)
32
33
47
112
4.68pp
0.05
33
30
49
112
4.26pp
0.10
27
35
50
112
3.80pp
0.50
13
46
53
112
2.29pp
Table 2: Effect of increasing the KL coefficient under KL-SFT. Counts use the δ=5 pp TOST margin and are computed over Qwen2.5-3B and Llama3.2-3B MATH and TACO sub-metric cells for SFT and KL-SFT.
Figure 5 : Concept-direction shift across post-training regimes. The y-axis shows the concept-distance ratio ∥μpos−μneg∥trained/∥μpos−μneg∥baseline at the AUC-selected layer. The dotted line marks the baseline ratio of 1.0 ; values below 1.0 indicate compression.
Figure 6 : Concept-distance shift versus signed behavioral alignment change for each probed concept. Each point is a trained model checkpoint.
Concept
Pearson r
CI
Holm sig.
Corrigibility
+0.65
[+0.29,+0.85]
✓
Sycophancy
−0.78
[−0.91,−0.49]
✓
Coordination
+0.57
[+0.19,+0.82]
✓
Power-seeking
−0.95
[−0.98,−0.88]
✓
Self-awareness
+0.81
[+0.60,+0.93]
✓
Stereotype
+0.69
[+0.38,+0.88]
✓
Table 3: Correlation between concept-direction shifts and behavioral drift. All reported Pearson correlations remain statistically significant after Holm–Bonferroni correction, with 95% confidence intervals given.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Data
Steps
LR
βtrain
Epochs
RLVR
MATH
150
5×10−6
0.005
≈2.6
RLVR
TACO
200
5×10−6
0.005
≈8.4
SFT
MATH
351
2×10−5
—
3
SFT
TACO
420
1×10−5
—
10
Appendix
Table 4: Run-specific training configuration. “Steps” is the optimizer-step count of the evaluated checkpoint.
Setting
GRPO
SFT
KL-SFT
Optimizer
AdamW ( β1=0.9 , β2=0.95 , ε=10−8 )
Weight decay
0.01
LR schedule
constant
cosine, 10% warmup
cosine, 10% warmup
Mini-batch
128 prompts
64 examples
64 examples
Micro-batch/GPU
8
4
1
Reference model
frozen πref
—
frozen πref
Appendix
Table 5: Settings shared within each training method. KL-SFT uses βtrain∈{0.05,0.10,0.50} and computes the full-distribution per-position DKL(πθ∥πref) .
Setting
MATH GRPO
TACO GRPO
Group size G
16 rollouts/prompt
Temperature
1.0
0.7
Top-p
1.0
0.95
Max prompt length
1024
1536
Max response length
1024
Clip ratio ε
0.2
Appendix
Table 6: GRPO rollout and reward configuration.
Figure 7 : The AUC curves for each of the seven concepts considered for each layer of each model. Higher AUC indicates stronger linear separability between the positive and negative examples for a given concept. These curves provide the validation basis for the concept-specific layer choices used in the residual-stream analyses.
Benchmark
N
Prompt
Decoding
Metric
Dir.
Aggregation
MMLU
14,042
0-shot
greedy
4-way accuracy
↑
mean over 57 subjects
GSM8K-CoT
1,319
8-shot CoT
greedy
numeric exact match
↑
single metric
Minerva-MATH
5,000
0-shot
greedy
symbolic equivalence
↑
mean over 7 subjects
HumanEval-instruct
164
0-shot
greedy
pass@1
↑
mean over tasks
MBPP-instruct
500
0-shot
greedy
pass@1
↑
mean over tasks
TACO
200
0-shot
sampled
unit-test pass rate
↑
mean over problems
Appendix
Table 7 : Evaluation benchmarks and headline metrics used in this work. Direction indicates whether higher or lower scores are preferred.