Beyond Average Safety: Chance-Constrained LLM Fine-tuning
Organizations: Department of Electrical and Computer Engineering Johns Hopkins University
Abstract
Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chance constraint contains a discontinuous indicator, we introduce a differentiable majorization of the violation rate, yielding a tractable conservative constraint. We then develop a constraint-aware gradient descent method that treats the majorized constraint as a safe set in parameter space and minimally modifies the fine-tuning direction to preserve feasibility. The resulting update admits a closed form and produces a tail-aware safety correction that emphasizes examples near or above the degradation threshold. We conduct an extensive set of experiments on harmful fine-tuning across three different tasks and three models and show that our approach consistently outperforms the baselines that exist in the literature. These results suggest that safety preservation in LLM fine-tuning is better viewed as a reliability-constrained optimization problem than as average-risk regularization.
Figures & tables
| Method | Alignment | SFT | Post-hoc | Mechanism |
| SFT | ✓ | Plain supervised fine-tuning on the poisoned task; worst-case lower bound w/o defensive intervention. | ||
| Vaccine [ 16 ] | ✓ | SAM-style perturbation of attention representations during alignment to flatten the loss landscape against fine-tuning shifts. | ||
| Booster [ 14 ] | ✓ | Adds a harmful-perturbation regularizer ( ) that perturbs parameters along the harmful gradient before each alignment step. | ||
| Lisa [ 15 ] | ✓ | Bi-state optimization alternating fine-tuning and alignment phases, with a proximal pull toward the aligned checkpoint. | ||
| SafeGrad [ 35 ] | ✓ | Gradient surgery resolving conflicts between the task gradient and a KL-based alignment gradient at every fine-tuning step. | ||
| AsFT [ 32 ] | ✓ | Anchors fine-tuning to the alignment direction by penalizing the orthogonal component of the LoRA update ( ). |
| Defense | SST-2 | AG News | GSM8K | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HS | HS qK | FA | HS | HS qK | FA | HS | HS qK | FA | ||
| Qwen-3.5-4B ( ) | Base | 0.430 | 0.707 | 0.915 | 0.430 | 0.707 | 1.000 | 0.430 | 0.707 | 0.280 |
| Aligned (SFT) | 0.350 | 0.738 | 0.940 | 0.350 | 0.738 | 0.890 | 0.350 | 0.738 | 0.170 | |
| SFT | 0.546 0.042 | 0.890 0.015 | 0.973 0.016 | 0.545 0.043 | 0.901 0.025 | 0.890 0.015 | 0.637 0.047 | 0.945 0.017 | 0.633 0.064 | |
| Vaccine | 0.688 0.032 | 0.962 0.008 | 0.963 0.038 | 0.680 0.025 | 0.960 0.002 | 0.932 0.033 | 0.697 0.023 | 0.968 0.009 | 0.378 0.172 | |
| Booster | 0.665 0.021 | 0.940 0.012 | 0.943 0.018 | 0.674 0.013 | 0.949 0.007 | 0.952 0.021 | 0.715 0.012 | 0.951 0.002 | 0.427 0.083 | |
| Defense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | |
| Base | HS = 0.430 HS q90 = 0.707 FA = 0.915 (constant in ) | |||||||||||
| Aligned (SFT) | HS = 0.350 HS q90 = 0.738 FA = 0.940 (constant in ) | |||||||||||
| SFT | 0.425 0.040 | 0.810 0.029 | 0.985 0.013 | 0.437 0.037 | 0.828 0.021 | 0.977 0.014 | 0.516 0.020 | 0.862 0.015 | 0.983 0.020 | 0.546 0.042 | 0.890 0.015 | 0.973 0.016 |
| Vaccine | 0.575 0.046 | 0.940 0.018 | 0.973 0.019 | 0.604 0.027 | 0.944 0.008 | 0.972 0.010 | 0.665 0.025 | 0.958 0.006 | 0.983 0.010 | 0.688 0.032 | 0.962 0.008 | 0.963 0.038 |
| Booster | 0.464 0.015 | 0.872 0.013 | 0.973 0.008 | 0.493 0.015 | 0.883 0.007 | 0.957 0.008 | 0.600 0.026 | 0.918 0.010 | 0.943 0.016 | 0.665 0.021 | 0.940 0.012 | 0.943 0.018 |
| Defense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | |
| Base | HS = 0.430 HS q90 = 0.707 FA = 1.000 (constant in ) | |||||||||||
| Aligned (SFT) | HS = 0.350 HS q90 = 0.738 FA = 0.890 (constant in ) | |||||||||||
| SFT | 0.434 0.024 | 0.844 0.004 | 0.917 0.026 | 0.445 0.010 | 0.866 0.006 | 0.912 0.028 | 0.517 0.008 | 0.882 0.023 | 0.895 0.040 | 0.545 0.043 | 0.901 0.025 | 0.890 0.015 |
| Vaccine | 0.584 0.024 | 0.936 0.008 | 0.947 0.062 | 0.613 0.015 | 0.947 0.011 | 0.930 0.066 | 0.684 0.021 | 0.967 0.002 | 0.943 0.028 | 0.680 0.025 | 0.960 0.002 | 0.932 0.033 |
| Booster | 0.476 0.029 | 0.884 0.006 | 0.975 0.018 | 0.496 0.016 | 0.891 0.004 | 0.968 0.021 | 0.597 0.010 | 0.918 0.004 | 0.973 0.006 | 0.674 0.013 | 0.949 0.007 | 0.952 0.021 |
| Defense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | HS | HS q90 | FA | |
| Base | HS = 0.430 HS q90 = 0.707 FA = 0.280 (constant in ) | |||||||||||
| Aligned (SFT) | HS = 0.350 HS q90 = 0.738 FA = 0.170 (constant in ) | |||||||||||
| SFT | 0.470 0.024 | 0.877 0.006 | 0.665 0.010 | 0.510 0.043 | 0.889 0.016 | 0.665 0.020 | 0.605 0.026 | 0.938 0.010 | 0.660 0.043 | 0.637 0.047 | 0.945 0.017 | 0.633 0.064 |
| Vaccine | 0.636 0.024 | 0.948 0.013 | 0.498 0.020 | 0.653 0.024 | 0.952 0.004 | 0.465 0.044 | 0.696 0.015 | 0.964 0.005 | 0.465 0.026 | 0.697 0.023 | 0.968 0.009 | 0.378 0.172 |
| Booster | 0.565 0.010 | 0.924 0.002 | 0.460 0.025 | 0.587 0.016 | 0.923 0.005 | 0.450 0.015 | 0.658 0.009 | 0.938 0.007 | 0.452 0.019 | 0.715 0.012 | 0.951 0.002 | 0.427 0.083 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Defense | Hyperparameters |
|---|---|
| Vaccine | , align epochs |
| Booster | , , align epochs |
| Lisa | align step , ft step , , guide |
| SaLoRA | projection strength |
| SafeLoRA | projection threshold (applied after plain SFT fine-tuning) |
| SafeGrad | (KL alignment weight) |
| Value | HS | FA | |||
| sweep, fixing | |||||
| 0 | 0.323 | 0.256 | 0.676 | 0.942 | 0.925 |
| 0.01 | 0.311 | 0.250 | 0.691 | 0.930 | 0.925 |
| 0.05 | 0.314 | 0.252 | 0.704 | 0.949 | 0.930 |
| 0.1 | 0.320 | 0.253 | 0.707 | 0.930 | 0.920 |
| 0.2 | 0.350 | 0.252 | 0.723 | 0.949 | 0.935 |
| Defense | HS | FA | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Base | 0.430 | 0.161 | 0.248 | 0.365 | 0.531 | 0.707 | 0.824 | 0.938 | 0.984 | 0.915 |
| Aligned (SFT) | 0.350 | 0.094 | 0.164 | 0.296 | 0.504 | 0.738 | 0.867 | 0.961 | 1.000 | 0.940 |
| SFT | 0.546 0.042 | 0.119 0.010 | 0.227 0.020 | 0.446 0.040 | 0.705 0.037 | 0.890 0.015 | 0.947 0.008 | 0.992 0.000 | 1.000 0.000 | 0.973 0.016 |
| Vaccine | 0.688 0.032 | 0.149 0.023 | 0.316 0.033 | 0.611 0.026 | 0.861 0.022 | 0.962 0.008 | 0.986 0.002 | 1.000 0.000 | 1.000 0.000 | 0.963 0.038 |
| Booster | 0.665 0.021 | 0.143 0.013 | 0.305 0.017 | 0.584 0.030 | 0.826 0.026 | 0.940 0.012 | 0.970 0.006 | 1.000 0.000 | 1.000 0.000 | 0.943 0.018 |
| Lisa | 0.489 0.051 | 0.115 0.009 | 0.206 0.022 | 0.392 0.046 | 0.651 0.045 | 0.845 0.037 | 0.925 0.012 | 0.983 0.014 | 1.000 0.000 | 0.968 0.003 |
| Defense | HS | FA | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Base | 0.430 | 0.161 | 0.248 | 0.365 | 0.531 | 0.707 | 0.824 | 0.938 | 0.984 | 1.000 |
| Aligned (SFT) | 0.350 | 0.094 | 0.164 | 0.296 | 0.504 | 0.738 | 0.867 | 0.961 | 1.000 | 0.890 |
| SFT | 0.545 0.043 | 0.113 0.010 | 0.216 0.024 | 0.449 0.049 | 0.730 0.035 | 0.901 0.025 | 0.948 0.011 | 0.992 0.008 | 1.000 0.000 | 0.890 0.015 |
| Vaccine | 0.680 0.025 | 0.154 0.012 | 0.311 0.025 | 0.617 0.023 | 0.857 0.006 | 0.960 0.002 | 0.982 0.002 | 0.999 0.002 | 1.000 0.000 | 0.932 0.033 |
| Booster | 0.674 0.013 | 0.155 0.006 | 0.319 0.015 | 0.597 0.027 | 0.836 0.008 | 0.949 0.007 | 0.979 0.002 | 1.000 0.000 | 1.000 0.000 | 0.952 0.021 |
| Lisa | 0.510 0.038 | 0.108 0.009 | 0.200 0.019 | 0.407 0.032 | 0.666 0.022 | 0.874 0.020 | 0.940 0.012 | 0.992 0.004 | 1.000 0.000 | 0.900 0.018 |
| Defense | HS | FA | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Base | 0.430 | 0.161 | 0.248 | 0.365 | 0.531 | 0.707 | 0.824 | 0.938 | 0.984 | 0.280 |
| Aligned (SFT) | 0.350 | 0.094 | 0.164 | 0.296 | 0.504 | 0.738 | 0.867 | 0.961 | 1.000 | 0.170 |
| SFT | 0.637 0.047 | 0.133 0.025 | 0.275 0.049 | 0.558 0.055 | 0.814 0.049 | 0.945 0.017 | 0.979 0.009 | 1.000 0.000 | 1.000 0.000 | 0.633 0.064 |
| Vaccine | 0.697 0.023 | 0.176 0.007 | 0.346 0.017 | 0.614 0.028 | 0.872 0.016 | 0.968 0.009 | 0.991 0.002 | 1.000 0.000 | 1.000 0.000 | 0.378 0.172 |
| Booster | 0.715 0.012 | 0.175 0.022 | 0.361 0.016 | 0.635 0.014 | 0.847 0.004 | 0.951 0.002 | 0.982 0.002 | 1.000 0.000 | 1.000 0.000 | 0.427 0.083 |
| Lisa | 0.532 0.032 | 0.096 0.003 | 0.195 0.014 | 0.438 0.039 | 0.739 0.039 | 0.913 0.016 | 0.958 0.010 | 0.997 0.002 | 1.000 0.000 | 0.290 0.041 |
| Defense | step (ms) | p90 (ms) | GPU (MB) | samples/s | SFT |
|---|---|---|---|---|---|
| SFT | 136.8 | 160.5 | 14512 | 28.16 | 1.00 |
| Lisa | 138.6 | 153.3 | 14797 | 28.30 | 1.00 |
| SafeGrad | 334.7 | 371.5 | 26137 | 11.83 | 2.38 |
| AsFT | 525.6 | 551.4 | 16083 | 7.55 | 3.73 |
| Ours | 282.2 | 306.1 | 16739 | 13.98 | 2.01 |
| Task | Method | HS | HS q90 | FA |
|---|---|---|---|---|
| AG News | Our method | 0.324 | 0.688 | 0.930 |
| Primal-Dual | 0.346 | 0.755 | 0.935 | |
| GSM8K | Our method | 0.303 | 0.645 | 0.140 |
| Primal-Dual | 0.380 | 0.802 | 0.155 | |
| SST-2 | Our method | 0.317 | 0.684 | 0.940 |
| Primal-Dual | 0.311 | 0.703 | 0.940 |
| SST-2 | AG News | GSM8K | ||||||||
| Method | HS | HS | FA | HS | HS | FA | HS | HS | FA | |
| Qwen-3.5-4B (q90) | SFT-Aligned | |||||||||
| Antibody-Aligned | ||||||||||
| SFT | ||||||||||
| Antibody | ||||||||||
| Ours on SFT-Aligned | ||||||||||