cs.LGJun 26, 2026

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment

Authors: Changyue Li, Jiaming He, Youliang Yuan, Jialin Wu, Boxi Yu, Zhicong Huang, Pinjia He

Organizations: The Chinese University of Hong Kong, Shenzhen · Ant Group · Lero the Research Ireland Centre for Software, University of Limerick

Abstract

Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

    Jul 13, 2026Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3Large Language Model SafetyLarge Language Model Fine-Tuning

  2. Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

    May 23, 2026Seokil Ham, Jaehyuk Jang, Wonjun Lee +1Large Language Model Fine-TuningJailbreak Attacks

  3. Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

    Apr 19, 2026Thong Bach, Truyen TranSafety AlignmentLarge Language Model Alignment