cs.AIOct 1, 2026

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Authors: Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang

Organizations: PRADA Lab, King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City, University of Macau · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · School of Computer Science and Engineering, University of Electronic Science and Technology of China

Abstract

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1.23 points. Source code can be found in the supplemental material.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

    Aug 16, 2026Md Messal Monem Miah, Adrita Anika, Zhiyuan Yu +1Jailbreak AttacksTrajectory-Aware Hidden-State Analyses

  2. THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models

    Jun 1, 2026Zhiqing Ma, Zhonghao Xu, Dong Yu +3Large Language Model JailbreaksJailbreak Attacks

  3. One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent

    May 7, 2026Xinjie Shen, Rongzhe Wei, Peizhi Niu +6AttackerMalicious Agents