cs.CLOct 8, 2026

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

Authors: Haitong Jiang, Chunlin Liu, Sihan Tang, Chan Wu, Xiaoqing Su, Yuhong Feng

Organizations: College of Computer Science and Software Engineering, Shenzhen University · School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen), Shenzhen, China

Abstract

Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at https://github.com/kevinjiang0121-cyber/IRIS.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

    Apr 29, 2026Mingqian Zheng, Malia Morgan, Liwei Jiang +2LLM Safety BenchmarksLanguage Model Safety Evaluation

  2. OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

    Jul 2, 2026Rheeya Uppaal, Seungwoo Lyu, Selina Sung +1LLM Safety BenchmarksLanguage Model Safety Evaluation