cs.CLOct 8, 2026

When History Helps and Hurts: Selective History Use across Multimodal Turns

Authors: Shuoyang Sun, Kerui Gu, Hao Fang, Shaoli Huang, Bin Chen

Organizations: Harbin Institute of Technology, Shenzhen · AgiBot · Tsinghua Shenzhen International Graduate School, Tsinghua University

Abstract

Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories

    Aug 6, 2026Xiaoqing Wu, Xingyu Fan, Feifei Li +1Tool-Use EvaluationLLM Tool Use

  2. M3^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

    Jun 5, 2026Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6Multimodal GroundingEfficient Multimodal Inference

  3. Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

    Jul 31, 2026Eileen Ye, Jiawen Tao, Yaoming Li +7Conversational MemoryLLM Evaluation