When History Helps and Hurts: Selective History Use across Multimodal Turns
Organizations: Harbin Institute of Technology, Shenzhen · AgiBot · Tsinghua Shenzhen International Graduate School, Tsinghua University
Abstract
Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.
Figures & tables
| Benchmark | Scope | Structure | History Design | |||
| Modality | Scale | Update | Direct | Role | Conflict | |
| LongMemEval | T | 500 q. | ✓ | ✗ | ✗ | ✗ |
| Lost in Conversation | T | 600 tasks | ✗ | ✓ | ✗ | ✗ |
| MMDU | V+T | 110 dlg. | ✗ | ✗ | ✗ | ✗ |
| MultiVerse | V+T | 647 dlg. | ✗ | ✗ | ✗ | ✗ |
| MMRC | V+T | 5,120 dlg. | ✓ | ✗ | ✗ | ✗ |
| Model | Operation | Accuracy (%) | Gain over Neither (pp) | ||||
|---|---|---|---|---|---|---|---|
| Neither | Question | Options | Both | Question | Options | ||
| Qwen3-Omni | Reconfirm | 97.7 | 94.5 | 99.2 | 100.0 | -3.1 | +1.6 |
| Reground | 57.8 | 72.7 | 99.2 | 98.4 | +14.8 | +41.4 | |
| MiniCPM-o 4.5 | Reconfirm | 85.9 | 90.6 | 98.4 | 97.7 | +4.7 | +12.5 |
| Reground | 29.7 | 42.2 | 92.2 | 96.1 | +12.5 | +62.5 | |
| Method | Qwen3-Omni | MiniCPM-o 4.5 | ||
|---|---|---|---|---|
| OpenQA | MCQ | OpenQA | MCQ | |
| Neutral | 71.4 | 78.8 | 59.8 | 59.6 |
| Restatement | 76.0 | 86.4 | 74.2 | 78.2 |
| Reweighting | 73.2 | 77.6 | 61.4 | 60.2 |
| Routed policy | 78.2 | 86.0 | 77.8 | 78.2 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Exposure indicator | Test tasks ( ) |
| Any input parent occurs in development | 3,072 |
| Any input parent occurs in gradient-used data | 1,875 |
| An exact input window occurs in development | 2,947 |
| An exact input window occurs in gradient-used data | 1,571 |
| Task has an earlier training/development tag | 2,268 |
| Target parent occurs in development | 918 |
| Display name | Scope; limit | Execution and recorded settings |
| Qwen3.8-Omni-Flash ( Alibaba Cloud, 2026b ) | A/V; 1,024 4,096 | Hosted qwen3.8-omni-flash ; temperature 0, reasoning effort none, 2 fps; companion WAV and silent video. |
| Qwen3-Omni ( Qwen Team, 2025a ) | A/V; 128 | Checkpoint: Qwen3-Omni-30B-A3B-Instruct . Local native runner; frozen formal decoding configuration. |
| Doubao-Seed-2.0-Lite ( Seed, 2026 ) | A/V; 1,024 4,096 | Hosted doubao-seed-2-0-lite-260428 ; temperature 0, 2 fps; audio, silent video, and text in order. |
| MiniCPM-o 4.5 ( Cui et al., 2026 ) | A/V; 256 | Recorded runner key: minicpmo_45_instruct . Local native runner; frozen formal decoding configuration. |
| MiMo-V2.5 ( XiaomiMiMo, 2026 ) | A/V; 4,096 | Recorded runner key: mimo_v2_5 . Local SGLang, tensor parallelism 4; configured adaptive cap 8,192. |
| Nemotron-3-Nano-Omni ( NVIDIA, 2026 ) | A/V; 1,024 | Model: Nemotron-3-Nano-Omni-30B . Local SGLang single-turn and vLLM multi-turn execution. |
| Task-carrying history | Evidence-carrying history | |||||||||||
| Model | Reconfirm | Reground | Retrieve | Rebind | ||||||||
| ST | MT | ST | MT | ST | MT | ST | MT | |||||
| Omni-modal | ||||||||||||
| Qwen3.8-Omni-Flash | 96.2 | 94.4 | 1.8 | 96.2 | 95.0 | 1.1 | 93.0 | 94.4 | -1.4 | 92.9 | 80.6 | 12.3 |
| Qwen3-Omni | 95.0 | 95.4 | -0.3 | 95.0 | 91.0 | 4.0 | 94.2 | 94.2 | 0.0 | 94.2 | 85.9 | 8.3 |
| Doubao-Seed-2.0-Lite | 97.4 | 95.4 | 2.1 | 97.4 | 93.1 | 4.3 | 95.4 | 94.7 | 0.6 | 95.4 | 75.8 | 19.5 |
| Task-carrying history | Evidence-carrying history | |||||||||||
| Model | Reconfirm | Reground | Retrieve | Rebind | ||||||||
| ST | MT | ST | MT | ST | MT | ST | MT | |||||
| Omni-modal | ||||||||||||
| Qwen3.8-Omni-Flash | 96.9 | 91.3 | 5.6 | 97.0 | 73.1 | 23.9 | 93.0 | 93.8 | -0.8 | 93.0 | 74.6 | 18.4 |
| Qwen3-Omni | 94.9 | 84.8 | 10.1 | 94.9 | 72.5 | 22.4 | 94.2 | 74.7 | 19.5 | 94.2 | 19.5 | 74.7 |
| Doubao-Seed-2.0-Lite | 97.8 | 82.9 | 14.9 | 97.8 | 77.9 | 19.8 | 95.4 | 76.5 | 18.9 | 95.4 | 14.7 | 80.6 |
| Task-carrying history | Evidence-carrying history | |||||||||||
| Model | Reconfirm | Reground | Retrieve | Rebind | ||||||||
| ST | MT | ST | MT | ST | MT | ST | MT | |||||
| Omni-modal | ||||||||||||
| Qwen3.8-Omni-Flash | 98.6 | 98.6 | 0.0 | 98.6 | 93.9 | 4.6 | 99.0 | 99.0 | 0.0 | 99.0 | 81.9 | 17.1 |
| Qwen3-Omni | 98.6 | 97.0 | 1.6 | 98.6 | 67.4 | 31.2 | 99.5 | 99.2 | 0.3 | 99.5 | 95.2 | 4.3 |
| Doubao-Seed-2.0-Lite | 99.4 | 98.7 | 0.6 | 99.4 | 71.8 | 27.5 | 99.5 | 98.9 | 0.6 | 99.5 | 86.7 | 12.8 |
| Task-carrying history | Evidence-carrying history | |||||||||||
| Model | Reconfirm | Reground | Retrieve | Rebind | ||||||||
| ST | MT | ST | MT | ST | MT | ST | MT | |||||
| Omni-modal | ||||||||||||
| Qwen3.8-Omni-Flash | 98.6 | 95.3 | 3.2 | 98.6 | 89.9 | 8.7 | 99.0 | 98.9 | 0.2 | 99.0 | 76.8 | 22.2 |
| Qwen3-Omni | 98.9 | 95.8 | 3.0 | 98.9 | 48.2 | 50.7 | 99.5 | 91.5 | 8.0 | 99.5 | 24.0 | 75.5 |
| Doubao-Seed-2.0-Lite | 99.2 | 82.2 | 17.0 | 99.2 | 62.6 | 36.6 | 99.5 | 91.2 | 8.3 | 99.5 | 29.1 | 70.4 |
| Model | OpenQA | MCQ | ||||
|---|---|---|---|---|---|---|
| Rule | Content | Rule | Content | |||
| Qwen3.8-Omni-Flash | 4994 | 9.5 | 7.6 | 4997 | 7.7 | 7.0 |
| Qwen3-Omni | 5000 | 13.9 | 17.3 | 5000 | 21.9 | 21.8 |
| Doubao-Seed-2.0-Lite | 5000 | 19.2 | 20.1 | 5000 | 10.2 | 21.7 |
| MiniCPM-o 4.5 | 5000 | -0.9 | 32.5 | 5000 | 42.5 | 39.5 |
| MiMo-V2.5 | 5000 | 3.2 | 28.0 | 5000 | 29.3 | 33.0 |
| Model | Condition | Original | Alternative | Disagreements | |
|---|---|---|---|---|---|
| Qwen3-Omni | ST | 80 | 94.60 | 82.79 | 17 |
| Qwen3-Omni | MT | 80 | 77.26 | 71.96 | 9 |
| MiniCPM-o 4.5 | ST | 80 | 93.08 | 83.50 | 2 |
| MiniCPM-o 4.5 | MT | 80 | 60.62 | 54.28 | 3 |
| Qwen3.8-Omni-Flash | ST | 80 | 94.76 | 81.86 | 14 |
| Qwen3.8-Omni-Flash | MT | 80 | 87.14 | 76.25 | 11 |
| Model | Interface | Reground recall | Source-fact ST | Source-fact MT |
|---|---|---|---|---|
| Qwen3-Omni | OpenQA | 1210/1250 | 656/712 | 532/712 |
| Qwen3-Omni | MCQ | 1226/1250 | 708/712 | 640/712 |
| MiniCPM-o 4.5 | OpenQA | 756/1250 | 644/712 | 476/712 |
| MiniCPM-o 4.5 | MCQ | 940/1250 | 692/712 | 582/712 |
| Qwen3.8-Omni-Flash | OpenQA | 1246/1248 | 652/712 | 610/712 |
| Qwen3.8-Omni-Flash | MCQ | 1239/1250 | 696/712 | 676/712 |
| Reground: explicit question restatement | ||||||
|---|---|---|---|---|---|---|
| Model | Interface | Native | Changed | (pp) | Rescue/harm | |
| Qwen3-Omni | OpenQA | 1250 | 81.76 | 95.20 | +13.44 | 180/12 |
| Qwen3-Omni | MCQ | 1250 | 57.76 | 98.80 | +41.04 | 514/1 |
| MiniCPM-o 4.5 | OpenQA | 1250 | 58.32 | 92.64 | +34.32 | 441/12 |
| MiniCPM-o 4.5 | MCQ | 1250 | 28.00 | 97.04 | +69.04 | 867/4 |
| Qwen3.8-Omni-Flash | OpenQA | 1249 | 84.07 | 96.40 | +12.33 | 159/5 |
| Model | Operation | |||||
|---|---|---|---|---|---|---|
| Qwen3-Omni | Reconfirm | 128 | 125 | 121 | 127 | 128 |
| Qwen3-Omni | Reground | 128 | 74 | 93 | 127 | 126 |
| Qwen3-Omni | Overall | 256 | 199 | 214 | 254 | 254 |
| MiniCPM-o 4.5 | Reconfirm | 128 | 110 | 116 | 126 | 125 |
| MiniCPM-o 4.5 | Reground | 128 | 38 | 54 | 118 | 123 |
| MiniCPM-o 4.5 | Overall | 256 | 148 | 170 | 244 | 248 |
| Model | Interface | Before swap | After swap | Adopted/abandoned | |
|---|---|---|---|---|---|
| Qwen3-Omni | OpenQA | 1024 | 1 | 374 | 373/0 |
| Qwen3-Omni | MCQ | 980 | 1 | 400 | 399/0 |
| MiniCPM-o 4.5 | OpenQA | 1024 | 2 | 97 | 96/1 |
| MiniCPM-o 4.5 | MCQ | 980 | 0 | 34 | 34/0 |
| Qwen3.8-Omni-Flash | OpenQA | 1022 | 2 | 136 | 134/0 |
| Qwen3.8-Omni-Flash | MCQ | 980 | 1 | 212 | 211/0 |
| Model | Target ST | Old ST | New ST | All selected | All three correct |
|---|---|---|---|---|---|
| Qwen3-Omni | 120/128 | 120/128 | 123/128 | 40/128 | 37/111 |
| MiniCPM-o 4.5 | 113/128 | 122/128 | 121/128 | 6/128 | 4/107 |
| Qwen3.8-Omni-Flash | 119/128 | 119/128 | 117/128 | 13/128 | 12/102 |
| Model | Interface | Operation | ST | Near | Far-N | Far-C | |
|---|---|---|---|---|---|---|---|
| Qwen3-Omni | OpenQA | Reground | 400 | 97.0 | 81.8 | 71.8 | 75.5 |
| Qwen3-Omni | MCQ | Reground | 400 | 99.8 | 43.8 | 43.2 | 47.2 |
| Qwen3-Omni | OpenQA | Rebind | 531 | 94.9 | 79.7 | 44.4 | 31.1 |
| Qwen3-Omni | MCQ | Rebind | 531 | 99.8 | 92.1 | 74.2 | 63.5 |
| Qwen3.8-Omni-Flash | OpenQA | Reground | 399 | 98.5 | 82.5 | 75.7 | 78.2 |
| Qwen3.8-Omni-Flash | MCQ | Reground | 400 | 99.5 | 92.2 | 89.5 | 91.0 |
| Operation | Original | Message | R/H | Explicit media | R/H | |
|---|---|---|---|---|---|---|
| Reconfirm | 32 | 31 | 26 | 0/5 | 25 | 0/6 |
| Reground | 32 | 26 | 24 | 1/3 | 24 | 3/5 |
| Retrieve | 32 | 30 | 30 | 0/0 | 30 | 0/0 |
| Rebind | 32 | 25 | 26 | 4/3 | 28 | 4/1 |
| Overall | 128 | 112 | 106 | 5/11 | 107 | 7/12 |
| Training objective | OpenQA | MCQ | ||
|---|---|---|---|---|
| ST | MT | ST | MT | |
| SFT | 90.8 | 74.8 | 96.8 | 80.8 |
| SFT+Inv | 90.0 | 75.0 | 96.8 | 80.6 |
| SFT+Inv+sensitivity | 90.8 | 74.4 | 96.8 | 81.0 |
| Metric / subset | OpenQA | MCQ | |||||
|---|---|---|---|---|---|---|---|
| Base | SFT | SFT+Inv | Base | SFT | SFT+Inv | ||
| Overall ST | 5000 | 94.60 | 94.84 | 94.92 | 99.12 | 99.12 | 99.12 |
| Overall MT | 5000 | 77.26 | 80.68 | 80.70 | 77.28 | 79.64 | 79.84 |
| Overall | 5000 | 17.34 | 14.16 | 14.22 | 21.84 | 19.48 | 19.28 |
| Reconfirm MT | 1250 | 90.08 | 91.68 | 91.92 | 96.40 | 95.76 | 95.68 |
| Reground MT | 1250 | 81.76 | 85.84 | 85.52 | 57.76 | 62.96 | 63.36 |
| Comparison | Interface | (pp) | Rescue/harm | Root interval |
|---|---|---|---|---|
| SFT vs. Base | OpenQA | +3.42 | 261/90 | |
| SFT vs. Base | MCQ | +2.36 | 163/45 | |
| SFT+Inv vs. SFT | OpenQA | +0.02 | 26/25 | |
| SFT+Inv vs. SFT | MCQ | +0.20 | 41/31 |
| Retained subset | OpenQA | MCQ | |||
|---|---|---|---|---|---|
| Gain | R/H | Gain | R/H | ||
| All Test tasks | 5,000 | +3.42 | 261/90 | +2.36 | 163/45 |
| Target parent disjoint from Dev | 4,082 | +3.58 | 218/72 | +2.13 | 127/40 |
| All input parents disjoint from Dev | 1,928 | +2.65 | 78/27 | +1.66 | 54/22 |
| Target parent disjoint from training | 4,514 | +3.57 | 239/78 | +2.10 | 139/44 |
| All input parents disjoint from training | 3,125 | +3.14 | 148/50 | +2.50 | 106/28 |
| Action | OpenQA | MCQ | ||||
|---|---|---|---|---|---|---|
| ST | MT | ST | MT | |||
| Restatement | 95.00 | 81.48 | 13.52 | 99.12 | 87.80 | 11.32 |
| Reweighting | 95.00 | 79.40 | 15.60 | 99.12 | 77.32 | 21.80 |
| Routed | 95.00 | 84.02 | 10.98 | 99.12 | 88.02 | 11.10 |
| Model | Interface | ST | Neutral | Restate | Reweight | Routed |
|---|---|---|---|---|---|---|
| Qwen3-Omni | OpenQA | 89.2 | 71.4 | 76.0 | 73.2 | 78.2 |
| Qwen3-Omni | MCQ | 96.8 | 78.8 | 86.4 | 77.6 | 86.0 |
| MiniCPM-o 4.5 | OpenQA | 89.2 | 59.8 | 74.2 | 61.4 | 77.8 |
| MiniCPM-o 4.5 | MCQ | 99.6 | 59.6 | 78.2 | 60.2 | 78.2 |