Long-MDR: Long-Context Reinforcement Learning for Multimodal Deep-Research Agents
Organizations: The Hong Kong University of Science and Technology (HKUST)
Abstract
The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumulate tens of thousands of tokens of multimodal context. Despite this trend, online RL for multimodal research agents remains largely confined to shorter contexts and interaction horizons. We push online RL training to 128k context and 75+ tool-interaction turns. To our knowledge, this is the first online multimodal deep-research RL study trained at 128k context, and the first trained with a 75 tool-turn horizon. Scaling to this regime exposes several practical limitations of conventional RL training. Early in training, weak policies make poor use of large interaction budgets, causing expensive rollouts with little reward improvement. Later, policy entropy can collapse before performance has saturated, prematurely ending useful learning. We introduce Long-MDR, a three-component training recipe designed specifically for this setting: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue. Together, these techniques improve both the learning efficiency and stability of long-horizon RL, enabling continued gains in a regime where direct training is slow and costly. At a 50-turn evaluation budget, our RL-trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B-9B agents.
Figures & tables
| Model | Size | MM Search [ 26 ] | Simple VQA [ 27 ] | Live VQA [ 28 ] | BC-VL [ 29 ] | BC-V 3 [ 30 ] | Vis Browse [ 31 ] |
| Proprietary models: direct answer | |||||||
| Claude-Opus-4.6 | – | 59.8 | 71.7 | 53.1 | 43.5 | 15.0 | 27.2 |
| Gemini-3.1-Pro | – | 64.2 | 64.1 | 65.0 | 41.4 | 19.3 | 23.7 |
| GPT-5.5 | – | 68.7 | 67.0 | 73.0 | 47.9 | 23.0 | 36.1 |
| Proprietary models: agentic workflow | |||||||
| Claude-Opus-4.6 | – | 81.7 | 81.7 | 88.0 | 63.2 | 55.0 | 62.7 |
| Training context | Training horizon | Table 1 evaluation horizon |
| 128k tokens | turns | 50 turns |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Round | Tool | Parameters | Returned excerpt |
|---|---|---|---|
| 1 | ReverseImageSearch | goal =Identify the event shown in the image, the location/hotel where it was held, and the people in the photos image_token =[REDACTED] | {"target":"Input 1 (<image_1>)","target_url":"https:[bin].r2.[bin].png","text_summary":"ReverseImageSearch candidate results for Input 1 (<i |
| 2 | CropImage | crop_config ={’image_1’: [[1, 1, 490, 990], [490, 1, 990, 990]]} | {"status":"success","message":"Successfully cropped 2 image(s). Results will appear as <obs_N> tags in the next message.","results":{"image_ |
| 3 | ReverseImageSearch | goal =[bin] image_token =[REDACTED] | {"results_per_image":[{"image_index":0,"target":"Input 1 (<obs_1>)","target_url":"https:[bin].r2.[bin].png","text_summary":"ReverseImageSear |
| 4 | TextSearch | query =Chanel Iman Pete Davidson wedding hotel location cliff side infinity pool red life rings | Titles owned by Atwell Public Library - Biblionix ( https://atwell.biblionix.com/catalog/titles/ ) Titles owned by Atwell Public Library. Thi |
| 5 | TextSearch | query =Chanel Iman pink off shoulder gown event | Inside Chanel Iman’s Final Wedding Dress Fitting Before … ( https://www.vogue.[bin] ) Also from the designer’s Fall 2018 bridal collection, |
| 6 | TextSearch | query =Chanel Iman Sterling Shepard wedding hotel location | Chanel Iman and Sterling Shepard’s Stunning Beverly Hills … ( https://www.brides.com/story/chanel-iman-sterling-shepard-wedding-photos ) Ch |