cs.CVOct 4, 2026

Long-MDR: Long-Context Reinforcement Learning for Multimodal Deep-Research Agents

Authors: On Tai Tang

Organizations: The Hong Kong University of Science and Technology (HKUST)

Abstract

The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumulate tens of thousands of tokens of multimodal context. Despite this trend, online RL for multimodal research agents remains largely confined to shorter contexts and interaction horizons. We push online RL training to 128k context and 75+ tool-interaction turns. To our knowledge, this is the first online multimodal deep-research RL study trained at 128k context, and the first trained with a 75 tool-turn horizon. Scaling to this regime exposes several practical limitations of conventional RL training. Early in training, weak policies make poor use of large interaction budgets, causing expensive rollouts with little reward improvement. Later, policy entropy can collapse before performance has saturated, prematurely ending useful learning. We introduce Long-MDR, a three-component training recipe designed specifically for this setting: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue. Together, these techniques improve both the learning efficiency and stability of long-horizon RL, enabling continued gains in a regime where direct training is slow and costly. At a 50-turn evaluation budget, our RL-trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B-9B agents.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

    Date pendingMinghao Guo, Meng Cao, Sui Zhao +10Deep Research AgentsLong-Horizon Agent Evaluation

  2. Context-Aware RL for Agentic and Multimodal LLMs

    Jun 15, 2026Peiyang Xu, Bangzheng Li, Sijia Liu +4Multimodal GroundingContrastive RL

  3. DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

    Apr 21, 2026Shengqin Wang, Wentao Yan, Huichi Zhou +4Reward ShapingRL for Language Model Reasoning