cs.CVAug 25, 2026

Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

Authors: Mingwen Zhang, Jisheng Dang, Minqiang Yang, Bimei Wang, Bin Hu, Tat-Seng Chua

Abstract

Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform their within-input peers, and a bootstrapped calibration loss steers temporal predictions toward verifier-preferred spans. Trained on source tasks and evaluated without target-dataset fine-tuning, a two-billion-parameter instantiation transfers zero-shot across grounded question answering, temporal grounding, and long-video question answering, reaching 28.7% intersection-over-union and 25.4% answer-grounding accuracy on a grounded-question-answering benchmark, 46.1% intersection-over-union on a temporal-grounding benchmark, and 54.1% on a long-video question-answering benchmark. Relative to a strong same-scale baseline, the gains are modest but consistent, with the clearest improvements on relevance-oriented metrics such as intersection-over-union and moderate-overlap recall. Within the tested benchmarks and transfer setting, the results support frozen verification as a training signal for evidence selection, while showing that strict boundary precision remains comparatively weaker. Code and models are available at https://anonymous.4open.science/r/MASIRL-E50C/

Explore similar work

CardsList
  1. SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

    Jun 23, 2026Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4Temporal Video GroundingVideo Reasoning

  2. Long-to-Short Video Evidence Reasoning for Grounded Question Answering

    Sep 14, 2026Kaiyan Chen, Junbin Xiao, Xun YangTemporal Video GroundingVideo QA

  3. Temporal-Aware Reasoning Optimization for Video Temporal Grounding

    Jun 8, 2026Minghang Zheng, Zihao Yin, Yi Yang +2Temporal Video GroundingTemporal Reasoning