cs.CVSep 29, 2026

MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding

Authors: Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li, Maosong Sun, Ruixuan Li

Organizations: Huazhong University of Science and Technology Wuhan, China · Tsinghua University Beijing, China

Abstract

Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task--sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

    May 27, 2026Xuanzhao Dong, Wenhui Zhu, Peijie Qiu +11Multimodal Large Language ModelsRecent Vision-Language Models

  2. Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

    May 21, 2026Changyuan Tian, Zhicong Lu, Huaxing Liu +7Multimodal ReasoningRecent Vision-Language Models