Organizations: School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University, Xi’an 710049, China · School of Computer Science and Technology, Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an 710049, China · State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.
Figures & tables
Figure 1: Comparison between conventional MLLM grounding and CoEvolve. Conventional methods compress spatial reasoning and localization into a terminal prediction, leaving intermediate spatial decisions implicit and difficult to correct. CoEvolve uses RER to construct a measurable progressive region trajectory and BDR to synchronously refine its coordinate fields while retaining the surrounding response structure.
Figure 2: Overview of CoEvolve. RER converts otherwise free-form grounding analysis into a progressive semantic–spatial state in which each reasoning step commits to an explicit region. The resulting response becomes an editable localization state for BDR. BDR performs bidirectional same-position refinement of its coordinate fields while retaining the surrounding response as context. Geometry-level objectives provide target geometry, and behavior-level objectives provide edit-preference signals.
RefCOCO
RefCOCO+
RefCOCOg
Method
val
testA
testB
val
testA
testB
val
test
Overall ↑
Qwen2-VL-7B ( Wang et al., 2024a )
91.7
93.6
87.3
85.8
90.5
79.5
87.3
87.8
87.94
Qwen2-VL-72B ( Wang et al., 2024a )
93.2
95.3
90.7
90.1
93.8
85.6
89.9
90.4
91.13
Qwen2.5-VL-7B ( Bai et al., 2025b )
90.0
92.5
85.4
84.2
89.1
76.9
87.2
87.2
86.56
Qwen2.5-VL-72B ( Bai et al., 2025b )
92.7
94.6
89.7
88.9
92.2
83.7
89.9
90.3
90.25
Qwen3-VL-8B-Instruct ( Bai et al., 2025a )
91.6
93.3
87.8
85.8
90.3
79.9
88.7
88.7
88.26
Table 1: Grounding performance on the eight standard RefCOCO-family splits, measured by Acc@0.5 . Overall is the arithmetic mean across the eight split-level results. It is computed from full-precision values when available and otherwise from the reported split-level values. Public methods follow their reported protocols; ‡ denotes our locally reproduced Qwen3.5-9B baseline. CoEvolve trains only on RefCOCOg train, without target-dataset fine-tuning. Best and second-best results are bolded and underlined.
RefCOCO
RefCOCO+
RefCOCOg
Aggregate Results
Method
testA
testB
testA
testB
test
mIoU ↑
Acc@0.7 ↑
Acc@0.9 ↑
Hi-R1 † ( Zhu et al., 2026b )
-
-
-
-
-
79.97
80.07
54.84
InternVL3.5-8B ‡ ( Wang et al., 2025 )
85.9/63.0
77.5/55.0
83.9/61.6
70.6/50.0
79.3/58.6
79.33
79.68
58.02
Qwen3-VL-4B ( Bai et al., 2025a ; Zhu et al., 2026c )
88.6/–
81.3/–
86.0/–
73.3/–
82.2/–
81.74
82.51
61.24
Qwen3-VL-8B ‡ ( Bai et al., 2025a )
88.8/66.3
81.1/58.4
86.4/64.9
74.8/54.0
83.0/61.7
82.39
83.09
61.35
Qwen3.5-9B ‡ ( Qwen Team, 2026 )
88.7/65.9
81.7/58.6
86.4/64.5
75.6/54.3
83.1/59.6
82.16
83.31
60.64
Table 2: Localization under stricter IoU thresholds on five test splits. Split entries report Acc@0.7/Acc@0.9 ; aggregate metrics are computed jointly across the splits. † denotes Hi-R1 aggregates reconstructed from its reported size groups; ‡ marks locally reproduced baselines (Appendix B.6 ). Bold and underline denote the best and second-best results.
Figure 3: Controlled box correction on five test splits with 30,969 deterministic anchors. Aggregate metrics are computed jointly over all anchors. CoEvolve BDR achieves the highest output mIoU, aggregate localization accuracies, mean IoU gain, improvement rate, and Repair@0.7 , while Qwen3.5-9B yields the lowest degradation rate. Higher values are better except for degradation rate.
Figure 4: Effect of a meaningful editable state and its associated anchor-aware objectives on five test splits. Aggregate metrics are computed jointly across the splits. Zero-Anchor BDR (ours) uses zero anchors; CoEvolve (ours) uses RER states and anchor-aware objectives. Gains are computed from unrounded aggregate values before display rounding.
State / editor
RER Trained?
Init. mIoU
Final mIoU
Δ mIoU
Init. Acc@0.9
Final Acc@0.9
Δ Acc@0.9
Structured Qwen3.5
No
63.48
–
–
24.28
–
–
Native output → BDR (ours)
No
82.16
84.57
+2.41
60.64
66.32
+5.68
RER → BDR (ours)
Yes
80.70
85.72
+5.02
59.81
68.64
+8.83
Table 3: Source-matched state construction and refinement on five test splits. Each source is paired with a separately trained BDR; metrics are computed jointly across the splits.
Method
Pr@0.5↑
Pr@0.7↑
Pr@0.9↑
mIoU ↑
LQVG ( Lan et al., 2024 )
83.41
75.91
43.53
74.02
LPVA ( Li et al., 2024 )
82.27
72.25
39.55
72.35
Qwen3.5-4B ( Qwen Team, 2026 ; Wang et al., 2026a )
53.10
39.40
–
49.00
InternVL3.5-8B ( Wang et al., 2025 ; Wang et al., 2026a )
64.90
50.50
–
56.90
RSGround-R1 ( Huang et al., 2026a )
71.80
58.70
–
63.40
GeoSearcher ( Wang et al., 2026a )
83.20
72.40
–
73.50
Table 4: Remote-sensing grounding on DIOR-RSVG. Public methods follow their reported protocols; RER and BDR are retrained on the DIOR-RSVG training split. † denotes our reproduced native baseline evaluated without DIOR-RSVG adaptation. Best and second-best results are bolded and underlined, respectively; dashes denote unreported results.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Supplementary training dynamics for RER. For visualization, we show the first 6K optimization steps; this window does not indicate the total training duration. From left to right: total reward; format and Step-format rewards; progressive region-evolution reward; and terminal box reward.
Proposal
Probability
Parameter range
Coordinate shift
0.25
s∼U(0.2,1.5) ; offsets in [−sw,sw] and [−sh,sh]
Centered shrink
0.25
width and height ratio r∼U(0.3,0.9)
Centered expansion
0.25
width and height ratio r∼U(1.3,3.0)
Random box
0.25
valid integer endpoints in [0,1000]
Appendix
Table 5: Geometric proposals used to construct BDR training perturbations. Probabilities refer to proposal selection at each rejection-sampling attempt; the accepted proportions may differ because proposals are filtered by the target IoU interval [0.2,0.5] .
Figure 6: Supplementary training dynamics for BDR. For visualization, we show the first 6K optimization steps; this window does not indicate the total training duration. From left to right: total and coordinate cross-entropy losses; coordinate regression and format losses; historical and final coordinate entropies; and general reconstruction and coordinate-wise preference loss.
Model
Native prompt interface
InternVL3.5-8B
Referring box with <ref> expression
Qwen3-VL-8B
JSON grounding with normalized box
Qwen3.5-9B
JSON grounding with normalized box
CogVLM-Grounding-17B
Caption-to-box grounding
Appendix
Table 6: Protocols for the locally reproduced standard-grounding baselines in Table 2 . “Native” denotes the prompt interface distributed or recommended for the corresponding model family.
Dataset
Split
Samples
Eight-split evaluation
Five-test evaluation
RefCOCO
val
10,834
✓
–
RefCOCO
testA
5,657
✓
✓
RefCOCO
testB
5,095
✓
✓
RefCOCO+
val
10,758
✓
–
RefCOCO+
testA
5,726
✓
✓
RefCOCO+
testB
4,889
✓
✓
Appendix
Table 7: Dataset composition used by the archived evaluations.
Dataset
Split
Samples
mIoU
Acc@0.5
Acc@0.7
Acc@0.9
RefCOCO
val
10,834
83.35
90.10
83.81
63.61
RefCOCO
testA
5,657
84.09
90.75
85.56
64.27
RefCOCO
testB
5,095
81.37
88.38
80.73
59.02
RefCOCO+
val
10,758
77.14
82.57
76.15
57.82
RefCOCO+
testA
5,726
79.32
85.23
79.03
59.10
RefCOCO+
testB
4,889
74.18
80.18
72.20
53.12
Appendix
Table 8: RER terminal localization on the eight standard splits. The aggregate row is computed jointly across the splits.
State source
Interface
RER adaptation
mIoU
Acc@0.7
Acc@0.9
Qwen3.5-9B
native output
No
82.16
83.31
60.64
Structured Qwen3.5-9B
structured
No
63.48
57.08
24.28
RER terminal answer (ours)
structured
Yes
80.70
80.56
59.81
Native-Output + BDR (ours)
native-output-derived
No
84.57
86.56
66.32
RER state + BDR (ours)
structured
Yes
85.72
87.94
68.64
Appendix
Table 9: Terminal grounding and source-matched refinement under the five-test metric aggregation. Metrics are computed jointly across the splits; invalid or unparseable outputs receive zero IoU, and BDR rows use the second refinement pass. Each refined state source uses a separately trained BDR, and the native and structured rows use their corresponding prompts in Section B.6 .
Method
RefCOCO A
RefCOCO B
RefCOCO+ A
RefCOCO+ B
RefCOCOg test
CoEvolve, mIoU / Acc@0.5
88.85 / 95.28
85.41 / 92.38
87.12 / 93.45
81.23 / 87.79
85.47 / 92.75
Zero-Anchor BDR, mIoU / Acc@0.5
87.91 / 94.22
84.24 / 90.62
86.00 / 92.19
78.89 / 84.27
81.34 / 89.79
CoEvolve, Acc@0.7 / Acc@0.9
92.65 / 74.24
87.05 / 67.79
90.67 / 73.11
82.33 / 64.68
86.86 / 65.15
Zero-Anchor BDR, Acc@0.7 / Acc@0.9
90.26 / 69.91
84.02 / 65.32
87.91 / 67.69
77.79 / 60.20
81.97 / 53.62
Appendix
Table 10: Split-level meaningful-state ablation. The upper block reports mIoU / Acc@0.5 and the lower block reports Acc@0.7 / Acc@0.9 .
Figure 8: A structured RER trajectory and its subsequent BDR coordinate edits for the query “broccoli on top right.” RER constructs a progressive localization trajectory, while BDR performs same-position coordinate refinement across multiple editable roles and substantially improves the final Answer. The image shows the first RER region, RER Answer, second-pass BDR Answer, and annotation. The lower matrix traces all coordinate states; amber cells mark edits from the preceding state.
Figure 9: Qualitative results on the RefCOCO family. Red solid boxes show model predictions, green dashed boxes show annotations, and badges report per-example IoU. The examples illustrate boundary correction, disambiguation among nearby instances, relational correction, and preservation of an accurate RER anchor.
Figure 10: Construction of the controlled-correction benchmark. The upper panel reports the number of examples and input mIoU for each test split; the dashed line marks the aggregate input mIoU over all examples. The lower panel reports the assignment across target IoU bands and perturbation types. Each example contributes one deterministic nonzero anchor.
State
mIoU
Pr@0.5
Pr@0.7
Pr@0.9
RER output
58.2248
64.2400
51.1200
23.0400
BDR Pass 1
75.4161
85.2533
74.6267
41.5200
BDR Pass 2
76.2295
86.2000
76.1600
43.4533
Appendix
Table 11: CoEvolve localization-state progression on DIOR-RSVG test expressions. RER and BDR are retrained on the DIOR-RSVG training split.
Figure 11: Qualitative results on DIOR-RSVG. Red solid boxes show model predictions and green dashed boxes show annotations. The first four rows show repair of a tiny object, a relationally specified instance, object extent, and a boundary further improved by the second BDR pass. The final row shows a failure in which refinement moves away from an initially accurate anchor. The rightmost column is a geometric crop around the annotation and final edit; no contrast enhancement is applied.
RefCOCO
RefCOCO+
RefCOCOg
Method
val
A
B
val
A
B
val
test
Macro-8
CRIS ( Wang et al., 2022 )
70.47
73.18
66.10
62.27
68.08
53.68
59.87
60.36
64.25
LAVT ( Yang et al., 2022 )
72.73
75.82
68.79
62.14
68.38
55.10
61.24
62.09
65.79
CGFormer ( Tang et al., 2023 )
74.75
77.30
70.64
64.54
71.00
57.14
64.68
65.09
68.14
ReLA ( Liu et al., 2023a )
73.82
76.48
70.18
66.04
71.02
57.65
65.00
65.97
68.27
MagNet ( Chng et al., 2024 )
75.24
78.24
71.05
66.16
71.32
58.14
65.36
66.03
68.94
Appendix
Table 12: Referring-expression segmentation cIoU on the eight standard splits. Public methods use their reported training and pretraining protocols; Macro-8 is the arithmetic mean of the eight split values. Best and second-best results are bold and underlined.
University of Chinese Academy of Sciences, Beijing, China · State Key Laboratory of Communication Content Cognition, Beijing, China · Peng Cheng Laboratory, Shenzhen, Guangdong, China +1
Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education School of Artificial Intelligence, Xidian University Xi’an, Shaanxi 710071, China