Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes
Authors: Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningwei Ouyang, Shaofeng Liang, Heyi Lin, Jinjing Zhu, Yang Shi, +4 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · Jiangnan University, Wuxi, China · Central China Normal University, Wuhan, China · Xi’an Jiaotong-Liverpool University, Suzhou, China · CUHK, Hong Kong SAR, China · Peking University, Beijing, China · Wuhan University, Wuhan, China · Fudan University, Shanghai, China · The Hong Kong University of Science and Technology, Hong Kong SAR, China
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and boxes separately leave this failure unpenalized. We trace the mismatch to the conventional answer-then-ground factorization, which commits to a numerical claim before any object is enumerated. To measure it, we build RoadSceneVQA-G, a benchmark of 34.7K question-answer pairs in which every free-form answer is linked to the set of boxes that witnesses it, and we propose the Answer-Grounding Consistency (AGC) evaluation suite. To address it, we introduce Enumerate-then-Answer (EtA), which reverses the generation order so that answer-evidence agreement becomes a property of the output structure, and Enumeration-Consistent Policy Optimization (ECPO), a reinforcement learning stage that uses the union of multiple rollouts as a recall teacher without ground-truth boxes. EtA raises say-point consistency from 26.6% to 93.7% and grounding F1 from 52.2% to 73.0%, and ECPO further increases F1 to 75.6% without per-box supervision. On gRefCOCO, the same framework outperforms the strongest compared method, indicating that it transfers beyond traffic scenes. The project is available at https://github.com/GuanRunwei/RoadSceneVQA-G.
Figures & tables
Fig. 1: Grounded VQA in a roadside scene. A direct VQA MLLM answers without evidence. A grounding model trained with naive GRPO states one car but localizes two (say-point mismatch). EtA with ECPO first localizes all evidence, then reads the answer from it.
Datasets
Venues
Domains
# Pairs
Open A.
Grd.
Multi.
∅ -grd.
Consist.
General VQA and visual grounding benchmarks
GQA [ 14 ]
CVPR 2019
General scenes
∼ 22M
×
✓
✓
×
×
RefCOCO+ [ 11 ]
ECCV 2016
General objects
∼ 20K
×
✓
×
×
×
gRefCOCO [ 15 ]
IJCV 2026
General objects
∼ 260K
×
✓
✓
✓
×
Traffic-scene VQA benchmarks
DriveBench [ 16 ]
ICCV 2025
Autonomous driving
∼ 60K
✓
×
×
×
×
TABLE I: Comparison of RoadSceneVQA-G with existing benchmarks. Open A. : free-form textual answers. Grd. : grounding annotations tied to the answer. Multi. : multi-object boxes per QA pair. ∅ -grd. : explicit no-ground cases. Consist. : answer-grounding consistency is directly measurable.
Fig. 2: Statistics of RoadSceneVQA-G. (a) Share of questions that require localization in each split. (b) Distribution of the number of ground-truth boxes per grounded question. (c) Count stated in the answer versus the number of ground-truth evidence boxes. (d) Box area as a percentage of the image area (log scale). (e) Spatial distribution of the referred targets over normalized image coordinates.
Split
Items
No-Grd.
Grd.
Multi-Box
Avg. Box
Train
30,059
59.6%
40.4%
32.4%
2.06
Test
4,677
15.8%
84.2%
40.1%
1.98
TABLE II: Composition of the RoadSceneVQA-G splits. No-Grd. : questions without evidence boxes. Grd. : questions with at least one evidence box. Multi-Box : share of grounded questions with at least two boxes. Avg. Box : mean number of evidence boxes per grounded question.
Tier
Category
Train
Test
Perception
Existence
24.7
11.0
Enumeration
8.7
27.8
Attribute recognition
41.5
31.0
Object identification
2.4
9.3
Scene inference
14.5
3.9
Reasoning
Spatial reasoning
4.1
9.8
TABLE III: Capability taxonomy of RoadSceneVQA-G.
Fig. 3: Overview of the proposed framework. Stage 1 (top): Enumerate-then-Answer (EtA) serializes the grounded objects before the answer, so the answer is read out from the enumerated evidence and an empty box set encodes the decision not to ground. Stage 2 (bottom): Enumeration-Consistent Policy Optimization (ECPO) samples G rollouts from the current policy, pools and clusters their predicted boxes by IoU, keeps the clusters supported by at least k distinct rollouts as the consensus set, and rewards each rollout by its coverage of this set. Group-relative advantages update the policy with KL regularization toward the frozen EtA-SFT reference.
Answer quality
Grounding
AGC (decision / consistency / joint)
Method
Ans.
METEOR
BERTScore
F1
AP
GDA
SPC
CntAcc
GCA
Zero-shot MLLMs: answer-capable but weak or unstable box enumeration
InternVL3.5-8B-HF [ 54 ]
✓
22.0
89.4
36.9
13.0
60.0
60.0
30.3
4.3
Qwen3.5-27B (thinking) [ 55 ]
✓
20.1
86.1
49.1
26.1
79.5
82.3
49.8
13.2
Qwen3.5-27B (non-thinking) [ 55 ]
✓
19.7
88.8
64.8
34.5
78.7
80.1
46.9
10.8
LLaVA-OneVision-2-8B [ 56 ]
✓
15.3
90.5
8.7
1.1
32.6
20.7
12.6
3.0
TABLE IV: Overall comparison on the RoadSceneVQA-G test split. Ans.: whether the method produces a VQA answer. N/A: the metric requires an answer, which VG specialists do not produce. AP denotes AP 50:95 , and GCA denotes GCA-l. Best results are in bold, and further metrics are in Appendix .
Method
testA
testB
VLT [ 65 ]
40.2 / 34.1
30.2 / 32.5
MDETR [ 66 ]
50.0 / 34.5
36.5 / 31.0
UNINEXT [ 67 ]
46.4 / 49.3
42.9 / 48.2
ReLA [ 68 ]
50.4 / 59.0
44.6 / 58.4
LLMDet [ 69 ]
62.8 / 78.9
57.3 / 76.2
Ours (InternVL3-1B [ 64 ] )
TABLE V: Generalized referring expression comprehension on gRefCOCO. Each entry reports Pr@F1=1 / N-acc. (no-target accuracy).
Method
Count MAE ↓
Rec@0.5 ↑
SPC ↑
MSS/PPM-CNN [ 70 ]
4.45
–
–
Direct-answer SFT
2.48
–
–
EtA-SFT
2.51
82.3
78.2
Vanilla GRPO [ 36 ]
2.53
82.3
78.5
ECPO
2.46
82.3
81.0
TABLE VI: Dense car counting on CARPK. Count MAE is computed per image, and Rec@0.5 is box recall at an IoU of 0.5.
Configuration
SPC
Prec
Rec
F1
CntAcc
GCA
Standard SFT (raw)
26.6
83.8
37.9
52.2
52.1
46.3
+ count rewriting
99.9
83.8
37.9
52.2
52.1
40.2
+ boxes-first prompting
2.9
9.6
78.6
17.1
18.5
12.4
EtA-SFT
93.7
82.6
65.4
73.0
65.9
51.2
EtA-SFT + ECPO
97.9
81.0
70.9
75.6
72.7
55.5
TABLE VII: Consistency controls on the standard SFT (answer-then-ground) checkpoint. Count rewriting replaces each stated count with the number of emitted boxes, leaving the boxes and thus all grounding metrics unchanged. Boxes-first prompting reorders the output at inference time without retraining.
Fig. 4: Qualitative comparison on RoadSceneVQA-G. Top: for a direction question, ECPO localizes the white van and answers correctly, Ferret emits no box and a wrong answer, and Groma localizes the van but answers incorrectly. Bottom: for a multi-object question, ECPO enumerates all three waiting vehicles and names them in its answer, Groma misses the white car, and Ferret emits 24 boxes.
Method
GT
F1
GDA
CntAcc
R@7+
Standard SFT
–
52.2
88.8
52.1
6.7
EtA-SFT
–
73.0
88.0
65.9
34.8
Vanilla GRPO [ 36 ]
–
73.7
95.1
72.0
36.8
Perception-R1 [ 62 ]
✓
75.6
95.2
72.3
37.2
VisionReasoner [ 63 ]
✓
75.6
95.0
72.0
37.1
ECPO (ours)
×
75.6
95.6
72.7
38.1
TABLE VIII: Comparison of reinforcement learning methods on RoadSceneVQA-G. All methods start from the same EtA-SFT checkpoint. The first two rows are the SFT references. GT: whether the reward uses per-box ground-truth annotations. R@7+: recall on samples with ≥7 objects.
Configuration
F1
Rec
GDA
CntAcc
Vanilla GRPO
73.7
69.8
95.1
72.0
+ group-union coverage
75.6
70.9
95.6
72.7
+ SPC only
75.4
70.5
94.7
71.6
+ union + SPC
75.9
71.3
95.9
73.0
+ union + GT recall
76.0
71.4
96.1
73.2
TABLE IX: Ablation of the ECPO reward terms. Each row modifies the reward of vanilla GRPO. Group-union coverage accounts for the gain. The last row adds a per-box ground-truth recall term and serves as a supervised reference.
Fig. 5: Qualitative comparison on gRefCOCO. Columns: ground truth, InternVL3.5-1B [ 54 ] , LLMDet [ 69 ] , and EtA+ECPO (ours). Top: a no-target expression, for which InternVL3.5-1B and LLMDet each output a spurious box, whereas our model correctly returns an empty set. Middle and bottom: collective multi-target expressions, for which our model enumerates the referred instances with tight boxes, whereas the baselines under-enumerate or leak onto visually similar distractors.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 6: Annotation workflow of RoadSceneVQA-G. (1) Candidate frames and their existing box annotations are loaded from the source dataset. (2) An LLM drafts candidate QA pairs. (3) Annotators edit each question and answer and select the evidence boxes in a dedicated interface. (4) Two reviewers reach consensus before an item enters the dataset.
Test
Train
All
Answer correct
100.0%
100.0%
100.0%
Evidence boxes correct
100.0%
100.0%
100.0%
Both correct
100.0%
100.0%
100.0%
Count matches box cardinality
100.0%
100.0%
100.0%
Appendix
TABLE X: Independent MLLM audit of annotation quality on 5,000 randomly sampled items. Given the calibrated detection rates of the judge (Table XVIII ), these rates are a conservative floor on top of the human consensus review rather than a standalone guarantee.
Category
Question
Answer
Boxes
Existence
Is there any pedestrian crossing the street?
Yes
7
Enumeration
How many cars are waiting at the intersection?
3
3
Attribute
What is the color of the car closest to the camera?
Black
1
Scene inference
What is the weather condition in the image?
Overcast
0
Appendix
TABLE XI: Representative items from the RoadSceneVQA-G test split. Boxes: number of evidence boxes bound to the answer. An empty set (0) marks a no-ground question.
Model
Setting
CIDEr
BERT-F1
Prec
Rec
F1
AP
GDA
GDA corr.
CntAcc
CntSup
Time
InternVL3.5-8B-HF [ 54 ]
zero-shot
76.5
89.4
40.4
33.9
36.9
13.0
60.0
2805/4677
60.0
25
–
InternVL3.5-8B-HF [ 54 ]
LoRA SFT
123.6
91.8
56.5
8.4
14.6
9.1
25.5
1193/4677
58.7
579
1.36 s
Qwen3.5-27B [ 55 ]
zero-shot, thinking
37.0
86.1
37.1
72.9
49.1
26.1
79.5
3719/4677
82.3
277
–
Qwen3.5-27B [ 55 ]
zero-shot, non-thinking
55.0
88.8
61.8
68.1
64.8
34.5
78.7
3682/4677
80.1
181
–
MiniCPM-o 4.5 [ 73 ]
zero-shot, thinking
45.7
85.7
9.0
2.5
3.9
0.4
45.8
2144/4677
38.5
143
–
MiniCPM-o 4.5 [ 73 ]
zero-shot, non-thinking
70.9
88.2
6.3
3.4
4.4
0.3
63.4
2964/4677
32.0
1265
–
Appendix
TABLE XII: Detailed zero-shot and LoRA-SFT MLLM baselines on RoadSceneVQA-G. Text metrics improve after standard LoRA SFT, but grounding recall and F1 often collapse because the models are not required to read the answer out of the box set. GDA corr.: raw number of correct grounding decisions. CntAcc is reported together with its support CntSup. Time: inference latency per sample.
Model
BLEU-4
CIDEr
ROUGE-L
METEOR
BERT-F1
Prec
Rec
F1
AP
GDA
SPC
CntAcc
GCA-s
GCA-l
Ferret-7B [ 22 ]
9.1
97.4
33.5
17.6
92.8
45.1
9.5
15.7
7.8
36.2
23.3
31.6
14.3
19.6
Shikra-7B [ 20 ]
10.4
110.3
34.6
20.0
92.4
56.8
10.9
18.3
14.3
42.3
28.1
37.8
15.1
22.5
Groma-7B [ 61 ]
12.6
131.9
35.9
25.1
92.6
66.9
57.4
61.8
39.5
95.4
97.8
71.6
25.3
46.1
EtA-SFT (ours)
13.9
146.7
39.7
26.6
92.6
82.6
65.4
73.0
54.4
88.0
93.7
65.9
29.6
51.2
ECPO (ours)
14.0
147.0
40.0
26.8
92.7
81.0
70.9
75.6
60.8
95.6
97.9
72.7
32.3
55.5
Appendix
TABLE XIII: Grounded MLLM baselines on RoadSceneVQA-G. All models are fine-tuned on the same training split with their native box syntax. GCA is reported in its strict (GCA-s) and lenient (GCA-l) variants.
Model
Runtime
Checkpoint
Pred. boxes
Prec
Rec
F1
AP 50:95
GDA
GDA corr.
Time
GroundingDINO-T [ 59 ]
official
Swin-T OGC
3,680
31.0
14.7
19.9
7.2
29.9
1398/4677
0.194 s
GroundingDINO-T [ 59 ]
MMDetection, zero-shot
Swin-T OGC
10,566
42.1
57.1
48.5
26.5
79.5
3716/4677
0.190 s
GroundingDINO-T [ 59 ]
MMDetection, fine-tuned
epoch 12
8,965
67.9
78.2
72.7
62.7
86.3
4038/4677
0.177 s
YOLO-World-v2-L 640 [ 60 ]
zero-shot
Obj365+GoldG
12,797
0.1
0.2
0.2
0.0
85.3
3988/4677
0.037 s
YOLO-World-v2-L 640 [ 60 ]
fine-tuned
epoch 24
5,730
23.3
17.2
19.8
3.0
51.4
2405/4677
0.037 s
DINO-X [ 58 ]
official API
–
16,180
29.7
61.8
40.2
27.9
76.5
3577/4677
1.934 s
Appendix
TABLE XIV: Visual grounding specialists on RoadSceneVQA-G. These methods are evaluated only as box predictors with the same QA-derived referring expressions as input. They produce no VQA answer, so answer-dependent metrics are not applicable.
Quantity
Value
Object share with s=0 / 1 / ≥k (%)
22.1 / 14.9 / 63.0
Marginal per-rollout detection rate p^
0.416
Union recall RU
77.9%
Independence prediction 1−(1−p^)G
96.0%
Consensus P / R / F1 (overall)
75.8 / 62.3 / 68.4
Consensus P / R / F1 (7+ bucket)
79.8 / 32.1 / 45.8
Appendix
TABLE XV: Rollout correlation and consensus-set quality, measured with G=6 rollouts on 2,000 test prompts. Support s counts the rollouts that recover each ground-truth object.
GDA
SPC
CntAcc
GCA-l
Method
raw
wtd
raw
wtd
raw
wtd
raw
wtd
Standard SFT
88.8
84.2
26.6
58.0
52.1
60.6
46.3
59.6
EtA-SFT
88.0
84.2
93.7
95.2
65.9
67.5
51.2
62.0
ECPO
95.6
92.9
97.9
97.2
72.7
78.4
55.5
68.9
Appendix
TABLE XVI: Importance-weighted test metrics. Each test sample is weighted by the ratio of its capability-category share between the training and test splits (raw: unweighted, wtd: weighted). Box-level F1 is computed over all matched pairs and is unaffected by per-sample weights, so it is omitted.
F1
GDA
SPC
CntAcc
GCA-l
Tier
Category
n
Std
EtA
ECPO
Std
EtA
ECPO
Std
EtA
ECPO
Std
EtA
ECPO
Std
EtA
ECPO
Perception
Existence
509
32.2
50.4
64.8
48.1
52.7
75.0
76.8
95.2
96.0
22.4
33.4
57.0
35.8
43.4
62.9
Enumeration
1344
33.1
70.2
71.3
90.8
91.5
98.0
15.9
93.5
99.0
7.7
43.9
46.1
26.5
42.2
44.0
Attribute recognition
1440
87.8
87.1
90.2
96.9
94.4
99.0
50.0 †
83.3 †
40.0 †
94.2
91.6
98.0
78.5
77.4
81.1
Object identification
476
54.9
68.2
70.1
98.1
88.0
96.2
5.3
89.3
82.9
59.4
65.5
70.8
13.7
14.7
15.8
Scene inference
182
–
–
–
100.0
100.0
100.0
–
–
–
–
–
–
82.4
80.8
81.9
Appendix
TABLE XVII: Per-category metrics on the test split (Std: standard SFT, EtA: EtA-SFT, ECPO: EtA-SFT + ECPO). Scene inference has no ground-truth boxes, so its grounding and consistency metrics are undefined. SPC on attribute recognition ( † ) has a small support because such answers rarely state a count.
Perturbation
Qwen3-VL-Plus
Qwen3-VL-235B
Wrong evidence boxes
26.7%
51.1%
Wrong count in answer
50.0%
47.4%
Extra spurious box
83.3%
83.1%
Appendix
TABLE XVIII: Calibration of the MLLM judge: detection rate of injected errors on 200 perturbed items for two judge scales of Qwen3-VL [ 24 ] .
Fig. 7: Qualitative comparison on CARPK. Boxes predicted by EtA-SFT, vanilla GRPO, and ECPO. ECPO recovers the most complete box set.
In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language Models (MLLMs) tend to over-attend to backgrounds, overwhelming crucial small objects during visual-language alignment, a failure mode we term 'critical evidence dilution.' Furthermore, existing visual question answering (VQA) datasets rarely expose this flaw, as they lack large-scale, distractor-heavy evaluations that require pinpointing local evidence. To bridge this evaluation and architecture gap, we introduce the Fine-Grained Traffic Reasoning Benchmark (FGTR-Bench) and the Text-Guided Small-Object Reasoning MLLM (TSR-MLLM). FGTR-Bench comprises 40,236 single-image Multiple-Choice Questions (MCQs) created via multi-agent generation, consistency checks, and expert audits, alongside a disjoint 4,947-sample blind test split. To resolve evidence dilution, TSR-MLLM, built on Qwen3-VL-4B, uses a query-conditioned Text-Guided Small-Object Focus (TG-SOF) map. Applied once at the decoder boundary, the map adds sparse Top-K gated residuals to the most question-relevant vision slots while leaving text tokens unchanged. Together with lightweight decoder adaptation, TSR-MLLM preserves single-pass inference without external detectors or image re-encoding. Under matched settings, TSR-MLLM outperforms the strongest 4B baseline by 2.1 points on FGTR-Bench (74.1% overall), with larger gains on evidence-local tracks. Furthermore, it remains competitive on DriveQA-V (CARLA Signs) under greedy decoding without task-specific fine-tuning.
Waikit Xiu, Qiang Lu, Zian Wang +4
1The University of Hong Kong · 2Sun Yat-Sen University · 3The Hong Kong University of Science and Technology (Guangzhou)
Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VISTAQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VISTAQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VISTAQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.
Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5
University of Waterloo · Stanford University · NVIDIA
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers should depend on scene-specific visual evidence but may instead be inferred from textual shortcuts. Through an audit of four public benchmarks, we find that several recent open-weight Vision-Language Models (VLMs) perform competitively, and sometimes better, without video input. On the MM-AU benchmark, removing video consistently improves accuracy, and adding more frames further degrades performance. To quantify visual dependence, we introduce two dataset-level diagnostics: Blind Gap, measuring above-chance text-only performance, and Visual Gain, measuring the marginal benefit of adding video. We further propose an instance-level Shortcut Score that combines text-only confidence with visual necessity signals, enabling continuous, training-free filtering of shortcut-prone questions. The resulting subsets reduce shortcut bias and improve visual grounding. Our findings reveal large differences in grounding quality across benchmarks and show that visually grounded evaluation, not just high accuracy, is essential in safety-critical VideoQA.
Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim +1
Technical University of Munich, Munich, Germany · Helmholtz Zentrum München, Munich, Germany · Munich Center for Machine Learning, Munich, Germany