Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94% of VQS answers correct, against 76% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS
Figures & tables
Figure 1: VQS improves every benchmark. Scores are for Qwen3-VL-2B. VISE loses on LogicVista and MMStar.
Figure 2: Majority-vote labels vs. computed answers. (a) In prior self-play, the solver’s majority answer becomes the label. Here three of five answers say 2 mugs, so the wrong answer is rewarded and the correct answer “1” is discarded. (b) VQS computes the answer instead. (1) A schema-constrained parser turns the image into a structured parse. A fixed template reads the parse (filter mugs, then count) to write a question and compute its answer “1” and the parser then paraphrases the question. (2) A visual checker verifies each claim behind the answer, a blind gate drops questions answerable without the image, and a difficulty band keeps those the solver gets right in some but not all of 8 rollouts. (3) GRPO rewards each rollout by exact match against the computed answer, so training reinforces the correct answer “1” and suppresses the wrong one “2” .
Figure 3: Label-free parser training. (1) The parser samples K structured parses of an unlabeled image under a schema-constrained decoder. (2) A visual checker verifies each parse one claim at a time. In the example, z2 wrongly calls the mug blue and places it to the right of the book, so it passes 6 of 8 claims. (3) Each parse is scored by its fraction of verified claims. We keep the best parse z1 (8/8) only because it beats the mean of all four parses (0.75) and contains at least m=4 claims. (4) The kept parses train the parser with SFT. The trained parser then feeds the templates that write questions and compute their answers for the solver.
Figure 4: Examples of structured scene parsing. The parser converts natural images, charts, diagrams, and infographics into domain-specific JSON records that preserve spatial relations, chart values and ticks, graph structure, and verbatim strings for deterministic QA construction.
Method
GQA
OK-VQA
InfoVQA
SQA
MMMU
MMB
ESB
LogicV
MMStar
SEED
Avg.
Qwen3-VL-2B-Instruct
Base
58.25
40.76
69.02
79.42
38.92
74.48
68.54
35.04
55.62
72.16
59.22
VisPlay
58.65 +0.40
41.15 +0.39
69.96 +0.94
80.74 +1.32
39.27 +0.35
74.52 +0.04
68.56 +0.02
34.93 -0.11
55.37 -0.25
71.62 -0.54
59.48 +0.26
Vision-Zero
58.98 +0.73
41.39 +0.63
70.93 +1.91
81.96 +2.54
39.58 +0.66
75.07 +0.59
69.72 +1.18
35.28 +0.24
55.58 -0.04
71.53 -0.63
60.00 +0.78
EvoLMM
59.01 +0.76
38.03 -2.73
70.69 +1.67
83.01 +3.59
39.08 +0.16
74.62 +0.14
69.32 +0.78
34.99 -0.05
55.50 -0.12
71.17 -0.99
59.54 +0.32
iReasoner
59.13 +0.88
38.13 -2.63
70.82 +1.80
83.12 +3.70
39.11 +0.19
74.75 +0.27
69.67 +1.13
35.09 +0.05
55.59 -0.03
71.25 -0.91
59.67 +0.45
Table 1: Main results across three scales. Every model trains on unlabeled images only, with no human question, answer or gold label at any stage. Subscripts give the change against that backbone’s own base, in green when positive and red when negative. VQS gives the largest average gain at all three scales , +3.18 at 2B, +2.32 at 4B and +2.43 at 8B, against +1.01, +1.04 and +0.66 for the strongest published method. InfoVQA stands for infographicsVQA, SQA for ScienceQA, MMB for MMBench, ESB for EmbSpatial, and LogicV for LogicVista.
Figure 5: What scale and coverage change. (a) Share of benchmark answers RL changes, measured per sample against each model’s base. (b) Smoothed training reward. (c) 2B gain over base against the number of template families.
Answer source
Correct (%)
Kept (%)
Templates + checker
94.4%
76.0
Majority vote
76.4%
92.0
Model judge
82.2%
90.8
Table 2: Human evaluation. Correct is accuracy among retained questions.
Figure 7: Parser target selection: quality and coverage. Each panel varies one setting around (K,δ,m)=(4,0.10,4) . Parse precision is the percentage of correct visual claims after training. Images retained is the percentage of source images yielding a selected training target.
Figure 8: Gains across cycles. Average gain over base per benchmark category for Qwen3-VL-2B.
Pipeline
Solver Acc.
Parser Prec.
Untrained parser
58.0
80.8
Full VQS
62.4
87.6
✗ constrained decoding
61.9
85.3
✗ fact-checker
60.9
87.6
✗ blind gate
61.8
87.6
✗ difficulty band
62.1
87.6
Table 4: Component ablations at 2B.
Rank
Family
Gain
▲ 1
Chart differences
7.4
▲ 2
Diagram connectivity
5.8
▼ 1
Object existence
-0.2
▼ 2
Attribute lookup
0.4
Table 5: Per-family gains. The two largest and two smallest.
Comparison
Setting
Solver acc.
Family difficulty
Easy
60.8
Middle
62.4
Hard
61.1
Generator
Fixed templates
62.4
Generated programs
61.4
Curriculum
Shuffled
61.3
Table 6: Question design at 2B. Each block changes one choice at an equal training budget. Solver accuracy averages the 10 benchmarks in Table 1 . Bold marks the VQS default.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Adapter
LoRA rank, alpha
64, 128
Target modules
all linear layers, .visual. excluded
Trainable visual parameters
0.00M
Gradient checkpointing
on
Optimiser
Appendix
Table 7: Solver training settings. Identical across 2B, 4B, and 8B except for the per-device micro-batch.
Category
Benchmark
What it tests
Perception
GQA
objects, attributes, and relations
InfoVQA
text and layout in infographics
EmbSpatial (ESB)
spatial relations in embodied scenes
Knowledge & Reasoning
OK-VQA
outside knowledge
ScienceQA (SQA)
school science
MMMU
college-level subjects
Appendix
Table 8: Benchmark categories. Each benchmark belongs to one category.
Checker
Relation to the parser
Accepts (%)
Trap YES (%)
κ vs. ours
Qwen3-VL-2B (ours)
same weights
94.4
1.2
–
Qwen3-VL-8B
same family, 4 × larger
93.3
1.6
0.837
InternVL3.5-8B
different family, 8B
93.1
4.0
0.810
Qwen2.5-VL-7B
previous Qwen generation
89.5
0.4
0.807
SmolVLM2-2.2B
different family, same size
86.1
5.7
0.660
Appendix
Table 9: Checker independence. All five checkers grade the same 1,417 claims. Accepts is the share of the 1,170 parser claims a checker calls true. Trap YES is the share of the 247 corrupted claims it wrongly accepts. κ is Cohen’s kappa with our checker over all claims.
Method
GQA
OK-VQA
InfoVQA
SQA
MMMU
MMB
ESB
LogicV
MMStar
SEED
Avg.
Gemma3-12B-It
Base
61.20
62.40
52.10
94.30
46.50
76.40
24.50
41.80
58.90
74.20
59.23
VQS (ours)
62.02 +0.82
62.71 +0.31
55.29 +3.19
95.23 +0.93
50.84 +4.34
79.06 +2.66
27.70 +3.20
42.87 +1.07
59.80 +0.90
74.47 +0.27
61.00 +1.77
InternVL3-8B-Instruct
Base
65.40
64.10
68.80
96.50
51.20
81.30
28.10
48.60
63.70
76.90
64.46
VQS (ours)
66.01 +0.61
64.36 +0.26
71.81 +3.01
96.71 +0.21
54.87 +3.67
83.17 +1.87
31.99 +3.89
49.42 +0.82
64.32 +0.62
77.08 +0.18
65.97 +1.51
Appendix
Table 10: Results on other model families. Every model trains on unlabeled images only. Subscripts give the change against that backbone’s base. Abbreviations follow Table 1 . VQS improves all ten benchmarks for every family , raising the average by +1.77 on Gemma3-12B-It, +1.51 on InternVL3-8B-Instruct and +1.84 on Llama-3.2-11B-Vision-Instruct.
Family
Lv
Template → example question
Answer key
Photos
existence
L2
Are there any ⟨ objects ⟩ in the image? → Are there any windows in the image?
no
relation_existence
L2
Are there any ⟨ objects ⟩⟨ relation ⟩ the ⟨ anchor ⟩ ? → Are there any chairs to the right of the painting?
yes
attribute_filtered_existence
L3
Is there a ⟨ attribute ⟩⟨ object ⟩ in the picture? → Is there white grass in the picture?
no
state_filtered_existence
L3
Do you see a ⟨ state ⟩⟨ object ⟩ in the picture? → Do you see an open oven door in the picture?
yes
gated_object_recognition
L3
Which kind of ⟨ category ⟩ is ⟨ attribute ⟩ ? → Which kind of furniture is green?
bench
Appendix
Table 11: Template families, with one question each. A slot in ⟨ angle brackets ⟩ is filled from the parse, and a part in [ square brackets ] appears only when the chart has several series. Each example is a question the family wrote for a real image, shown after paraphrasing where the paraphrase passed the guard, with the key its program computed. We checked every key against its image.
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high-quality training data or the cost of curating them. We show that a VLM can continuously improve itself as a visual questioner without any external supervision. We propose a self-evolving framework that uses a VLM itself as both a proposer and a filter to produce harder, more informative, and visual-centric questions, while maintaining their exploration diversity to avoid training collapse. These questions are then used to train the VLM in both questioner and answerer modes. To evaluate the questioner, we introduce an agentic protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across various backbone VLMs show that our method substantially enhances the quality and substantially expands the difficulty boundary of autonomous question generation. Under the same budget, our self-supervision is more effective than training on the static source data. Moreover, the self-evolving questioner remains a competitive or even better answerer.
Yijun Liang, Hengguang Zhou, Ming Li +3
University of Maryland, College Park · University of California, Los Angeles · Peking University +2
Vision-language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large-scale human-constructed supervision for post-training. Such supervision is costly to obtain, especially for reasoning-intensive multimodal tasks where questions, answers, and feedback signals must be carefully designed. This motivates self-evolving learning, where a model improves itself through a dual-role closed loop: a questioner autonomously poses questions and a solver learns to solve them. However, we observe that current VLM self-evolving methods still face three major challenges: coarse-grained role alternation delays the interaction between question generation and solver adaptation; generated questions can progressively degrade in quality; and question types may collapse toward a narrow distribution. These issues limit the efficiency and reliability of self-evolution. Thus, we propose \textbf{RISE}, a reliable self-evolving framework for vision-language models. RISE is built on three complementary designs: fine-grained role alternation, which shortens the feedback loop between the questioner and the solver to improve efficiency; a quality supervisor, which improves question validity and pseudo-label reliability; and skill-aware dynamic balancing, which mitigates mode collapse and maintains broad skill coverage during evolution. Together, these components enable more reliable and effective self-evolution from unlabeled images. Experiments on two VLM backbones across seven benchmarks show that RISE consistently improves the base models, yielding broad and sustained gains. Our code is publicly available at https://github.com/AMAP-ML/RISE.
Manual annotation of high-quality visual question answering with grounding (VQA-G) datasets, which pair visual questions with evidential grounding, is crucial for advancing vision-language models (VLMs), but remains unscalable. Existing automated methods are often hindered by two key issues: (1) inconsistent data fidelity due to model hallucinations; (2) brittle verification mechanisms based on simple heuristics. To address these limitations, we introduce AutoVQA-G, a self-improving agentic framework for automated VQA-G annotation. AutoVQA-G employs an iterative refinement loop where a Consistency Evaluation module uses Chain-of-Thought (CoT) reasoning for fine-grained visual verification. Based on this feedback, a memory-augmented Prompt Optimization agent analyzes critiques from failed samples to progressively refine generation prompts. Our experiments show that AutoVQA-G generates VQA-G datasets with superior visual grounding accuracy compared to leading multimodal LLMs, offering a promising approach for creating high-fidelity data to facilitate more robust VLM training and evaluation. Code: https://github.com/rohnson1999/AutoVQA-G
Rongsheng Hu, Runwei Guan, Yicheng Di +2
School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, China · Thrust of Artificial Intelligence, HKUST(GZ), Guangzhou, China