Program-Verified Self-Evolution for Vision-Language Models
Organizations: LG AI Research · MBZUAI · University of Illinois at Chicago · Australian National University
Abstract
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94% of VQS answers correct, against 76% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS
Figures & tables
| Method | GQA | OK-VQA | InfoVQA | SQA | MMMU | MMB | ESB | LogicV | MMStar | SEED | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-2B-Instruct | |||||||||||
| Base | 58.25 | 40.76 | 69.02 | 79.42 | 38.92 | 74.48 | 68.54 | 35.04 | 55.62 | 72.16 | 59.22 |
| VisPlay | 58.65 +0.40 | 41.15 +0.39 | 69.96 +0.94 | 80.74 +1.32 | 39.27 +0.35 | 74.52 +0.04 | 68.56 +0.02 | 34.93 -0.11 | 55.37 -0.25 | 71.62 -0.54 | 59.48 +0.26 |
| Vision-Zero | 58.98 +0.73 | 41.39 +0.63 | 70.93 +1.91 | 81.96 +2.54 | 39.58 +0.66 | 75.07 +0.59 | 69.72 +1.18 | 35.28 +0.24 | 55.58 -0.04 | 71.53 -0.63 | 60.00 +0.78 |
| EvoLMM | 59.01 +0.76 | 38.03 -2.73 | 70.69 +1.67 | 83.01 +3.59 | 39.08 +0.16 | 74.62 +0.14 | 69.32 +0.78 | 34.99 -0.05 | 55.50 -0.12 | 71.17 -0.99 | 59.54 +0.32 |
| iReasoner | 59.13 +0.88 | 38.13 -2.63 | 70.82 +1.80 | 83.12 +3.70 | 39.11 +0.19 | 74.75 +0.27 | 69.67 +1.13 | 35.09 +0.05 | 55.59 -0.03 | 71.25 -0.91 | 59.67 +0.45 |
| Answer source | Correct (%) | Kept (%) |
|---|---|---|
| Templates + checker | 94.4% | 76.0 |
| Majority vote | 76.4% | 92.0 |
| Model judge | 82.2% | 90.8 |
| Parser target | Parse Prec. | Answer Acc. | Solver Acc. |
|---|---|---|---|
| Untrained | 80.8 | 86.8 | 58.0 |
| Random | 82.1 | 91.0 | 61.5 |
| Whole parse | 85.0 | 94.0 | 62.0 |
| Per claim | 87.6 | 96.1 | 62.4 |
| Pipeline | Solver Acc. | Parser Prec. |
|---|---|---|
| Untrained parser | 58.0 | 80.8 |
| Full VQS | 62.4 | 87.6 |
| ✗ constrained decoding | 61.9 | 85.3 |
| ✗ fact-checker | 60.9 | 87.6 |
| ✗ blind gate | 61.8 | 87.6 |
| ✗ difficulty band | 62.1 | 87.6 |
| Rank | Family | Gain |
|---|---|---|
| 1 | Chart differences | 7.4 |
| 2 | Diagram connectivity | 5.8 |
| 1 | Object existence | -0.2 |
| 2 | Attribute lookup | 0.4 |
| Comparison | Setting | Solver acc. |
| Family difficulty | Easy | 60.8 |
| Middle | 62.4 | |
| Hard | 61.1 | |
| Generator | Fixed templates | 62.4 |
| Generated programs | 61.4 | |
| Curriculum | Shuffled | 61.3 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Adapter | |
| LoRA rank, alpha | 64, 128 |
| Target modules | all linear layers, .visual. excluded |
| Trainable visual parameters | 0.00M |
| Gradient checkpointing | on |
| Optimiser | |
| Category | Benchmark | What it tests |
| Perception | GQA | objects, attributes, and relations |
| InfoVQA | text and layout in infographics | |
| EmbSpatial (ESB) | spatial relations in embodied scenes | |
| Knowledge & Reasoning | OK-VQA | outside knowledge |
| ScienceQA (SQA) | school science | |
| MMMU | college-level subjects |
| Checker | Relation to the parser | Accepts (%) | Trap YES (%) | vs. ours |
|---|---|---|---|---|
| Qwen3-VL-2B (ours) | same weights | 94.4 | 1.2 | – |
| Qwen3-VL-8B | same family, 4 larger | 93.3 | 1.6 | 0.837 |
| InternVL3.5-8B | different family, 8B | 93.1 | 4.0 | 0.810 |
| Qwen2.5-VL-7B | previous Qwen generation | 89.5 | 0.4 | 0.807 |
| SmolVLM2-2.2B | different family, same size | 86.1 | 5.7 | 0.660 |
| Method | GQA | OK-VQA | InfoVQA | SQA | MMMU | MMB | ESB | LogicV | MMStar | SEED | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma3-12B-It | |||||||||||
| Base | 61.20 | 62.40 | 52.10 | 94.30 | 46.50 | 76.40 | 24.50 | 41.80 | 58.90 | 74.20 | 59.23 |
| VQS (ours) | 62.02 +0.82 | 62.71 +0.31 | 55.29 +3.19 | 95.23 +0.93 | 50.84 +4.34 | 79.06 +2.66 | 27.70 +3.20 | 42.87 +1.07 | 59.80 +0.90 | 74.47 +0.27 | 61.00 +1.77 |
| InternVL3-8B-Instruct | |||||||||||
| Base | 65.40 | 64.10 | 68.80 | 96.50 | 51.20 | 81.30 | 28.10 | 48.60 | 63.70 | 76.90 | 64.46 |
| VQS (ours) | 66.01 +0.61 | 64.36 +0.26 | 71.81 +3.01 | 96.71 +0.21 | 54.87 +3.67 | 83.17 +1.87 | 31.99 +3.89 | 49.42 +0.82 | 64.32 +0.62 | 77.08 +0.18 | 65.97 +1.51 |
| Family | Lv | Template example question | Answer key |
|---|---|---|---|
| Photos | |||
| existence | L2 | Are there any objects in the image? Are there any windows in the image? | no |
| relation_existence | L2 | Are there any objects relation the anchor ? Are there any chairs to the right of the painting? | yes |
| attribute_filtered_existence | L3 | Is there a attribute object in the picture? Is there white grass in the picture? | no |
| state_filtered_existence | L3 | Do you see a state object in the picture? Do you see an open oven door in the picture? | yes |
| gated_object_recognition | L3 | Which kind of category is attribute ? Which kind of furniture is green? | bench |