Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundry-10k consistently improves perception benchmarks across three open-source backbones (e.g., +6.7% on MMVP-pair and +10.5% on CV-Bench-3D for Qwen2.5-VL-3B-Instruct) while preserving broader capabilities and showing positive data scaling. The same synthetic supervision also yields consistent gains under reinforcement learning (RL) across all three backbones, and the framework remains effective under open-source synthesis and self-verification. Our findings demonstrate that automated synthetic supervision offers an effective and scalable path toward systematic VLM training.
Figures & tables
Figure 1: Training VLMs with synthetic images. VisionFoundry uses a task-keyword-only pipeline that requires no reference text or images: it generates task-aware T2I prompts and question–answer pairs, synthesizes images, and uses the resulting supervision to improve capabilities such as visual perception in VLMs. Compared with traditional web-image collection pipelines, this approach provides more controllable and task-targeted supervision.
Figure 2: VisionFoundry overview. Given only task keywords, an LLM builds an adaptive concept pool; compositional sampling forms entities for T2I prompts; and a T2I model synthesizes images that a frontier VLM verifies to yield a high-quality VQA dataset.
Figure 3: VisionFoundry-10K examples. Randomly selected examples from our dataset, covering all 10 tasks. The task names are annotated at the top and serve as the only valid input to the pipeline. Each panel shows a generated image, its corresponding question, and a ground-truth answer.
Benchmark
Qwen2.5-VL-3B Instruct
MiMo-VL-7B SFT
Llama-3.2-11B Vision-Instruct
Baseline
Synth
Baseline
Synth
Baseline
Synth
Visual Perception Benchmarks
MMVPpair ( Tong et al., 2024b )
35.3
42.0
43.3
57.3
42.7
46.7
MMVPsingle ( Tong et al., 2024b )
64.3
68.3
66.7
77.7
70.3
71.7
CV-Bench-2D ( Tong et al., 2024a )
67.3
72.4
74.3
79.0
70.4
71.7
CV-Bench-3D ( Tong et al., 2024a )
66.0
76.5
72.3
83.7
74.4
75.3
Table 1: Main Benchmark Results. Main results on 13 benchmarks with 15 reported metrics across three VLMs. Models trained on VisionFoundry-10K improve visual perception performance while maintaining general capabilities. We report the mean score over four independent runs. Visual perception benchmarks are highlighted at the top.
Model
MMVPpair
MMVPsingle
CV-Bench-2D
CV-Bench-3D
RealWorldQA
Qwen2.5-VL-3B-Instruct
46.0 (+10.7)
71.3 (+7.0)
72.9 (+5.6)
72.3 (+6.3)
67.2 (+2.2)
Llama-3.2-11B-Vision-Instruct
45.3 (+2.6)
70.7 (+0.4)
72.4 (+2.0)
76.3 (+1.9)
63.8 (+0.8)
MiMo-VL-7B-SFT
48.0 (+4.7)
73.3 (+6.6)
78.7 (+4.4)
75.8 (+3.5)
69.2 (+3.3)
Table 2: RL results on visual perception benchmarks. Evaluation of GRPO across three VLM backbones on the 10-task VisionFoundry-10K dataset. Numbers in parentheses indicate absolute gains over each model’s baseline.
Figure 4: Data-Size Effects of Synthetic Supervision. We report performance trends as the amount of VisionFoundry synthetic VQA supervision increases. “10k” corresponds to the full VisionFoundry-10K dataset (10k samples), while “500/1k/2k/5k” are random subsets sampled from it. The training recipe matches the main experimental setting. Overall, results show an upward trend on visual perception benchmarks as synthetic data increases.
Figure 5: Equal-sized mixture vs. pure natural data. We compare training on an equal-sized mix of VisionFoundry-10K samples and natural data against training on a size-matched natural-only subset. In the figure, the bold-bordered first row denotes visual perception benchmarks, while the second row reports general-purpose benchmarks. The mixed setting yields consistently higher accuracy on visual perception benchmarks while maintaining comparable performance on general-purpose evaluations.
Figure 6: Epoch trade-off. The first row uses 1k samples from one randomly selected task, while the second row uses the full VisionFoundry-10K . With only a single-task 1k subset, performance typically converges after around 8 epochs. With the full 10-task set, convergence is reached with fewer training epochs. The dashed line indicates the baseline.
Figure 7: Task-wise gains and cross-benchmark divergence. (a): Visualization of five representative task-specific 1k subsets; complete results for all ten task-specific models are reported in Appendix Tables 7 and 8 . Darker colors indicate larger gains, while lighter colors indicate smaller gains or regressions. (b): Per-benchmark response divergence across tasks, where a larger top-to-bottom accuracy span indicates stronger task sensitivity. While visual perception benchmarks generally benefit from synthetic supervision, the magnitude and direction of gains remain benchmark-dependent.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Epochs
Batch
ηViT
ηadapter
ηLLM
LLM
Qwen2.5-VL-3B-Instruct
1
128
5×10−7
5×10−6
5×10−6
Unfrozen
MiMo-VL-7B-SFT
1
128
5×10−7
2.5×10−6
2.5×10−6
Unfrozen
Llama-3.2-11B-Vision-Instruct
1
128
5×10−7
5×10−6
N/A
Frozen
Appendix
Table 3: Main-experiment hyperparameters. Training settings for Qwen2.5-VL-3B-Instruct, MiMo-VL-7B-SFT, and Llama-3.2-11B-Vision-Instruct in the primary comparison (all runs use 1 epoch and global batch size 128). ηViT , ηadapter , and ηLLM denote learning rates for the vision encoder, adapter modules, and LLM backbone, respectively; LLM indicates whether the LLM backbone is frozen or unfrozen.
Model
Trainable modules
Reward model
Epoch
LR
Batch
Rollouts
Details
Qwen2.5-VL-3B-Instruct
Full model
Qwen2.5-3B
1
4×10−6
128
4
GRPO, step 70
Llama-3.2-11B-Vision-Instruct
Vision + adapters
Qwen2.5-3B
1
1×10−6
32
4
GRPO, step 281
MiMo-VL-7B-SFT
Full model
Qwen2.5-3B
1
1×10−6
64
4
GRPO, step 140
Appendix
Table 4: RL hyperparameters. Experimental configurations for GRPO post-training on VisionFoundry-10K . All runs use a KL coefficient of 0.01 and four rollouts per prompt.
Figure 8: Confusion matrix for the random-sampling audit. Rows are manual correctness labels, and columns are verifier outputs. Values summarize the audited trajectory.
Split
Accuracy
Precision
Recall
F1
Overall audited set
92.1
99.0
90.8
94.7
Appendix
Table 5: Verifier performance on the audited sample set. Accuracy, precision, recall, and F1 are computed by comparing verifier pass/fail outputs against manual labels over all audited candidates.
Metric
Definition
Value
Final valid proportion
generated correct & filter pass over all audited candidates
70.7%
Acceptance rate
retained / generated candidates
71.4%
Rejection rate
rejected / generated candidates
28.6%
False accept rate
generated wrong but filter pass / generated wrong candidates
3.2%
False reject rate
generated correct but filter fail / generated correct candidates
9.2%
Human–judge agreement
Cohen’s κ on audited subset
0.794
Appendix
Table 6: Pipeline quality metrics on the audited trajectory. The table reports final valid proportion, acceptance/rejection rates, false accept/reject rates, and human–judge agreement (Cohen’s κ ).
Task
MVP-S
MVP-P
CV2D
CV3D
RWQA
BLK
MMS
OCR
OaD
64.0
36.0
68.4
73.9
66.4
49.1
55.8
82.4
VaP
65.0
36.0
68.4
73.4
65.6
48.9
55.8
81.8
PaRC
66.0
38.0
67.9
73.5
66.1
49.5
55.4
82.2
SR
64.3
36.0
67.9
72.7
66.1
49.0
55.3
82.3
SaC
64.7
36.7
68.3
71.5
65.6
48.9
55.4
82.0
SPC
65.0
36.0
67.9
73.7
66.0
48.5
55.7
82.4
Appendix
Table 7: Task-wise benchmark performance across 10 task-specialized models (Part I). Scores (%) on MVP-S, MVP-P, CV2D, CV3D, RWQA, BLK, MMS, and OCR; higher is better.
Task
LEGO
MMMU
MMB
MVis
SSP
MMSI
3DSR
OaD
15.7
50.0
76.2
64.2
20.7
8.8
35.2
VaP
16.3
49.0
75.9
65.2
20.5
8.7
35.2
PaRC
16.4
48.8
76.8
64.6
20.6
9.4
35.3
SR
16.0
48.1
76.2
63.3
20.2
9.4
34.9
SaC
16.1
48.8
76.1
64.4
20.7
9.7
35.2
SPC
15.8
48.4
76.3
63.9
20.4
9.5
35.3
Appendix
Table 8: Task-wise benchmark performance across 10 task-specialized models (Part II). Scores (%) on LEGO, MMMU, MMB, MVis, SSP, MMSI, and 3DSR; higher is better.
Figure 9: Verification is helpful under a matched 1k-sample training budget. The figure compares baseline, finetuning without verification, and finetuning with verification ( VisionFoundry ) on representative benchmarks. The verified branch performs better than the non-verified variant, with the clearest gains on visual perception evaluations.
Figure 10: Natural images with synthetic QA vs. synthetic images with synthetic QA under matched 1k-data training. With the same model and setup, the synthetic-image setting shows stronger overall performance, with particularly clear gains on visual perception benchmarks while remaining comparable on the rest.
Figure 11: Strict control experiment with matched synthetic QA. Natural images and caption-conditioned synthetic images are compared under identical synthetic QA supervision and finetuning setup. Synthetic images achieve consistently stronger performance, with the largest gain on CV-Bench-3D.
Figure 12: Open-source vs. proprietary T2I models under matched 1k-data training. Replacing the proprietary image generator in VisionFoundry with the open-source Qwen-Image-2512 still yields consistent improvements over the no-finetuning baseline. The proprietary T2I model remains slightly stronger overall, but the open-source variant retains most of the benefit, especially on visual perception benchmarks.
Figure 13: Open-source vs. proprietary verifiers under matched 1k-data training. With the T2I model fixed to Qwen-Image-2512, replacing Gemini-3-Pro with the open-source Qwen2.5-VL-3B-Instruct verifier still yields strong downstream gains. The two verifier choices trade wins across benchmarks rather than showing a one-sided dominance, suggesting that VisionFoundry ’s improvements are not simply inherited from the verifier ceiling.
Setting
MMVPpair
MMVPsingle
CV-Bench-2D
CV-Bench-3D
RealWorldQA
Qwen2.5-VL-3B baseline
35.3
64.3
67.3
66.0
65.0
Self-verifier 10K SFT
40.7
68.7
69.3
75.2
65.9
Improvement
+5.4
+4.4
+2.0
+9.2
+0.9
Appendix
Table 9: Full-scale self-verification on visual perception benchmarks. Performance of Qwen2.5-VL-3B-Instruct trained on the 10K self-verified VisionFoundry variant, where both image synthesis (Qwen-Image-2512) and verifier (Qwen2.5-VL-3B-Instruct) are performed using open-source models.
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.
Matteo Farina, Vishaal Udandarao, Thao Nguyen +33
University of Trento · 2Tübingen AI Center, University of Tübingen · University of Cambridge +12
Achieving human-like reasoning in Vision-Language Models (VLMs) remains a long-standing challenge. Recent approaches leverage Chain-of-Thought (CoT) rationales generated by human annotators or proprietary models, which are costly and difficult to scale. Self-training offers a promising alternative but often suffers from visual hallucinations and language shortcuts because rationales are filtered only by answer correctness without verifying visual perception. We propose a perception-verified self-training framework that enforces visually grounded reasoning. Our method employs a CoT template (caption-reasoning-conclusion) that disentangles perception from reasoning, enabling independent verification of visual understanding. To compensate for the absence of ground-truth captions, we introduce PerceptEval, an unsupervised method that evaluates caption quality based on its alignment with visual and textual elements in the image. Using caption verification together with answer correctness, we partition the data into easy, medium, and hard subsets and design a two-stage curriculum learning strategy. Stage 1 trains on easy samples, while Stage 2 enhances medium samples by regenerating reasoning conditioned on verified captions and retaining only those with correct conclusions. This ensures training is performed exclusively on perceptually grounded reasoning, reducing hallucinations and language shortcuts. Extensive experiments across diverse domains and models demonstrate improvements of up to 16% over standard self-training baselines, showing that our framework provides a scalable and cost-effective solution for advancing multimodal reasoning without manually annotated CoT rationales.
Sourabh Sharma, Sonam Gupta, Sadbhawna
Malaviya National Institute of Technology Jaipur, India · IBM Software Innovation Lab, India
Vision-Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual differences. This paper presents a training-free framework for the 5th DataCV Challenge Task 1 at CVPR 2026, addressing this perception-versus-memory conflict through three complementary strategies:(1) illusion-aware image preprocessing that weakens illusion-inducing context via type-specific transformations (edge extraction, color isolation, morphological processing, and reference-line overlay), (2) anti-illusion prompt engineering guiding VLMs toward qualitative visual comparison, and (3) multi-vote ensemble that further improves robustness. Our method achieves 90.48% accuracy on the official 630-image test set using Claude (claude-opus-4-6) with 5-vote majority ensemble, and 98.41% on a human-verified subset. The approach requires no finetuning, relying solely on visual manipulation and prompt design. Our solution secured 2nd place in the challenge, only 0.47% behind the 1st-place solution. Code is available at https://github.com/jasminezz/sf-illusion-aware-vlm.git.