VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
Organizations: Princeton University · New York University
Abstract
Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundry-10k consistently improves perception benchmarks across three open-source backbones (e.g., +6.7% on MMVP-pair and +10.5% on CV-Bench-3D for Qwen2.5-VL-3B-Instruct) while preserving broader capabilities and showing positive data scaling. The same synthetic supervision also yields consistent gains under reinforcement learning (RL) across all three backbones, and the framework remains effective under open-source synthesis and self-verification. Our findings demonstrate that automated synthetic supervision offers an effective and scalable path toward systematic VLM training.
Figures & tables
| Benchmark | Qwen2.5-VL-3B Instruct | MiMo-VL-7B SFT | Llama-3.2-11B Vision-Instruct | |||
|---|---|---|---|---|---|---|
| Baseline | Synth | Baseline | Synth | Baseline | Synth | |
| Visual Perception Benchmarks | ||||||
| ( Tong et al., 2024b ) | 35.3 | 42.0 | 43.3 | 57.3 | 42.7 | 46.7 |
| ( Tong et al., 2024b ) | 64.3 | 68.3 | 66.7 | 77.7 | 70.3 | 71.7 |
| CV-Bench-2D ( Tong et al., 2024a ) | 67.3 | 72.4 | 74.3 | 79.0 | 70.4 | 71.7 |
| CV-Bench-3D ( Tong et al., 2024a ) | 66.0 | 76.5 | 72.3 | 83.7 | 74.4 | 75.3 |
| Model | CV-Bench-2D | CV-Bench-3D | RealWorldQA | ||
|---|---|---|---|---|---|
| Qwen2.5-VL-3B-Instruct | 46.0 (+10.7) | 71.3 (+7.0) | 72.9 (+5.6) | 72.3 (+6.3) | 67.2 (+2.2) |
| Llama-3.2-11B-Vision-Instruct | 45.3 (+2.6) | 70.7 (+0.4) | 72.4 (+2.0) | 76.3 (+1.9) | 63.8 (+0.8) |
| MiMo-VL-7B-SFT | 48.0 (+4.7) | 73.3 (+6.6) | 78.7 (+4.4) | 75.8 (+3.5) | 69.2 (+3.3) |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Epochs | Batch | LLM | |||
|---|---|---|---|---|---|---|
| Qwen2.5-VL-3B-Instruct | 1 | 128 | Unfrozen | |||
| MiMo-VL-7B-SFT | 1 | 128 | Unfrozen | |||
| Llama-3.2-11B-Vision-Instruct | 1 | 128 | N/A | Frozen |
| Model | Trainable modules | Reward model | Epoch | LR | Batch | Rollouts | Details |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-3B-Instruct | Full model | Qwen2.5-3B | 1 | 128 | 4 | GRPO, step 70 | |
| Llama-3.2-11B-Vision-Instruct | Vision + adapters | Qwen2.5-3B | 1 | 32 | 4 | GRPO, step 281 | |
| MiMo-VL-7B-SFT | Full model | Qwen2.5-3B | 1 | 64 | 4 | GRPO, step 140 |
| Split | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| Overall audited set | 92.1 | 99.0 | 90.8 | 94.7 |
| Metric | Definition | Value |
|---|---|---|
| Final valid proportion | generated correct & filter pass over all audited candidates | 70.7% |
| Acceptance rate | retained / generated candidates | 71.4% |
| Rejection rate | rejected / generated candidates | 28.6% |
| False accept rate | generated wrong but filter pass / generated wrong candidates | 3.2% |
| False reject rate | generated correct but filter fail / generated correct candidates | 9.2% |
| Human–judge agreement | Cohen’s on audited subset | 0.794 |
| Task | MVP-S | MVP-P | CV2D | CV3D | RWQA | BLK | MMS | OCR |
|---|---|---|---|---|---|---|---|---|
| OaD | 64.0 | 36.0 | 68.4 | 73.9 | 66.4 | 49.1 | 55.8 | 82.4 |
| VaP | 65.0 | 36.0 | 68.4 | 73.4 | 65.6 | 48.9 | 55.8 | 81.8 |
| PaRC | 66.0 | 38.0 | 67.9 | 73.5 | 66.1 | 49.5 | 55.4 | 82.2 |
| SR | 64.3 | 36.0 | 67.9 | 72.7 | 66.1 | 49.0 | 55.3 | 82.3 |
| SaC | 64.7 | 36.7 | 68.3 | 71.5 | 65.6 | 48.9 | 55.4 | 82.0 |
| SPC | 65.0 | 36.0 | 67.9 | 73.7 | 66.0 | 48.5 | 55.7 | 82.4 |
| Task | LEGO | MMMU | MMB | MVis | SSP | MMSI | 3DSR |
|---|---|---|---|---|---|---|---|
| OaD | 15.7 | 50.0 | 76.2 | 64.2 | 20.7 | 8.8 | 35.2 |
| VaP | 16.3 | 49.0 | 75.9 | 65.2 | 20.5 | 8.7 | 35.2 |
| PaRC | 16.4 | 48.8 | 76.8 | 64.6 | 20.6 | 9.4 | 35.3 |
| SR | 16.0 | 48.1 | 76.2 | 63.3 | 20.2 | 9.4 | 34.9 |
| SaC | 16.1 | 48.8 | 76.1 | 64.4 | 20.7 | 9.7 | 35.2 |
| SPC | 15.8 | 48.4 | 76.3 | 63.9 | 20.4 | 9.5 | 35.3 |
| Setting | CV-Bench-2D | CV-Bench-3D | RealWorldQA | ||
|---|---|---|---|---|---|
| Qwen2.5-VL-3B baseline | 35.3 | 64.3 | 67.3 | 66.0 | 65.0 |
| Self-verifier 10K SFT | 40.7 | 68.7 | 69.3 | 75.2 | 65.9 |
| Improvement | +5.4 | +4.4 | +2.0 | +9.2 | +0.9 |