MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning
Organizations: Zhejiang University · Shanghai Artificial Intelligence Laboratory, OpenDataLab · Shanghai Jiao Tong University · Tongji University
Abstract
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
Figures & tables
| Dataset | Distill-model | #Samples | Tokens | Mean | Std. | Median | P25 | P75 | P95 | Min | Max |
| MVR-SFT-355K | Qwen3.5-9B | 354,886 | 631.72M | 1,780.06 | 2,203.43 | 900 | 436 | 2,195 | 6,316 | 26 | 16,374 |
| MVR-SFT-124K | Qwen3.5-27B | 124,010 | 309.30M | 2,494.13 | 3,290.11 | 1,031 | 440 | 2,998 | 10,414 | 45 | 16,375 |
| MVR-SFT-49K | Qwen3.5-122B | 48,913 | 103.04M | 2,106.55 | 3,120.21 | 670 | 340 | 2,313 | 9,657 | 39 | 16,381 |
| MVR-SFT-528K | All | 527,809 | 1.04B | 1,978.09 | 2,607.77 | 906 | 426 | 2,359 | 7,625 | 26 | 16,381 |
| Benchmarks | Closed-source VLMs | Open-weight VLMs | Open-source VLMs | Ours | |||||||
| Gemini-3 Flash | GPT-5.1 | Qwen3.5-9B (Instruct) | Qwen3.5-35B A3B (Instruct) | InternVL3.5 30B-A3B | InternVL3.5 241B-A28B | OMR 7B | MFR 4B | MFR 8B | MVR 4B | MVR 9B | |
| OSWorld-G | 65.4 | 62.6 | 60.4 | 62.8 | 45.7 | 54.9 | 32.8 | 41.4 | 48.3 | 58.5 | 64.8 |
| CV-Bench-2D | 84.6 | 83.0 | 82.9 | 79.4 | 80.3 | 81.2 | 76.9 | 79.2 | 80.2 | 82.4 | 83.2 |
| CV-Bench-3D | 92.8 | 91.6 | 91.3 | 93.1 | 87.8 | 89.3 | 84.0 | 89.1 | 90.8 | 91.8 | 92.4 |
| MMBench-EN | 91.5 | 89.6 | 90.5 | 92.1 | 85.2 | 87.8 | 88.3 | 88.8 | 89.5 | 90.6 | 92.2 |
| RealWorldQA | 80.9 | 79.1 | 76.8 | 76.7 | 72.3 | 75.2 | 68.8 | 74.6 | 75.2 | 77.4 | 77.4 |
| Benchmark | 4B Models | 9B Models | ||||||||
| Base | Inst. | Think. | SFT | MOPD | Base | Inst. | Think. | SFT | MOPD | |
| OSWorld-G | 41.5 | 52.0 | 54.2 | 54.0 | 58.5 | 43.2 | 60.4 | 62.6 | 59.9 | 64.8 |
| CV-Bench-2D | 80.9 | 82.7 | 82.1 | 80.8 | 82.4 | 81.8 | 82.9 | 83.1 | 80.3 | 83.2 |
| CV-Bench-3D | 91.2 | 90.5 | 91.9 | 90.6 | 91.8 | 91.6 | 91.3 | 92.6 | 91.8 | 92.4 |
| MMBench-EN | 89.8 | 90.2 | 90.0 | 90.1 | 90.6 | 91.0 | 90.5 | 90.8 | 91.7 | 92.2 |
| RealWorldQA | 75.2 | 75.0 | 77.9 | 73.9 | 77.4 | 75.3 | 76.8 | 79.0 | 75.2 | 77.4 |
| Scale | Stage | Real-World Visual Reasoning | Analytical Visual Reasoning | |||||||||||||||
| OSW | CV2D | CV3D | MMB | RWA | Count | Avg. | SFE | MMMU | SQA | MVis | MVer | LVis | VLog | CQA | CharXiv | Avg. | ||
| 4B | SFT | 54.0 | 80.8 | 90.6 | 90.1 | 73.9 | 88.7 | 79.7 | 17.8 | 70.7 | 96.2 | 80.4 | 78.3 | 65.1 | 28.1 | 83.9 | 63.6 | 64.9 |
| RW Expert | 59.5 | 82.9 | 92.3 | 91.2 | 78.6 | 91.4 | 82.7 | 17.0 | 67.6 | 95.9 | 80.5 | 77.1 | 64.9 | 27.8 | 83.1 | 64.7 | 64.3 | |
| Ana. Expert | 53.7 | 78.9 | 90.6 | 90.1 | 74.3 | 87.7 | 79.2 | 19.6 | 71.4 | 97.2 | 82.5 | 79.8 | 66.7 | 29.1 | 86.7 | 66.7 | 66.6 | |
| MOPD | 58.5 | 82.4 | 91.8 | 90.6 | 77.4 | 91.5 | 82.0 | 19.8 | 71.8 | 96.8 | 82.4 | 79.7 | 66.4 | 29.0 | 86.2 | 66.9 | 66.6 | |
| vs. SFT | 4.5 | 1.6 | 1.2 | 0.5 | 3.5 | 2.8 | 2.3 | 2.0 | 1.1 | 0.6 | 2.0 | 1.4 | 1.3 | 0.9 | 2.3 | 3.3 | 1.7 | |
| Subset Name | Samples | Tokens | Subset Name | Samples | Tokens |
| SciMM [ 40 ] | 247,292 | 291,185,096 | vero_captioning_IF [ 46 ] | 4,506 | 2,755,634 |
| MMR1 [ 22 ] | 77,116 | 322,402,209 | PRISM [ 55 ] | 3,831 | 11,606,008 |
| vero_spatial_action [ 46 ] | 30,153 | 48,377,896 | Euclid30K [ 25 ] | 3,038 | 14,260,100 |
| FineVision [ 60 ] | 22,799 | 42,814,429 | WaltonColdStart [ 58 ] | 2,932 | 7,034,074 |
| vero_chart_ocr [ 46 ] | 21,998 | 19,464,798 | ViRL39K [ 54 ] | 2,236 | 6,383,976 |
| deepvision_math [ 49 ] | 19,550 | 88,001,303 | LLaVA-CoT [ 65 ] | 2,018 | 906,254 |
| SFT Subset | Mastery=0.25 | Mastery=0.50 | Mastery=0.75 | Mastery=1.00 | Weighted Mean |
| MVR-SFT-49K (122B) | 69.26% | 19.53% | 7.61% | 3.61% | 0.3639 |
| MVR-SFT-124K (27B) | 49.74% | 25.50% | 15.17% | 9.59% | 0.4616 |
| MVR-SFT-355K (9B) | 40.49% | 26.23% | 19.06% | 14.23% | 0.5176 |
| Subtype | 122B Teacher | 27B Teacher | 9B Teacher | Overall |
| Numeric | 12,058 (24.65%) | 37,806 (30.49%) | 110,705 (31.19%) | 160,569 (30.42%) |
| Short Phrase | 16,985 (34.72%) | 41,626 (33.57%) | 89,283 (25.16%) | 147,894 (28.02%) |
| Multiple Choice | 4,665 (9.54%) | 19,954 (16.09%) | 72,450 (20.42%) | 97,069 (18.39%) |
| Entity Name | 7,188 (14.70%) | 16,611 (13.39%) | 51,856 (14.61%) | 75,655 (14.33%) |
| Yes/No | 748 (1.53%) | 2,562 (2.07%) | 13,504 (3.81%) | 16,814 (3.19%) |
| Position | 6,358 (13.00%) | 2,720 (2.19%) | 5,045 (1.42%) | 14,123 (2.68%) |
| Domain | #Samples | Tokens | Mean | Std. | Median | P75 | P95 | Max |
| Science | 175,007 | 208.90M | 1,193.7 | 1,365.4 | 681 | 1,420 | 3,723 | 16,369 |
| Mathematics | 130,677 | 518.85M | 3,970.5 | 3,739.1 | 2,669 | 5,875 | 12,188 | 16,381 |
| Logic/Game/Puzzle | 47,140 | 131.24M | 2,784.1 | 2,448.5 | 2,107 | 3,919 | 7,421 | 16,375 |
| Chart/Table/Doc | 109,250 | 121.95M | 1,116.3 | 1,314.4 | 588 | 1,393 | 3,708 | 16,039 |
| General | 32,165 | 15.75M | 489.5 | 635.3 | 277 | 513 | 1,532 | 13,725 |
| Spatial | 23,470 | 42.31M | 1,802.6 | 2,182.1 | 978 | 2,238 | 6,344 | 16,374 |
| RL Expert | 0.125 | 0.250 | 0.375 | 0.500 | 0.625 | 0.750 | 0.875 | Mean | Median |
| 4B Analytical RL | 44 | 1,955 | 6,654 | 7,633 | 2,620 | 194 | 25 | 0.4502 | 0.500 |
| 4B Real-World RL | 0 | 1,348 | 2,939 | 3,270 | 2,624 | 1,416 | 0 | 0.4981 | 0.500 |
| 9B Analytical RL | 47 | 2,085 | 7,097 | 8,142 | 2,795 | 207 | 27 | 0.4503 | 0.500 |
| 9B Real-World RL | 0 | 1,393 | 3,036 | 3,377 | 2,711 | 1,463 | 0 | 0.4981 | 0.500 |
| Overall | 91 | 6,781 | 19,726 | 22,422 | 10,750 | 3,280 | 52 | 0.4681 | 0.500 |
| Question Type | 4B Analytical | 4B Real-World | 9B Analytical | 9B Real-World |
| Multiple Choice | 11,269 (58.92%) | 4,975 (42.90%) | 11,669 (57.20%) | 5,002 (41.75%) |
| Numeric | 5,250 (27.45%) | 254 (2.19%) | 5,007 (24.54%) | 287 (2.40%) |
| Counting | 396 (2.07%) | 3,404 (29.35%) | 446 (2.19%) | 3,482 (29.07%) |
| Position | 22 (0.12%) | 1,943 (16.75%) | 37 (0.18%) | 2,052 (17.13%) |
| Short Phrase | 1,107 (5.79%) | 560 (4.83%) | 1,813 (8.89%) | 659 (5.50%) |
| Entity Name | 778 (4.07%) | 175 (1.51%) | 978 (4.79%) | 205 (1.71%) |