Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
Figures & tables
Figure 1 : Overview of MMVistaReason benchmark performance. Left: MMVistaReason achieves strong scores against representative baselines; right: MMVistaReason-SFT-528K and MMVistaReason-RL-63K cover diverse reasoning domains and complementary training groups.
Figure 2 : Overview of the MMVistaReason Recipes. It introduces a unified pipeline covering data cleaning, reasoning trace construction, reinforcement learning, and multi-teacher integration, enabling scalable training of reliable multimodal reasoning models from open data.
Figure 3 : Data flow and statistics of MMVistaReason. (a) Data filtering and balancing from the raw pool to SFT and RL data. (b) Statistics of MMVistaReason datasets.
Dataset
Distill-model
#Samples
Tokens
Mean
Std.
Median
P25
P75
P95
Min
Max
MVR-SFT-355K
Qwen3.5-9B
354,886
631.72M
1,780.06
2,203.43
900
436
2,195
6,316
26
16,374
MVR-SFT-124K
Qwen3.5-27B
124,010
309.30M
2,494.13
3,290.11
1,031
440
2,998
10,414
45
16,375
MVR-SFT-49K
Qwen3.5-122B
48,913
103.04M
2,106.55
3,120.21
670
340
2,313
9,657
39
16,381
MVR-SFT-528K
All
527,809
1.04B
1,978.09
2,607.77
906
426
2,359
7,625
26
16,381
Table 1 : Statistics of MMVistaReason (MVR)-SFT-528K and its sub-datasets. Each sub-dataset is distilled from the teacher model that first solves the sample in the cascaded rejection sampling pipeline. Token statistics are computed over model responses using the Qwen3.5 tokenizer.
Benchmarks
Closed-source VLMs
Open-weight VLMs
Open-source VLMs
Ours
Gemini-3 Flash
GPT-5.1
Qwen3.5-9B (Instruct)
Qwen3.5-35B A3B (Instruct)
InternVL3.5 30B-A3B
InternVL3.5 241B-A28B
OMR 7B
MFR 4B
MFR 8B
MVR 4B
MVR 9B
OSWorld-G
65.4
62.6
60.4
62.8
45.7
54.9
32.8
41.4
48.3
58.5
64.8
CV-Bench-2D
84.6
83.0
82.9
79.4
80.3
81.2
76.9
79.2
80.2
82.4
83.2
CV-Bench-3D
92.8
91.6
91.3
93.1
87.8
89.3
84.0
89.1
90.8
91.8
92.4
MMBench-EN
91.5
89.6
90.5
92.1
85.2
87.8
88.3
88.8
89.5
90.6
92.2
RealWorldQA
80.9
79.1
76.8
76.7
72.3
75.2
68.8
74.6
75.2
77.4
77.4
Table 2 : Comparison of MMVistaReason (MVR) models with representative closed-source, open-weight, and open-source VLMs across diverse multimodal benchmarks.
Benchmark
4B Models
9B Models
Base
Inst.
Think.
SFT
MOPD
Base
Inst.
Think.
SFT
MOPD
OSWorld-G
41.5
52.0
54.2
54.0
58.5
43.2
60.4
62.6
59.9
64.8
CV-Bench-2D
80.9
82.7
82.1
80.8
82.4
81.8
82.9
83.1
80.3
83.2
CV-Bench-3D
91.2
90.5
91.9
90.6
91.8
91.6
91.3
92.6
91.8
92.4
MMBench-EN
89.8
90.2
90.0
90.1
90.6
91.0
90.5
90.8
91.7
92.2
RealWorldQA
75.2
75.0
77.9
73.9
77.4
75.3
76.8
79.0
75.2
77.4
Table 3 : Performance comparison across different model scales and training stages on multimodal benchmarks. The best result within each model scale is highlighted in bold. Inst. and Think. denote the instruct and thinking modes of the Qwen3.5 models, respectively.
Figure 4 : Data scale and performance trade-off across different student model capacities.
Figure 5 : Benchmark-wise score changes of different reasoning trajectory selection strategies over random correct trajectory selection for 4B (top) and 9B (bottom) student models.
Scale
Stage
Real-World Visual Reasoning
Analytical Visual Reasoning
OSW
CV2D
CV3D
MMB
RWA
Count
Avg.
SFE
MMMU
SQA
MVis
MVer
LVis
VLog
CQA
CharXiv
Avg.
4B
SFT
54.0
80.8
90.6
90.1
73.9
88.7
79.7
17.8
70.7
96.2
80.4
78.3
65.1
28.1
83.9
63.6
64.9
RW Expert
59.5
82.9
92.3
91.2
78.6
91.4
82.7
17.0
67.6
95.9
80.5
77.1
64.9
27.8
83.1
64.7
64.3
Ana. Expert
53.7
78.9
90.6
90.1
74.3
87.7
79.2
19.6
71.4
97.2
82.5
79.8
66.7
29.1
86.7
66.7
66.6
MOPD
58.5
82.4
91.8
90.6
77.4
91.5
82.0
19.8
71.8
96.8
82.4
79.7
66.4
29.0
86.2
66.9
66.6
Δ vs. SFT
↑ 4.5
↑ 1.6
↑ 1.2
↑ 0.5
↑ 3.5
↑ 2.8
↑ 2.3
↑ 2.0
↑ 1.1
↑ 0.6
↑ 2.0
↑ 1.4
↑ 1.3
↑ 0.9
↑ 2.3
↑ 3.3
↑ 1.7
Table 4 : Benchmark-wise and domain-level performance of complementary RL experts across different model scales. RW and Ana. denote Real-World and Analytical experts, respectively. Benchmark names are abbreviated due to space constraints. Green and red backgrounds indicate gains and drops over the corresponding SFT model. The Δ rows report the improvement of MOPD over SFT.
Figure 6 : Comparison of Mixed-RL and MOPD under different KL directions across model scales. Colors indicate changes over MVR-SFT for 4B (top) and 9B (bottom) models.
Subset Name
Samples
Tokens
Subset Name
Samples
Tokens
SciMM [ 40 ]
247,292
291,185,096
vero_captioning_IF [ 46 ]
4,506
2,755,634
MMR1 [ 22 ]
77,116
322,402,209
PRISM [ 55 ]
3,831
11,606,008
vero_spatial_action [ 46 ]
30,153
48,377,896
Euclid30K [ 25 ]
3,038
14,260,100
FineVision [ 60 ]
22,799
42,814,429
WaltonColdStart [ 58 ]
2,932
7,034,074
vero_chart_ocr [ 46 ]
21,998
19,464,798
ViRL39K [ 54 ]
2,236
6,383,976
deepvision_math [ 49 ]
19,550
88,001,303
LLaVA-CoT [ 65 ]
2,018
906,254
Table 5 : Detailed source composition of MMVistaReason-SFT-528K . We report the number of samples and response tokens for each source subset.
SFT Subset
Mastery=0.25
Mastery=0.50
Mastery=0.75
Mastery=1.00
Weighted Mean
MVR-SFT-49K (122B)
69.26%
19.53%
7.61%
3.61%
0.3639
MVR-SFT-124K (27B)
49.74%
25.50%
15.17%
9.59%
0.4616
MVR-SFT-355K (9B)
40.49%
26.23%
19.06%
14.23%
0.5176
Table 6 : Difficulty distribution of different SFT subsets generated by cascaded teacher distillation.
Subtype
122B Teacher
27B Teacher
9B Teacher
Overall
Numeric
12,058 (24.65%)
37,806 (30.49%)
110,705 (31.19%)
160,569 (30.42%)
Short Phrase
16,985 (34.72%)
41,626 (33.57%)
89,283 (25.16%)
147,894 (28.02%)
Multiple Choice
4,665 (9.54%)
19,954 (16.09%)
72,450 (20.42%)
97,069 (18.39%)
Entity Name
7,188 (14.70%)
16,611 (13.39%)
51,856 (14.61%)
75,655 (14.33%)
Yes/No
748 (1.53%)
2,562 (2.07%)
13,504 (3.81%)
16,814 (3.19%)
Position
6,358 (13.00%)
2,720 (2.19%)
5,045 (1.42%)
14,123 (2.68%)
Table 7 : Question subtype distribution of MVR-SFT-528K constructed by different teacher models. Each entry reports the number of samples and the corresponding ratio within each subset.
Domain
#Samples
Tokens
Mean
Std.
Median
P75
P95
Max
Science
175,007
208.90M
1,193.7
1,365.4
681
1,420
3,723
16,369
Mathematics
130,677
518.85M
3,970.5
3,739.1
2,669
5,875
12,188
16,381
Logic/Game/Puzzle
47,140
131.24M
2,784.1
2,448.5
2,107
3,919
7,421
16,375
Chart/Table/Doc
109,250
121.95M
1,116.3
1,314.4
588
1,393
3,708
16,039
General
32,165
15.75M
489.5
635.3
277
513
1,532
13,725
Spatial
23,470
42.31M
1,802.6
2,182.1
978
2,238
6,344
16,374
Table 8 : Domain-wise response-token statistics of MVR-SFT-528K. Token statistics are computed over model responses using the Qwen3.5 tokenizer for all seven reasoning domains.
RL Expert
0.125
0.250
0.375
0.500
0.625
0.750
0.875
Mean
Median
4B Analytical RL
44
1,955
6,654
7,633
2,620
194
25
0.4502
0.500
4B Real-World RL
0
1,348
2,939
3,270
2,624
1,416
0
0.4981
0.500
9B Analytical RL
47
2,085
7,097
8,142
2,795
207
27
0.4503
0.500
9B Real-World RL
0
1,393
3,036
3,377
2,711
1,463
0
0.4981
0.500
Overall
91
6,781
19,726
22,422
10,750
3,280
52
0.4681
0.500
Table 9 : Difficulty distribution of 4B and 9B RL experts measured by SFT rollout pass rate.
Question Type
4B Analytical
4B Real-World
9B Analytical
9B Real-World
Multiple Choice
11,269 (58.92%)
4,975 (42.90%)
11,669 (57.20%)
5,002 (41.75%)
Numeric
5,250 (27.45%)
254 (2.19%)
5,007 (24.54%)
287 (2.40%)
Counting
396 (2.07%)
3,404 (29.35%)
446 (2.19%)
3,482 (29.07%)
Position
22 (0.12%)
1,943 (16.75%)
37 (0.18%)
2,052 (17.13%)
Short Phrase
1,107 (5.79%)
560 (4.83%)
1,813 (8.89%)
659 (5.50%)
Entity Name
778 (4.07%)
175 (1.51%)
978 (4.79%)
205 (1.71%)
Table 10 : Question type distribution of 4B and 9B RL experts.
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China · The University of Sydney, Australia · Nanyang Technological University, Singapore