Same Reward, Different Skills: When Multimodal RL Learns to Look
Organizations: University of the Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Independent Researcher
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
Figures & tables
| Evaluated with the real image | Test image removed | |||
| Trained with | gain over base | share of Real ’s gain | gray canvas | no image |
| base model (accuracy) | 0.175 | — | 0.090 | 0.068 |
| Real | 0.238 | 1 (reference) | 0.019 | 0.036 |
| Caption | 0.171 | 0.72 | 0.024 | 0.037 |
| trained without visual information | ||||
| None | 0.137 | 0.57 | 0.015 | 0.044 |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| benchmark (format) | with image | blind | chance | naive retention | corrected retention | |
|---|---|---|---|---|---|---|
| BLINK (MC, pooled) | 1901 | 0.493 | 0.409 | 0.377 | 0.829 [0.781, 0.879] | 0.271 [0.089, 0.454] |
| MathVerse (MC, pooled) | 2180 | 0.465 | 0.394 | 0.260 | 0.848 [0.807, 0.889] | 0.655 [0.568, 0.743] |
| MathVerse (free-form) | 1755 | 0.055 | 0.019 | 0.000 | 0.340 [0.240, 0.458] | 0.340 [0.240, 0.458] |
| MMVP (MC, ) | 300 | 0.660 | 0.500 | 0.500 | 0.758 [0.668, 0.854] | 0.000 [ 0.417, 0.342] |
| MMMU dev+val (MC, pooled) | 988 | 0.506 | 0.413 | 0.263 | 0.816 [0.764, 0.873] | 0.617 [0.512, 0.729] |
| MMMU dev+val (free-form) | 62 | 0.097 | 0.048 | 0.000 | 0.500 [0.000, 1.000] | 0.500 [0.000, 1.000] |
| model | pool | with image | blind | chance | naive retention | corrected retention | |
|---|---|---|---|---|---|---|---|
| Gemma-3 | ViRL39K free-form | 2789 | 0.430 | 0.312 | 0.000 | 0.727 [0.690, 0.765] | 0.727 [0.690, 0.765] |
| Gemma-3 | ViRL39K multiple choice | 1215 | 0.135 | 0.086 | 0.268 | 0.634 [0.533, 0.744] | undefined |
| InternVL3-9B | ViRL39K free-form | 2789 | 0.269 | 0.130 | 0.000 | 0.485 [0.439, 0.533] | 0.485 [0.439, 0.533] |
| InternVL3-9B | ViRL39K multiple choice | 1215 | 0.294 | 0.205 | 0.268 | 0.697 [0.619, 0.782] | 2.439 [ 17.961, 6.505] |
| level | question | visual cue | gold (twin a) | gold (twin b) |
|---|---|---|---|---|
| L1 (arrow cue; position given) | Point K8 has the smallest -coordinate. What is the -coordinate of point K8? | offset arrow at K8 | ||
| L2 (target named; identity given) | Point K8 has the smallest -coordinate. What is the -coordinate of point K8? | none | ||
| L3 (discover ground read; nothing given) | Consider the point with the smallest -coordinate. What is its -coordinate? | none | ||
| probe (identification only) | Which labeled point has the smallest -coordinate? | none | K8 | K8 |
| model | discovery 8-pt | 12-pt | 20-pt | probe 8-pt | 12-pt | 20-pt |
|---|---|---|---|---|---|---|
| 3B, ViRL39K-trained (step 100); two seeds | ||||||
| untrained 3B | 0.330 | 0.260 | 0.245 | 0.705 | 0.680 | 0.630 |
| Real , two-seed mean | 0.383 | 0.338 | 0.330 | 0.728 | 0.700 | 0.610 |
| seeds 1 / 2 | 0.380 / 0.385 | 0.330 / 0.345 | 0.330 / 0.330 | 0.725 / 0.730 | 0.710 / 0.690 | 0.620 / 0.600 |
| Caption , two-seed mean | 0.370 | 0.323 | 0.340 | 0.728 | 0.715 | 0.635 |
| seeds 1 / 2 | 0.365 / 0.375 | 0.325 / 0.320 | 0.345 / 0.335 | 0.715 / 0.740 | 0.710 / 0.720 | 0.610 / 0.660 |
| condition | cued readout: gain [95% CI] | find-and-bind | header-cued table |
|---|---|---|---|
| untrained 3B (pair accuracy) | 0.320 | 0.455 | 0.867 |
| Real | 0.143 [0.102, 0.188] | 0.018 [ 0.007, 0.044] | 0.019 [ 0.002, 0.042] |
| seeds 1 / 2 / 3 | 0.157 / 0.127 / 0.147 | 0.012 / 0.022 / 0.022 | 0.030 / 0.013 / 0.013 |
| Caption | 0.107 [0.066, 0.149] | 0.006 [ 0.018, 0.030] | 0.021 [ 0.001, 0.044] |
| seeds 1 / 2 / 3 | 0.130 / 0.113 / 0.077 | 0.003 / 0.000 / 0.013 | 0.023 / 0.027 / 0.013 |
| None | 0.108 [0.070, 0.148] | 0.013 [ 0.039, 0.013] | 0.023 [0.000, 0.049] |
| (4) caption-only accuracy | |||||||
|---|---|---|---|---|---|---|---|
| set | density | discovery (L3) | accuracy | limit ( L3) | outcome | (5) artifact screen | criteria met |
| development | 8-pt | 0.660 | 0.580 | 0.330 | not met | met | 4 of 5 |
| 12-pt | 0.575 | 0.445 | 0.288 | not met | not met: 0.557 in 1 of 3 seeds | 3 of 5 | |
| 20-pt | 0.470 | 0.185 | 0.235 | met | not met: 0.565 in 1 of 3 seeds | 4 of 5 | |
| confirmatory | 8-pt | 0.585 | 0.570 | 0.293 | not met | met | 4 of 5 |
| 12-pt | 0.560 | 0.410 | 0.280 | not met | met | 4 of 5 | |
| task | pairs | role | untrained 3B | untrained 7B [95% CI] |
|---|---|---|---|---|
| grounding task (named-point coordinate pairs) | 600 | find-and-bind (primary) | 0.455 | 0.768 [0.733, 0.802] |
| header-cued table task | 300 | cued control, high baseline | 0.867 | 0.993 [0.983, 1.000] |
| cued-readout task (marked-plot pairs) | 300 | cued readout, location marked | 0.320 | 0.673 [0.620, 0.727] |
| twin of the grounding task | 600 | regenerated twin | – | 0.728 [0.692, 0.762] |
| twin of the cued-readout task | 300 | twin control | – | 0.623 [0.570, 0.677] |
| twin of the header-cued table task | 300 | twin control | – | 0.997 [0.990, 1.000] |
| experiment | backbone, vision encoder | training images | reward | training data | steps; evaluated at | seeds | runs |
|---|---|---|---|---|---|---|---|
| Geometry3K access conditions (§ 2.1 , § 3 ) | 3B, frozen | Real / Gray / None / Caption , one condition per run | answer correctness | Geometry3K, filtered (1,288 items) | 100 | 3 | 3 |
| 7B access pair (§ 2.1 , § 3 ) | 7B, frozen | Real ; Gray | answer correctness | Geometry3K | 100 | 1 | 1 |
| long-horizon runs (§ 3 ) | 3B, trainable | real images | accuracy format | Geometry3K, unfiltered | 400, in segments; evaluated at 100, 150, 200, 300, 400 | 2 | 2 |
| resolvable training, standard reward (§ 6 ) | 7B, frozen | real images | accuracy format | constructed corpus, 2,880 items (discovery and probe, both twins) | seeds 0 and 1: 100, in segments (seed 0 evaluated at 10, 20, 30, 50, 75, 100; seed 1 at 30 and 100). Seeds 2 and 3: 30 | 4 | 5 (two of seed 0) |
| Gray -trained control (§ 6.2 ) | 7B, frozen | gray canvases | as above | same corpus | 30; evaluated at 10, 20, 30 (seed 0) or 30 (seeds 1–3) | 4 | 4 |
| named-target ablation (§ 6.1 ; Table G.6 ) | 7B, frozen | real images | as above | same scene programs, with the find-and-bind (L2) question in place of the discovery question; probe unchanged | 30 | 3 | 3 |
| condition | seed | acc. real | acc. gray | acc. none | acc. caption | gain real | gain gray | gain none | gain caption | share |
|---|---|---|---|---|---|---|---|---|---|---|
| base model | – | 0.175 | 0.090 | 0.068 | 0.210 | – | – | – | – | – |
| Real | 1 | 0.423 | 0.106 | 0.108 | 0.311 | 0.248 | 0.017 | 0.040 | 0.102 | – |
| 2 | 0.419 | 0.112 | 0.098 | 0.301 | 0.245 | 0.022 | 0.030 | 0.091 | – | |
| 3 | 0.398 | 0.110 | 0.106 | 0.326 | 0.223 | 0.020 | 0.038 | 0.117 | – | |
| Caption | 1 | 0.361 | 0.118 | 0.112 | 0.318 | 0.186 | 0.028 | 0.043 | 0.108 | 0.752 |
| 2 | 0.346 | 0.108 | 0.106 | 0.291 | 0.171 | 0.018 | 0.038 | 0.082 | 0.701 |
| task | pairs | untrained 7B | Real [95% CI] | Gray [95% CI] |
|---|---|---|---|---|
| grounding task | 600 | 0.768 | 0.035 [0.013, 0.058] | 0.012 [ 0.010, 0.033] |
| twin of the grounding task | 600 | 0.728 | 0.038 [0.017, 0.062] | 0.013 [ 0.010, 0.037] |
| cued-readout task | 300 | 0.673 | 0.030 [ 0.007, 0.067] | 0.017 [ 0.013, 0.047] |
| twin of the cued-readout task | 300 | 0.623 | 0.013 [ 0.030, 0.060] | 0.013 [ 0.023, 0.050] |
| header-cued table task | 300 | 0.993 | 0.000 [ 0.010, 0.010] | 0.003 [ 0.010, 0.000] |
| twin of the header-cued table task | 300 | 0.997 | 0.000 [ 0.010, 0.010] | 0.007 [ 0.017, 0.000] |
| trained with | step | real image: correct [95% CI] | gain over untrained [95% CI] | gray canvas: correct [95% CI] |
|---|---|---|---|---|
| untrained 7B | – | 141/601 = 0.235 [0.201, 0.270] | – | 48/601 = 0.080 [0.058, 0.103] |
| Real | 100 | 290/601 = 0.483 [0.443, 0.521] | 0.248 [0.203, 0.291] (149/601) | 75/601 = 0.125 [0.100, 0.151] |
| Gray | 100 | 257/601 = 0.428 [0.388, 0.466] | 0.193 [0.153, 0.235] (116/601) | 79/601 = 0.131 [0.103, 0.160] |
| questions | untrained | Real gain | Caption gain | None gain | Gray gain | |
|---|---|---|---|---|---|---|
| all questions | 601 | 0.175 | 0.238 [0.202, 0.275] | 0.171 [0.136, 0.207] | 0.137 [0.103, 0.171] | 0.125 [0.092, 0.158] |
| extractable | 497 | 0.211 | 0.200 [0.160, 0.240] | 0.140 [0.100, 0.179] | 0.107 [0.069, 0.145] | 0.101 [0.064, 0.139] |
| not extractable | 104 | 0.000 | 0.423 [0.343, 0.503] | 0.324 [0.250, 0.401] | 0.279 [0.212, 0.349] | 0.237 [0.176, 0.301] |
| result (section; table) | seeds | runs | individual seeds or runs | 95% intervals resample |
|---|---|---|---|---|
| blind recovery shares, 3B (§ 2.1 ; Table 1 ) | 3 | 3 | Table D.2 , per seed | point estimates over seeds |
| Gray recovery share, 7B (§ 2.1 ; Table D.4 ) | 1 | 1 | – | benchmark items |
| operation-level gains, 3B (§ 2.3 ; Table B.3 ) | 3 | 3 | Table B.3 , per seed | grounding pairs |
| discovery, ViRL39K-trained 3B (§ 2.3 ; Table B.2 ) | 2 | 2 | Table B.2 , per seed | point estimates over seeds |
| grounding loss under prolonged training (§ 3 ; Table E.1 ) | 2 | 2 | Table E.1 , per run | grounding pairs |
| discovery on the confirmatory set (§ 6.1 ; Tables G.2 , G.6 ) | 4 | 5 | Table G.2 (seed 0, per run); Table G.6 (seeds 1–3; seed 1 also at step 100) | scene programs (seed 0); seeds, -based (Table G.7 ) |
| condition | seed | step 1: overall / format / accuracy | step 30: overall / format / accuracy | mean accuracy, steps 19–30 |
|---|---|---|---|---|
| Real | 1 | 0.5975 / 0.8658 / 0.3292 | 0.9858 / 1.0000 / 0.9717 | 0.975 |
| Real | 2 | 0.5900 / 0.8875 / 0.2925 | 0.9863 / 1.0000 / 0.9725 | 0.982 |
| Real | 3 | 0.6095 / 0.8825 / 0.3365 | 0.9930 / 0.9992 / 0.9867 | 0.970 |
| Gray -trained | 1 | 0.4350 / 0.8683 / 0.0017 | 0.5175 / 1.0000 / 0.0350 | 0.033 |
| Gray -trained | 2 | 0.4413 / 0.8800 / 0.0025 | 0.5238 / 0.9992 / 0.0483 | 0.041 |
| Gray -trained | 3 | 0.4425 / 0.8800 / 0.0050 | 0.5104 / 1.0000 / 0.0208 | 0.024 |
| seed | step | benchmark [95% CI] | grounding [95% CI] | vs base [95% CI] |
|---|---|---|---|---|
| 1 | 0 | 0.175 [0.145, 0.206] | 0.455 [0.415, 0.495] | – |
| 100 | 0.431 [0.393, 0.471] | 0.477 [0.437, 0.517] | 0.022 [ 0.008, 0.052] | |
| 150 | 0.463 [0.424, 0.504] | 0.465 [0.425, 0.505] | 0.010 [ 0.020, 0.040] | |
| 200 | 0.483 [0.444, 0.524] (peak) | 0.450 [0.412, 0.490] | 0.005 [ 0.035, 0.025] | |
| 300 | 0.464 [0.424, 0.506] | 0.440 [0.400, 0.480] | 0.015 [ 0.047, 0.017] | |
| 400 | 0.436 [0.398, 0.474] | 0.407 [0.367, 0.447] | 0.048 [ 0.082, 0.017] |
| run | degraded pairs | pairs gained | wrong members |
|---|---|---|---|
| seed 1 | 51 | 24 | 52 |
| seed 2 | 49 | 22 | 53 |
| seed 3 | 45 | 23 | 46 |
| run | step | 8-pt: level [95% CI]; [95% CI] | 12-pt | 20-pt | probe 8 / 12 / 20-pt |
|---|---|---|---|---|---|
| untrained 7B | 0 | 0.660 [0.580, 0.740] | 0.575 [0.485, 0.660] | 0.470 [0.385, 0.555] | 0.940 / 0.910 / 0.840 |
| run 1 | 30 | 0.965 [0.935, 0.990] 0.305 [0.230, 0.380] | 0.890 [0.830, 0.945] 0.315 [0.230, 0.400] | 0.800 [0.730, 0.865] 0.330 [0.255, 0.410] | 1.000 / 0.975 / 0.975 |
| 100 | 200/200 (100 programs) 0.340 [0.260, 0.420] | 0.940 [0.895, 0.980] 0.365 [0.280, 0.455] | 0.910 [0.860, 0.950] 0.440 [0.360, 0.525] | 1.000 / 0.990 / 0.995 | |
| run 2 | 30 | 0.960 [0.925, 0.985] 0.300 [0.225, 0.380] | 0.860 [0.795, 0.920] 0.285 [0.205, 0.370] | 0.795 [0.730, 0.860] 0.325 [0.245, 0.405] | 1.000 / 0.960 / 0.960 |
| 100 | 0.995 [0.985, 1.000] 0.335 [0.255, 0.415] | 0.950 [0.915, 0.980] 0.375 [0.290, 0.465] | 0.905 [0.860, 0.945] 0.435 [0.350, 0.520] | 0.995 / 0.985 / 0.990 |
| real images | gray canvas | |||||
|---|---|---|---|---|---|---|
| run | step | density | level [95% CI] | vs untrained [95% CI] | level [95% CI] | criteria met |
| untrained 7B | 0 | 8-pt | 0.585 [0.500, 0.665] | – | 0/200 (100 programs) | 4 of 5 |
| untrained 7B | 0 | 12-pt | 0.560 [0.470, 0.645] | – | 0/200 (100 programs) | 4 of 5 |
| untrained 7B | 0 | 20-pt | 0.425 [0.335, 0.515] | – | 0/200 (100 programs) | 5 of 5 |
| Real , run 1 | 30 | 8-pt | 0.935 [0.895, 0.970] | 0.350 [0.275, 0.430] | 0/200 (100 programs) | 4 of 5 |
| Real , run 1 | 30 | 12-pt | 0.860 [0.805, 0.910] | 0.300 [0.225, 0.380] | 0/200 (100 programs) | 4 of 5 |
| gray-canvas count | |||||||
|---|---|---|---|---|---|---|---|
| set | density | task | Real [95% CI] | control [95% CI] | [95% CI] | Real | control |
| dev | 8-pt | L3 | 0.305 [0.230, 0.380] | 0.030 [ 0.005, 0.065] | 0.275 [0.205, 0.350] | 0/200 | 11/200 |
| dev | 8-pt | probe | 0.060 [0.025, 0.105] | 0.005 [ 0.020, 0.035] | 0.055 [0.020, 0.095] | 0/200 | 0/200 |
| dev | 12-pt | L3 | 0.315 [0.230, 0.400] | 0.035 [ 0.080, 0.005] | 0.350 [0.265, 0.435] | 0/200 | 9/200 |
| dev | 12-pt | probe | 0.065 [0.025, 0.110] | 0.005 [ 0.035, 0.020] | 0.070 [0.030, 0.120] | 0/200 | 0/200 |
| dev | 20-pt | L3 | 0.330 [0.255, 0.410] | 0.030 [ 0.070, 0.010] | 0.360 [0.275, 0.445] | 0/200 | 8/200 |
| run | step | grounding task (600 pairs) | twin (600) | cued-readout task (300) | header-cued table (300) |
|---|---|---|---|---|---|
| untrained 7B | 0 | 0.768 [0.733, 0.802] | 0.728 [0.692, 0.762] | 0.673 [0.620, 0.727] | 0.993 [0.983, 1.000] |
| run 1 | 30 | 0.930 [0.908, 0.950] 0.162 [0.130, 0.193] | 0.900 0.172 | 0.680 [0.627, 0.733] 0.007 [ 0.023, 0.037] | 0.997 [0.990, 1.000] 0.003 [0.000, 0.010] |
| 100 | 0.948 [0.930, 0.965] 0.180 [0.147, 0.215] | 0.937 [0.917, 0.955] 0.208 [0.175, 0.243] | 0.707 [0.657, 0.757] 0.033 [0.000, 0.070] | 0.997 [0.990, 1.000] 0.003 [0.000, 0.010] | |
| run 2 | 30 | 0.935 [0.915, 0.955] 0.167 [0.137, 0.198] | 0.905 0.177 | 0.690 [0.640, 0.743] 0.017 [ 0.013, 0.047] | 0.997 [0.990, 1.000] 0.003 [0.000, 0.010] |
| 100 | 0.962 [0.947, 0.977] 0.193 [0.162, 0.227] | – | 0.727 [0.677, 0.777] 0.053 [0.020, 0.087] | 0.997 [0.990, 1.000] 0.003 [0.000, 0.010] |
| task | untrained | Real run 1, 30 | Real run 2, 30 | Gray -trained control, 30 | Real run 1, 100 |
|---|---|---|---|---|---|
| L1 cued readout | 0.610 [0.525, 0.695] | 0.820 [0.755, 0.880] 0.210 [0.140, 0.285] | 0.810 [0.745, 0.870] 0.200 [0.130, 0.270] | 0.640 [0.555, 0.720] 0.030 [ 0.010, 0.075] | 0.890 [0.835, 0.935] 0.280 [0.200, 0.360] |
| L2 find-and-bind | 0.620 [0.530, 0.705] | 0.900 [0.850, 0.945] 0.280 [0.200, 0.360] | 0.875 [0.815, 0.930] 0.255 [0.180, 0.330] | 0.680 [0.600, 0.760] 0.060 [0.020, 0.105] | 0.950 [0.915, 0.980] 0.330 [0.250, 0.415] |
| L3 discovery | 0.425 [0.335, 0.515] | 0.740 [0.665, 0.815] 0.315 [0.235, 0.400] | 0.720 [0.645, 0.795] 0.295 [0.215, 0.375] | 0.385 [0.300, 0.475] 0.040 [ 0.085, 0.000] | 0.875 [0.815, 0.930] 0.450 [0.365, 0.540] |
| identification probe | 0.825 [0.755, 0.890] | 0.975 [0.940, 1.000] 0.150 [0.090, 0.215] | 0.975 [0.940, 1.000] 0.150 [0.090, 0.215] | 0.850 [0.785, 0.905] 0.025 [0.000, 0.050] | 0.985 [0.960, 1.000] 0.160 [0.100, 0.225] |
| seed | condition | discovery, confirmatory | discovery, untouched | grounding task | twin | target-switch, conf. | target-switch, untouched | find-and-bind |
|---|---|---|---|---|---|---|---|---|
| – | untrained 7B | 0.425 (85/200) | 0.471 (628/1,332) | 0.768 | 0.728 | 13/50 | 0.318 | 0.620 |
| 0 | Real , run 1 | 0.740 | 0.791 | 0.930 | 0.900 | 30/50 | 0.586 | 0.900 |
| 0 | Real , run 2 (same seed) | 0.720 | 0.770 | 0.935 | 0.905 | 27/50 | 0.562 | 0.875 |
| 0 | Gray -trained | 0.385 | 0.440 | 0.770 | 0.730 | 12/50 | 0.252 | 0.680 |
| 1 | Real | 0.760 (152/200) | 0.792 (1,055/1,332) | 0.907 | 0.885 | 30/50 | 0.583 | 0.915 |
| 1 | Real , step 100 | 0.895 (179/200) | 0.900 (1,199/1,332) | 0.942 | 0.925 | – | – | 0.960 |
| quantity | endpoint | seeds | per-seed values | mean | seed SD | range | seed-level 95% interval |
|---|---|---|---|---|---|---|---|
| Real level | discovery, confirmatory | 0–3 | 0.740, 0.760, 0.730, 0.755 | 0.746 | 0.014 | 0.030 | – |
| Real Gray | discovery, confirmatory | 0–3 | 0.355, 0.345, 0.380, 0.335 | 0.354 | 0.019 | 0.045 | [ 0.323, 0.384] |
| Real Gray | discovery, untouched | 0–3 | 0.351, 0.341, 0.355, 0.309 | 0.339 | 0.021 | 0.046 | [ 0.306, 0.372] |
| Real untrained | discovery, confirmatory | 0–3 | 0.315, 0.335, 0.305, 0.330 | 0.321 | 0.014 | – | – |
| Real untrained | discovery, untouched | 0–3 | 0.320, 0.321, 0.297, 0.291 | 0.307 | 0.015 | – | – |
| Gray untrained | discovery, confirmatory | 0–3 | 0.040, 0.010, 0.075, 0.005 | 0.033 | 0.032 | – | – |
| subset | members | untrained | Real | Gray -trained |
|---|---|---|---|---|
| probe correct | 165 | 0.485 (80/165) | 0.764 (126/165) | 0.436 (72/165) |
| probe wrong | 35 | 0.143 (5/35) | 0.629 (22/35) | 0.143 (5/35) |
| all | 200 | 0.425 (85/200) | 0.740 (148/200) | 0.385 (77/200) |
| seed | test input | discovery, confirmatory | discovery, untouched | grounding task | gray test images |
|---|---|---|---|---|---|
| 1 | image only | 0.550 (110/200) | 0.589 (784/1,332) | 0.820 (492/600) | 0/200 |
| 2 | image only | 0.515 (103/200) | 0.553 (737/1,332) | 0.802 (481/600) | 0/200 |
| 3 | image only | 0.570 (114/200) | 0.595 (793/1,332) | 0.812 (487/600) | 0/200 |
| 1 | image and caption | 0.645 (129/200) | 0.669 (891/1,332) | – | – |
| 2 | image and caption | 0.610 (122/200) | 0.646 (861/1,332) | – | – |
| 3 | image and caption | 0.655 (131/200) | 0.682 (908/1,332) | – | – |
| quantity | endpoint | per-seed values (1, 2, 3) | mean | seed SD | seed-level 95% interval | Real reference |
|---|---|---|---|---|---|---|
| shortcut level | discovery, confirmatory | 0.550, 0.515, 0.570 | 0.545 | 0.028 | [0.476, 0.614] | 0.760, 0.730, 0.755 |
| shortcut untrained | discovery, confirmatory | 0.125, 0.090, 0.145 | 0.120 | 0.028 | [ 0.051, 0.189] | 0.335, 0.305, 0.330 |
| Real shortcut | discovery, confirmatory | 0.210, 0.215, 0.185 | 0.203 | 0.016 | [ 0.163, 0.243] | Real Gray : 0.345, 0.380, 0.335 |
| shortcut level | discovery, untouched | 0.589, 0.553, 0.595 | 0.579 | 0.023 | [0.523, 0.635] | 0.792, 0.769, 0.762 |
| Real shortcut | discovery, untouched | 0.203, 0.215, 0.167 | 0.195 | 0.025 | [ 0.132, 0.258] | Real Gray : 0.341, 0.355, 0.309 |
| shortcut untrained | grounding task | 0.052, 0.033, 0.043 | 0.043 | 0.009 | [ 0.020, 0.066] | 0.139, 0.154, 0.125 |
| in-band | 8-pt: level; gain [95% CI] | 12-pt | 20-pt (never trained) | grounding [95% CI] | |||
|---|---|---|---|---|---|---|---|
| 0 | 0.126 | 0.094 | 0.752 | 0.740 0.080 [0.030, 0.135] | 0.570 0.005 [ 0.035, 0.030] | 0.460 0.010 [ 0.040, 0.025] | 0.797 [0.763, 0.828] |
| 0.165 | 0.229 | 0.900 | 0.935 0.275 [0.200, 0.355] | 0.805 0.230 [0.155, 0.305] | 0.745 0.275 [0.205, 0.350] | 0.908 [0.885, 0.930] | |
| 0.199 | 0.240 | 0.926 | 0.950 0.290 [0.215, 0.370] | 0.850 0.275 [0.195, 0.360] | 0.740 0.270 [0.200, 0.345] | 0.927 [0.905, 0.947] | |
| 1 | 0.249 | 0.251 | – | 0.965 0.305 [0.230, 0.380] | 0.890 0.315 [0.230, 0.400] | 0.800 0.330 [0.255, 0.410] | 0.930 [0.908, 0.950] |
| endpoint | slope on [95% CI] | slope on [95% CI] |
|---|---|---|
| 8-pt discovery | 1.670 [1.187, 2.190] | 1.441 [1.005, 1.907] |
| 12-pt discovery | 2.446 [1.813, 3.088] | 1.942 [1.432, 2.471] |
| 20-pt discovery (never trained) | 2.482 [1.886, 3.083] | 2.079 [1.565, 2.596] |
| grounding task | 1.009 [0.787, 1.233] | 0.861 [0.669, 1.054] |
| discovery, real images | identification probe | discovery | ||||||
|---|---|---|---|---|---|---|---|---|
| model | step | density | level [95% CI] | vs untrained [95% CI] | real | gray canvas | gray canvas | grounding task [95% CI] |
| 3B | 0 | 8-pt | 0.330 [0.255, 0.405] | – | 0.705 | 0.000 | 0.080 | 0.455 [0.415, 0.495] |
| 12-pt | 0.260 [0.190, 0.340] | – | 0.680 | 0.000 | 0.025 | |||
| 20-pt | 0.245 [0.175, 0.320] | – | 0.630 | 0.000 | 0.055 | |||
| 10 | 8-pt | 0.475 [0.395, 0.555] | 0.145 [0.065, 0.225] | 0.875 | 0.000 | 0.080 | – | |
| 12-pt | 0.425 [0.340, 0.510] | 0.165 [0.085, 0.245] | 0.755 | 0.000 | 0.025 | |||
| benchmark | model | with image | no image | ||
|---|---|---|---|---|---|
| MathVista testmini | untrained 7B | 0.682 (682/1,000) | 0.366 (366/1,000) | 0.179 | 0.372 |
| MathVista testmini | , step 30 | 0.532 (532/1,000) | 0.258 (258/1,000) | 0.179 | 0.224 |
| MathVista testmini | , step 30 | 0.696 (696/1,000) | 0.367 (367/1,000) | 0.179 | 0.364 |
| BLINK | untrained 7B | 0.536 (1,018/1,901) | 0.416 (791/1,901) | 0.377 | 0.247 |
| BLINK | , step 30 | 0.457 (869/1,901) | 0.363 (690/1,901) | 0.377 | 0.175 |
| BLINK | , step 30 | 0.547 (1,039/1,901) | 0.420 (798/1,901) | 0.377 | 0.252 |
| condition | seed | rows | step-1 accuracy | step-1 format | mixed, step 1 | nonzero advantage, step 1 | mixed, steps 1–5 | nonzero advantage, steps 1–5 | all correct, step 1 | all wrong, step 1 |
|---|---|---|---|---|---|---|---|---|---|---|
| Real | 1 | discovery | 0.223 | 0.857 | 0.633 | 0.808 | 0.645 | 0.698 | 0.008 | 0.358 |
| Real | 1 | probe | 0.435 | 0.875 | 0.892 | 0.933 | 0.612 | 0.622 | 0.025 | 0.083 |
| Real | 2 | discovery | 0.178 | 0.883 | 0.575 | 0.800 | 0.632 | 0.690 | 0.008 | 0.417 |
| Real | 2 | probe | 0.407 | 0.892 | 0.875 | 0.942 | 0.598 | 0.610 | 0.017 | 0.108 |
| Real | 3 | discovery | 0.245 | 0.875 | 0.667 | 0.842 | 0.655 | 0.718 | 0.008 | 0.325 |
| Real | 3 | probe | 0.428 | 0.890 | 0.892 | 0.933 | 0.622 | 0.637 | 0.025 | 0.083 |
| component | scale |
|---|---|
| constructed evaluations | a 1,200-pair suite and its regenerated 1,200-pair twin; discovery scenes at 8, 12 and 20 points, with 450 confirmatory scene programs; an untouched replication set of 333 scene programs per pair role (Table G.6 ); 160 premise groups (an untrained nearest-point relation task); 100 label-swap diagnostic pairs |
| constructed training data | 2,880 discovery and probe prompts; a second corpus of 3,000 pairs and 300 answer-preserving twins |
| public image corpora | Geometry3K (1,288 training, 601 test items); ViRL39K (23,542 training, 4,239 held-out items) |
| public benchmarks | seven, more than 10,000 items, audited with and without the image at 3B and 7B (3B: Table A.1 ) |
| models | Qwen2.5-VL-3B and 7B trained; Gemma-3 and InternVL3-9B audited (Table A.2 ) |
| training runs | over thirty GRPO runs of 30 to 400 steps, one to five runs per condition (main-text runs: Table D.1 ) |