Selective Transfer of RL Updates for Visual Reasoning
Organizations: University of Pennsylvania · Independent Researcher · University of Electronic Science and Technology of China · Chongqing University
Abstract
Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at https://anonymous.4open.science/r/selective-rl.
Figures & tables
| Method | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|
| H1: Qwen3-VL-8B-Instruct | |||||
| Receiver | |||||
| Task Arithmetic | 76.10_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.10}} | 44.50_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.50}} | 39.47_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.97}} | 65.67_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.64}} | 54.67_{{\color[rgb]{0,0.5547,0.6133}\downarrow 0.13}} |
| Full interpolation | 76.70_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.70}} | 45.30_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 3.30}} | 40.46_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.96}} | 66.07_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.04}} | 55.93_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.13}} |
| TIES | 76.40_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.40}} | 44.80_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.80}} | 39.80_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.30}} | 65.77_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.74}} | 55.53_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.73}} |
| DARE | 76.30_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.30}} | 45.00_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 3.00}} | 40.13_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.63}} | 65.97_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.94}} | 55.73_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.93}} |
| Construction | Math Vista | Math Verse | DynaMath | MMStar | ||
|---|---|---|---|---|---|---|
| (a) Composition: initialization displacement and RL update | ||||||
| Receiver | 0.00 | 0.00 | 75.00 | 42.00 | 65.03 | 54.80 |
| Initialization only | 0.05 | 0.00 | 74.70 | 41.50 | 64.65 | 54.07 |
| Direct full | 0.00 | 0.05 | 76.10 | 44.50 | 65.67 | 54.67 |
| Direct rank-1 | 0.00 | 0.05 | 79.20 | 50.20 | 69.06 | 58.80 |
| Interpolation full | 0.05 | 0.05 | 76.70 | 45.30 | 66.07 | 55.93 |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| 1 | Input: recipient , actual pre-RL checkpoint , donor , matrix support , coefficients . |
|---|---|
| 2 | Verify aligned tensor coordinates and consistent support and updates for tied parameters. |
| 3 | . // copy all tensors |
| 4 | for each do |
| 5 | ; . // isolate RL |
| 6 | . |
| 7 | if then // zero RL term |
| Configuration | H1: primary | H2: cross-family | H3: cross-backbone |
|---|---|---|---|
| Receiver | Qwen3-VL-8B-Instruct | LLaVA-NeXT-LLaMA3-8B | Idefics2-8B |
| LM family | Qwen3 | LLaMA3 | Mistral |
| Layers / matrices | 36 / 252 | 32 / 224 | 32 / 224 |
| Donor seed | 42 | 42 | 42 |
| Selected step | 120 | 100 | 110 |
| Training budget | 150 steps | 150 steps | 150 steps |
| Benchmark | Metric | |
|---|---|---|
| MathVista | 1,000 | Accuracy (%) |
| MathVerse | 1,000 | Accuracy (%) |
| MathVision | 304 | Accuracy (%) |
| DynaMath | 5,010 | Accuracy (%) |
| MMStar | 1,500 | Accuracy (%) |
| TextVQA | 1,000 | VQA soft score |
| MathVista-val | DynaMath-val | |
|---|---|---|
| 0.01 | 76.90 | 66.10 |
| 0.025 | 79.00 | 68.35 |
| 0.05 | 80.20 | 69.40 |
| 0.075 | 79.90 | 69.10 |
| 0.1 | 79.50 | 68.75 |
| Method | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|
| H1: Qwen3-VL-8B-Instruct | |||||
| Receiver | |||||
| Task Arithmetic | 76.10_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.10}} | 44.50_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.50}} | 39.47_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.97}} | 65.67_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.64}} | 54.67_{{\color[rgb]{0,0.5547,0.6133}\downarrow 0.13}} |
| Full interpolation | 76.70_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.70}} | 45.30_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 3.30}} | 40.46_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.96}} | 66.07_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.04}} | 55.93_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.13}} |
| TIES | 76.40_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.40}} | 44.80_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.80}} | 39.80_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.30}} | 65.77_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.74}} | 55.53_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.73}} |
| DARE | 76.30_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 1.30}} | 45.00_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 3.00}} | 40.13_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 2.63}} | 65.97_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.94}} | 55.73_{{\color[rgb]{0.832,0.3633,0.1758}\uparrow 0.93}} |
| Construction | Math Vista | Math Verse | Math Vision | DynaMath | MMStar | ||
|---|---|---|---|---|---|---|---|
| (a) Composition: initialization displacement and RL update | |||||||
| Receiver | 0.00 | 0.00 | 75.00 | 42.00 | 37.50 | 65.03 | 54.80 |
| Initialization only | 0.05 | 0.00 | 74.70 | 41.50 | 37.17 | 64.65 | 54.07 |
| Direct full | 0.00 | 0.05 | 76.10 | 44.50 | 39.47 | 65.67 | 54.67 |
| Direct rank-1 | 0.00 | 0.05 | 79.20 | 50.20 | 47.04 | 69.06 | 58.80 |
| Interpolation full | 0.05 | 0.05 | 76.70 | 45.30 | 40.46 | 66.07 | 55.93 |
| Rank | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|
| 1 | 81.00 | 52.00 | 49.01 | 70.06 | 59.67 |
| 2 | 80.70 | 51.60 | — | 69.66 | 59.27 |
| 4 | 80.30 | 50.80 | 47.70 | 69.06 | 58.67 |
| 8 | 79.50 | 49.40 | — | 68.26 | 58.00 |
| 16 | 78.30 | 47.60 | — | 67.27 | 57.00 |
| Full | 76.70 | 45.30 | 40.46 | 66.07 | 55.93 |
| MathVista Full | MathVista Rank-1 | DynaMath Full | DynaMath Rank-1 | MMStar Full | MMStar Rank-1 | |
|---|---|---|---|---|---|---|
| 0 | 75.00 | 75.00 | 65.03 | 65.03 | 54.80 | 54.80 |
| 0.01 | 75.40 | 77.40 | 65.27 | 66.67 | 55.33 | 56.60 |
| 0.025 | 76.20 | 79.90 | 65.71 | 69.26 | 55.67 | 58.60 |
| 0.05 | 76.70 | 81.00 | 66.07 | 70.06 | 55.93 | 59.67 |
| 0.1 | 76.00 | 80.50 | 65.47 | 69.66 | 54.93 | 59.27 |
| 0.2 | 74.40 | 78.90 | 63.87 | 68.26 | 53.00 | 57.73 |
| Support | Matrices | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|---|
| QKV only | 108 | 78.10 | 47.60 | — | — | 57.47 |
| Attention only | 144 | 79.40 | 49.30 | — | 68.52 | 58.47 |
| MLP only | 108 | 78.80 | 48.60 | — | 68.88 | 57.93 |
| O + Down only | 72 | 77.90 | 46.90 | — | 67.54 | 57.20 |
| Attention + MLP | 252 | 81.00 | 52.00 | 49.01 | 70.06 | 59.67 |
| Configuration | Attention median | MLP median | Overall | Q1 | Q3 | Median amplification |
|---|---|---|---|---|---|---|
| H1 Qwen RL | 0.24 | 0.15 | 0.18 | 0.11 | 0.31 | 2.36 |
| H2 LLaMA RL | 0.17 | 0.11 | 0.13 | 0.08 | 0.24 | 2.77 |
| H3 Mistral RL | 0.12 | 0.08 | 0.09 | 0.05 | 0.17 | 3.33 |
| Module | Mean | Median | Q1 | Q3 |
|---|---|---|---|---|
| q_proj | 0.27 | 0.25 | 0.17 | 0.35 |
| k_proj | 0.24 | 0.22 | 0.15 | 0.31 |
| v_proj | 0.25 | 0.23 | 0.16 | 0.33 |
| o_proj | 0.21 | 0.19 | 0.12 | 0.28 |
| gate_proj | 0.16 | 0.14 | 0.09 | 0.22 |
| up_proj | 0.15 | 0.13 | 0.08 | 0.21 |
| Layer region | Median | Median amplification |
|---|---|---|
| Early 1-12 | 0.13 | 2.77 |
| Middle 13-24 | 0.22 | 2.13 |
| Late 25-36 | 0.19 | 2.29 |
| Normalization | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|
| None | 75.20 | 42.60 | 38.16 | 65.27 | 54.80 |
| Nuclear-norm match | 80.00 | 50.10 | — | 69.16 | 58.67 |
| Frobenius match | 81.00 | 52.00 | 49.01 | 70.06 | 59.67 |
| Reference control | Math Vista | Math Verse | Math Vision | DynaMath | MMStar |
|---|---|---|---|---|---|
| Full, actual pre-RL | 76.70 | 45.30 | 40.46 | 66.07 | 55.93 |
| Rank-1, actual pre-RL | 81.00 | 52.00 | 49.01 | 70.06 | 59.67 |
| Full, earlier reference | 75.60 | 43.80 | — | 65.41 | 54.93 |
| Rank-1, earlier reference | 77.80 | 46.90 | — | 67.11 | 56.40 |
| MathVista | DynaMath | MMStar | |
|---|---|---|---|
| 0.05 | 75.20 | 65.27 | 54.80 |
| 0.1 | 76.40 | 65.87 | 55.47 |
| 0.2 | 77.60 | 66.91 | 56.33 |
| 0.5 | 78.50 | 67.52 | 56.87 |
| 1 | 77.80 | 66.95 | 56.40 |
| 2 | 74.90 | 63.95 | 53.87 |
| Control | MathVista | DynaMath | MMStar |
|---|---|---|---|
| Full, global beta | 76.70 | 66.07 | 55.93 |
| Full, rank-1 matrixwise gains | 77.60 | 66.93 | 56.33 |
| Rank-1, tuned global beta | 78.50 | 67.52 | 56.87 |
| Rank-1, Frobenius matched | 81.00 | 70.06 | 59.67 |
| Direction seed | MathVista | DynaMath | MMStar |
|---|---|---|---|
| 0 | 74.60 | 64.43 | 53.27 |
| 1 | 73.80 | 63.87 | 52.87 |
| 2 | 74.10 | 64.07 | 53.13 |
| 3 | 73.40 | 63.55 | 52.47 |
| 4 | 74.70 | 64.51 | 53.60 |
| 5 | 73.90 | 63.91 | 52.93 |
| Reconstruction | MATH | GSM8K | Accuracy retention (%) | Gain retention (%) |
|---|---|---|---|---|
| Pre-RL | 62.00 | 84.99 | — | — |
| Donor | 76.20 | 90.98 | 100.00 | 100.00 |
| rank-1 | 75.80 | 90.74 | 99.48 | 97.18 |
| rank-2 | 75.40 | 90.60 | 98.95 | 94.37 |
| rank-4 | 76.00 | 90.82 | 99.74 | 98.59 |
| rank-8 | 75.60 | 90.67 | 99.21 | 95.77 |
| Stage | Source MATH | MathVista Full | MathVista Rank-1 | DynaMath Full | DynaMath Rank-1 | MMStar Full | MMStar Rank-1 |
|---|---|---|---|---|---|---|---|
| Early | 69.60 | 75.80 | 77.90 | 65.47 | 66.91 | 55.33 | 56.47 |
| Middle | 75.60 | 76.60 | 79.90 | 65.87 | 68.72 | 55.80 | 58.33 |
| Late | 76.20 | 76.70 | 81.00 | 66.07 | 70.06 | 55.93 | 59.67 |
| Source | Transfer | MathVista | DynaMath | MMStar | Median |
|---|---|---|---|---|---|
| RL | Full | 76.70 | 66.07 | 55.93 | — |
| RL | Rank-1 | 81.00 | 70.06 | 59.67 | 0.18 |
| Offline SFT | Full | 77.10 | 66.67 | 56.27 | — |
| Offline SFT | Rank-1 | 78.00 | 67.15 | 56.13 | 0.10 |
| Offline distillation | Full | 77.60 | 67.27 | 56.67 | — |
| Offline distillation | Rank-1 | 78.80 | 68.18 | 56.80 | 0.12 |
| Seed | Full MathVista | Selective MathVista | MathVista gain | DynaMath gain | MMStar gain |
|---|---|---|---|---|---|
| 17 | 76.60 | 80.70 | |||
| 42 | 76.70 | 81.00 | |||
| 73 | 76.50 | 80.50 | |||
| Mean SD |
| Task | Gain (pp) | 95% CI | |||||
|---|---|---|---|---|---|---|---|
| MathVista | 727 | 40 | 83 | 150 | 1000 | [2.14, 6.46] | |
| MathVerse | 410 | 43 | 110 | 437 | 1000 | [4.31, 9.09] | |
| MathVision | 108 | 15 | 41 | 140 | 304 | [3.82, 13.29] | |
| DynaMath | 3197 | 113 | 313 | 1387 | 5010 | [2.10, 5.86] | |
| MMStar | 800 | 39 | 95 | 566 | 1500 | [2.23, 5.23] |
| Statistic | Full | Selective-RL |
|---|---|---|
| All-ten-correct (%) | 29.1 | 34.3 |
| Median correct variants | 7.0 | 8.0 |
| Method | Mean tokens | P95 tokens | Repeat | Extract fail | At cap |
|---|---|---|---|---|---|
| Receiver | 460 | 1550 | 8 | 5 | 6 |
| Full | 575 | 2200 | 15 | 8 | 14 |
| IP-Merging | 590 | 2050 | 11 | 6 | 10 |
| FRISM | 640 | 2200 | 7 | 4 | 8 |
| STAR | 520 | 1820 | 8 | 4 | 7 |
| AdaRank | 495 | 1740 | 6 | 4 | 6 |
| Full error category | Count | Share (%) |
|---|---|---|
| Repetition / loop | 8 | 9.6 |
| Token-cap truncation | 6 | 7.2 |
| Extraction pathology | 4 | 4.8 |
| Reasoning / perception | 65 | 78.3 |
| Statistic | Receiver | Full | Selective-RL |
|---|---|---|---|
| Audited examples | 200.0 | 200.0 | 200.0 |
| Answer/extraction disagreement | 6.0 | 7.0 | 5.0 |
| False-positive extraction | 2.0 | 3.0 | 2.0 |
| False-negative extraction | 4.0 | 4.0 | 3.0 |
| Disagreement (%) | 3.0 | 3.5 | 2.5 |
| Benchmark | Receiver | Full | Selective-RL | vs Full |
|---|---|---|---|---|
| TextVQA | 82.03 | 81.06 | 83.64 | |
| MMMU | 60.00 | 59.20 | 62.80 | |
| OCRBench | 78.60 | 78.00 | 78.70 | |
| POPE | 86.90 | 86.30 | 86.80 |
| Method | SVD | Other | Write | Calib. | Total | Search (GPU-h) |
|---|---|---|---|---|---|---|
| Receiver | 0 | 0 | 0 | 0 | 0 | 0.0 |
| Full | 0 | 0 | 35 | 0 | 35 | 3.0 |
| TIES | 0 | 40 | 35 | 0 | 75 | 3.0 |
| DARE | 0 | 27 | 35 | 0 | 62 | 3.0 |
| TSV | 260 | 90 | 35 | 0 | 385 | 3.0 |
| STAR | 285 | 75 | 35 | 0 | 395 | 3.0 |
| Method | Dense ratio | Added module | VRAM (GB) |
|---|---|---|---|
| Receiver | 1.00 | No | 24.60 |
| Full | 1.00 | No | 24.60 |
| STAR | 1.00 | No | 24.60 |
| AdaRank | 1.00 | No | 24.60 |
| FRISM | 1.00 | No | 24.60 |
| SELECTIVE-RL | 1.00 | No | 24.60 |