What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
Organizations: Meituan · Tsinghua University
Abstract
Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of , with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at https://github.com/wangq-dev/CDMScaling.
Figures & tables
| Variant | 8B MAE | 72B-Base MAE | 72B-Inst MAE |
| 1.266% | 2.374% | 0.476% | |
| 1.259% | 2.055% | 0.394% | |
| Type | ( ) | ( ) | ( ) | MAE (%) | |
| Base | 0.212 | 1.60 | 0.78 | 1.61 | 1.33 |
| Chat | 0.254 | 1.49 | 1.04 | 1.47 | 1.11 |
| Family | ( ) | ( ) | ( ) | MAE (%) | |
| Llama | 0.361 | 0.28 | 4.55 | 1.22 | 0.90 |
| Gemma | 0.324 | 2.46 | 4.51 | 2.13 | 1.02 |
| Falcon3 | 0.288 | 1.12 | 1.33 | 1.17 | 0.60 |
| Qwen3 | 0.246 | 1.40 | 2.19 | 1.26 | 1.11 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Model | Type | Size |
| Qwen3 | Qwen3-0.6B-Base | Base | 0.6B |
| Qwen3-0.6B | Instruct | 0.6B | |
| Qwen3-1.7B-Base | Base | 1.7B | |
| Qwen3-1.7B | Instruct | 1.7B | |
| Qwen3-4B-Base | Base | 4B | |
| Qwen3-4B | Instruct | 4B |
| Category | Benchmark | Source |
| Information Extraction | CrossNER | [ Liu et al., 2021 ] |
| Knowledge | hindu_knowledge | BB |
| mmlu_stem | MMLU | |
| mmlu_humanities | MMLU | |
| mmlu_other | MMLU | |
| dark_humor_detection | BB |
| Category | Benchmark |
| General VQA | MMBench_DEV_EN_V11 [ Liu et al., 2024c ] |
| MMStar [ Chen et al., 2024 ] | |
| SEEDBench_IMG [ Li et al., 2023b ] | |
| SEEDBench2 [ Li et al., 2023a ] | |
| SEEDBench2_Plus [ Li et al., 2024d ] | |
| MME [ Fu et al., 2026 ] |
| Rank | Benchmark | Rank | Benchmark | ||
| Positive Transfer ( ) | Positive Transfer (Continued) | ||||
| 1 | dyck_languages_hard | +0.119 | 21 | checkmate_in_one | +0.019 |
| 2 | hindu_knowledge | +0.108 | 22 | undo_permutation | +0.019 |
| 3 | play_dialog_same_or_different | +0.082 | 23 | mmlu_other | +0.016 |
| 4 | word_unscrambling | +0.063 | 24 | causal_judgement_hard | +0.015 |
| 5 | mmlu_stem | +0.057 | 25 | winowhy | +0.015 |
| Stage-1 | Stage-2 | ||
| Vision | Resolution | 384 | 384 {(1 1),…,(6 6)} |
| #tokens | 729 | Max 10 729 | |
| Data | Samples | 5M | 12M |
| Model | Trainable | Projector | Full Model |
| Training | Batch Size | 512 | 512 |
| LR | |||
| Category | ( ) | ( ) | ( ) | MAE (%) | |
| General VQA | 0.322 | 1.502 | 1.056 | 1.501 | 1.552 |
| Document Understanding | 0.228 | 2.601 | 0.917 | 2.601 | 1.974 |
| STEM Puzzle | 0.212 | 1.665 | 0.098 | 1.665 | 1.583 |
| Alignment | 0.245 | 0.728 | 2.309 | 0.728 | 0.996 |
| Stable Positive Transfer | |||||||||
| Benchmark | Mean | Std. | Pos. | Neg. | Benchmark | Mean | Std. | Pos. | Neg. |
| dyck_languages_hard | 0.374 | 0.156 | 6 | 0 | symbol_interpretation | 0.118 | 0.038 | 6 | 0 |
| matrixshapes | 0.341 | 0.106 | 5 | 0 | disfl_qa | 0.116 | 0.048 | 6 | 0 |
| hindu_knowledge | 0.282 | 0.091 | 6 | 0 | wic | 0.116 | 0.109 | 5 | 1 |
| word_unscrambling | 0.261 | 0.052 | 6 | 0 | mmlu_humanities | 0.106 | 0.062 | 5 | 1 |
| winogrande | 0.217 | 0.097 | 5 | 0 | winowhy | 0.105 | 0.144 | 5 | 1 |
| Family | Backbone | (B) | Interpretation | |
| LLaMA-3.2 | Llama-3.2-1B | +0.828 | 85.2 | Crossover |
| Gemma-2 | Gemma-2-9B | +0.903 | 234.5 | Crossover |
| Qwen3 | Qwen3-30B-A3B | +1.117 | 587.8 | Large-scale crossover |
| Qwen3 | Qwen3-4B | +0.084 | – | Base-favorable |
| Falcon3 | Falcon3-3B | +0.024 | Weak / unstable | |
| Falcon3 | Falcon3-10B | -0.098 | – | Chat-favorable |