cs.CLJun 24, 2026

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Authors: Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai

Organizations: Meituan · Tsinghua University

Abstract

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score SS extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of SS, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at https://github.com/wangq-dev/CDMScaling.

Explore similar work

CardsList
  1. Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

    Oct 1, 2026Xinye Zhao, Yunkai Dang, Yunchen Wu +1Vision-Language ModelsEfficient VLM Inference

  2. On Test-Time Scaling for Vision-Language Models

    Jun 27, 2026Fawaz Sammani, Tzoulio Chamiti, Nikos DeligiannisVLM EvaluationTest-Time Scaling for VLMs

  3. ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

    Aug 4, 2026Yang Yang, Qinyu Zhao, Mouxiang Chen +5Vision-Language ModelsMultimodal Large Language Models