cs.CVOct 1, 2026

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

Authors: Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li

Organizations: School of Intelligence Science and Technology, Nanjing University

Abstract

Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.

Figures & tables

Appendix figures & tables45 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Test-Time Scaling for Vision-Language Models

    Jun 27, 2026Fawaz Sammani, Tzoulio Chamiti, Nikos DeligiannisTest-Time ScalingRecent Vision-Language Models

  2. AVIS: Adaptive Test-Time Scaling for Vision-Language Models

    Jun 10, 2026Ahmadreza Jeddi, Minh Ngoc Le, Amirhossein Kazerouni +8Fast InferenceEfficient Inference

  3. The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

    Jun 5, 2026Lujun Li, Lama Sleem, Niccolo Gentile +4Recent Vision-Language ModelsFine-Grained Perception