Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Organizations: School of Intelligence Science and Technology, Nanjing University
Abstract
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
Figures & tables
| Task type | Mean | Median | Floor | Saturation |
| Perception | 0.073 | 0.001 | 72% | 20% |
| Reasoning | 0.155 | 0.001 | 53% | 33% |
| Domain | Skill | ||
| D1 | Documents, Charts, Slides | S1 | Attribute Recognition |
| D2 | Aerial, Satellite | S2 | Text Reading |
| D3 | Vehicles, Driving | S3 | Counting |
| D4 | Indoor | S4 | Spatial Relation |
| D5 | Outdoor | S5 | Object Identification |
| D6 | People, Surveillance | S6 | Scene Reasoning |
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
| Version | Input strategy | Visual token mechanism | Effect on visual token count |
| InternVL 1.0 | Fixed input size | ViT features projected to the language backbone | Fixed count |
| InternVL 1.5 | Dynamic tiling | Up to 40 tiles of | Changes in whole tiles |
| InternVL2.5 | Dynamic tiling | Pixel unshuffle to 256 tokens per tile | Changes in whole tiles |
| InternVL3 | Dynamic tiling | 256 tokens per tile, V2PE positions | Changes in whole tiles |
| InternVL3.5 | Dynamic tiling | 256 tokens per tile, ViR only in Flash variants | Changes in whole tiles |
| Version | Input strategy | Visual token mechanism | Effect on visual token count |
| Qwen-VL | Fixed input | Cross-attention adapter to a fixed number of tokens | Fixed count |
| Qwen2-VL | Dynamic resolution | 2D-RoPE and MLP patch merger | Changes with input size in small steps |
| Qwen2.5-VL | Dynamic resolution | Window attention, same patch merger | Changes with input size in small steps |
| Qwen3-VL | Dynamic resolution | SigLIP-2, DeepStack, same patch merger | Changes with input size in small steps |
| Family | Series | Evaluated sizes | Visual frontend |
| InternVL | InternVL2.5 | 1B, 2B, 4B, 8B, 26B, 38B | Dynamic tiling |
| InternVL | InternVL3 | 1B, 2B, 8B, 9B, 14B, 38B | Dynamic tiling |
| InternVL | InternVL3.5 | 1B, 2B, 4B, 8B, 14B, 38B | Dynamic tiling |
| QwenVL | Qwen2.5-VL | 3B, 7B, 32B, 72B | Dynamic resolution |
| QwenVL | Qwen3-VL | 2B, 4B, 8B, 32B | Dynamic resolution |
| Benchmark | Native size ( ) | Input sizes (long edge, px) |
| HR-Bench-4K | ||
| HR-Bench-8K | ||
| MME-RealWorld | ||
| TreeBench | ||
| V * Bench |
| Group | Separable-law parameters | Fit diagnostics | |||||||
| Series | Benchmark | SSE | |||||||
| InternVL2.5 | HR-Bench | 0.001 | 0.308 | 0.149 | 1.24 | 4.66 | 0.031 | 42 | 0.92 |
| InternVL2.5 | MME-RealWorld | 0.001 | 0.527 | 0.058 | 11.41 | 0.043 | 36 | 0.77 | |
| InternVL2.5 | TreeBench | 0.001 | 0.582 | 0.057 | 9.65 | 0.030 | 36 | 0.73 | |
| InternVL2.5 | V * Bench | 0.001 | 0.089 | 0.381 | 56.31 | 0.203 | 36 | 0.59 | |
| InternVL3 | HR-Bench | 0.001 | 0.262 | 0.183 | 1.05 | 3.87 | 0.025 | 42 | 0.94 |
| Form | SSE | BIC | SSE | BIC | |
| Separable | 62 | 1.221 | — | — | |
| Multiplicative | 42 | 1.870 | |||
| Coupled | 83 | 1.182 | |||
| CES | 42 | 2.123 | |||
| -norm | 63 | 1.913 |
| Form | Pooled SSE | Median group | Groups with lower SSE |
| Separable | 1.088 | 0.892 | 19/20 |
| Multiplicative | 1.254 | 0.840 | 1/20 |
| Objective | Relative (%) | ||
| Unweighted | |||
| Weighted by precision |
| Exponents | Held-out axis | Separable | Coupled | RMSE (pp) |
| All cells | ||||
| Shared | Backbone size | |||
| Shared | Input size | |||
| Per-group | Backbone size | |||
| Per-group | Input size | |||
| Interior levels | ||||
| Held-out design | Eligible decisions | Mean regret (pp) | Same choice (%) | |
| Separable | Bounded | |||
| Backbone holdout | 297 | 0.48 | 0.48 | 100.0 |
| Input-size holdout | 411 | 0.66 | 0.67 | 99.3 |
| Joint holdout | 2,379 | 0.76 | 0.75 | 98.3 |
| Joint holdout, two-axis tradeoff | 1,284 | 0.88 | 0.87 | 97.1 |
| Family | Law | Median | Median | |||
| InternVL | Separable | 0.196 | 10.000 | 0 | 0 | 9 |
| InternVL | Bounded | 0.113 | 10.000 | 5 | 0 | 9 |
| QwenVL | Separable | 0.779 | 0.174 | 1 | 0 | 0 |
| QwenVL | Bounded | 0.664 | 0.074 | 2 | 3 | 0 |
| Held-out RMSE (pp) | Predictions outside | ||||
| Design | Held-out level | Separable | Bounded | Separable | Bounded |
| Backbone size | Interior | 4.47 | 4.43 | 0 | 0 |
| Backbone size | Endpoint | 7.38 | 7.18 | 6 | 0 |
| Input size | Interior | 4.56 | 4.45 | 0 | 0 |
| Input size | Endpoint | 8.29 | 11.79 | 0 | 0 |
| Joint holdout | Interior | 4.59 | 4.54 | 0 | 0 |
| Series | Benchmark | bootstrap interval | at search bound (%) | ||
| InternVL2.5 | HR-Bench | 3.2 | |||
| MME-RealWorld | 100.0 | ||||
| TreeBench | 95.4 | ||||
| V * Bench | 99.0 | ||||
| InternVL3 | HR-Bench | 1.8 | |||
| Series | Benchmark | ||||
| InternVL2.5 | HR-Bench | 0.17 | 0.001 | 0.91 | |
| InternVL2.5 | MME-RealWorld | 0.07 | 0.001 | 0.69 | |
| InternVL2.5 | TreeBench | 0.06 | 0.001 | 0.72 | |
| InternVL2.5 | V * Bench | 0.49 | 0.001 | 0.44 | |
| InternVL3 | HR-Bench | 0.21 | 0.001 | 0.94 | |
| InternVL3 | MME-RealWorld | 0.22 | 0.300 | 0.65 |
| Parameterization | Pooled RMSE (pp) | at upper bound | |
| Cell mean | |||
| Question moment |
| Series | Perception | Reasoning | Gap |
| InternVL2.5 | 0.001 | 0.067 | |
| InternVL3 | 0.097 | 0.187 | |
| InternVL3.5 | 0.180 | 0.194 | |
| Qwen2.5-VL | 0.001 | 0.187 | |
| Qwen3-VL | 0.085 | 0.141 | |
| Pooled | 0.073 | 0.155 |
| Series | Benchmark | Task | Exponent | |
| InternVL2.5 | HR-Bench | Cross | 0.099 | 5.18 |
| Single | 0.221 | 9.50 | ||
| MME-RealWorld | Perception | 0.068 | ||
| Reasoning | 0.065 | 7.05 | ||
| TreeBench | Perception | 0.115 | ||
| 22.3 / 52.5 / 25.2 | 20.0 / 54.8 / 25.2 | 17.4 / 57.4 / 25.2 | |
| 22.3 / 47.0 / 30.7 | 20.0 / 49.4 / 30.7 | 17.4 / 51.9 / 30.7 | |
| 22.3 / 44.4 / 33.2 | 20.0 / 46.8 / 33.2 | 17.4 / 49.3 / 33.2 |
| Ceiling Bound | Scaling Bound | Easy | ||||
| Series | Accuracy | Slope | Accuracy | Slope | Accuracy | Slope |
| InternVL2.5 | ||||||
| InternVL3 | ||||||
| InternVL3.5 | ||||||
| Qwen2.5-VL | ||||||
| Qwen3-VL | ||||||
| Criterion | ||||
| Low performance | 37.5 | 30.6 | 39.1 | 40.0 |
| Small endpoint change | 20.8 | 13.1 | 20.8 | 13.0 |
| Small late change | 23.5 | 18.3 | 23.4 | 18.1 |
| Small fitted slope | 19.0 | 14.1 | 18.3 | 13.9 |
| All three together | 16.4 | 9.5 | 16.3 | 8.2 |
| Group | Trajectories | Ceiling Bound | ||
| HR-Bench | 4,000 | 23.0 | 4.5 | 3.9 |
| MME-RealWorld | 8,900 | 39.3 | 12.4 | 10.8 |
| TreeBench | 1,975 | 40.8 | 9.7 | 8.3 |
| V * Bench | 955 | 17.6 | 2.4 | 1.6 |
| Perception | 10,525 | 30.0 | 7.9 | 6.7 |
| Reasoning | 5,305 | 42.1 | 12.7 | 11.1 |
| Skill marginals | Domain marginals | ||||
| Skill | Mean | Pairings | Domain | Mean | Pairings |
| S1 Attribute Recognition | 0.432 | 5 | D1 Documents/Charts/Slides | 0.041 | 4 |
| S2 Text Reading | 0.001 | 6 | D2 Aerial/Satellite | 0.227 | 6 |
| S3 Counting | 0.250 | 6 | D3 Vehicles/Driving | 0.204 | 6 |
| S4 Spatial Relation | 0.213 | 5 | D4 Indoor | 0.124 | 6 |
| S5 Object Identification | 0.044 | 6 | D5 Outdoor | 0.214 | 6 |
| Pairing | Domain Skill | |
| D2 S6 | Aerial/Satellite Scene Reasoning | 0.780 |
| D6 S4 | People/Surveillance Spatial Relation | 0.580 |
| D3 S3 | Vehicles/Driving Counting | 0.540 |
| D5 S1 | Outdoor Attribute Recognition | 0.540 |
| D2 S1 | Aerial/Satellite Attribute Recognition | 0.520 |
| D5 S4 | Outdoor Spatial Relation | 0.480 |
| Resampling | Complete | percentile interval | (% replicates) | ||
| (%) | (%) | ||||
| Ordinary cluster | – | – | – | ||
| Positive-weight cluster | – | – | – | ||
| Series | Backbone capacity exponent | Interior | ||
| Median | Range | At bound | ||
| InternVL | ||||
| InternVL2.5 | 0.10 | – | 0 | 4.66 |
| InternVL3 | 0.19 | – | 0 | 3.87 |
| InternVL3.5 | 0.44 | – | 0 | 4.93 |
| QwenVL | ||||
| Series | Benchmark | Questions | Gain | interval | Holm |
| InternVL3.5-8B | HR-Bench | ||||
| InternVL3.5-8B | MME-RealWorld | ||||
| InternVL3.5-8B | TreeBench | ||||
| InternVL3.5-8B | V * Bench | ||||
| Qwen3-VL-8B | HR-Bench | ||||
| Qwen3-VL-8B | MME-RealWorld |
| FLOPs of the language backbone | Latency | |||||||
| Scope | ||||||||
| Pooled | 650 | 1.00 | 0.77 | 0.995 | 0.07 | 0.54 | 0.366 | |
| InternVL2.5 | 150 | 1.00 | 0.96 | 1.000 | -0.04 | -0.15 | 0.019 | |
| InternVL3 | 150 | 1.00 | 0.96 | 1.000 | 0.14 | -0.53 | 0.029 | |
| InternVL3.5 | 150 | 1.00 | 0.96 | 1.000 | 0.24 | 1.04 | 0.150 | |
| Qwen2.5-VL | 100 | 1.00 | 0.75 | 0.986 | 0.03 | 0.43 | 0.764 | |
| Rule | Mean | 90th pct. | Worst |
| Direct selection | 0.76 | 2.74 | 19.37 |
| Projection | 2.33 | 7.71 | 34.55 |
| Largest token count | 4.09 | 12.23 | 27.00 |
| Largest backbone | 7.21 | 20.14 | 42.93 |
| Log-center | 7.56 | 16.22 | 38.74 |
| Rule | Mean | Median | 90th pct. | Worst | Exact hit |
| Projection | 1.98 | 0.00 | 6.34 | 11.52 | 51.25% |
| Largest token count | 6.51 | 6.10 | 13.45 | 19.37 | 27.50% |
| Log-center | 8.05 | 7.70 | 16.24 | 20.00 | 1.25% |
| Largest backbone | 15.40 | 14.23 | 28.32 | 41.36 | 0.00% |
| Benchmark | Projection | Largest token count | Log-center | Largest backbone |
| HR-Bench | 0.00 | 5.63 | 11.06 | 14.23 |
| MME-RealWorld | 0.07 | 3.54 | 6.52 | 17.77 |
| TreeBench | 2.49 | 6.09 | 4.10 | 4.85 |
| V * Bench | 0.00 | 7.85 | 9.16 | 27.23 |
| Threshold | Pooled | InternVL | QwenVL |
| pp | 52.50% | 52.08% | 53.13% |
| pp | 58.75% | 58.33% | 59.38% |
| pp | 65.00% | 64.58% | 65.63% |
| pp | 85.00% | 81.25% | 90.63% |
| pp | 96.25% | 93.75% | 100.00% |
| pp | 98.75% | 97.92% | 100.00% |
| Series | Benchmark | Mean | Median | 90th pct. | Worst | Hits |
| InternVL2.5 | HR-Bench | 2.16 | 0.00 | 6.04 | 8.63 | 3/4 |
| InternVL2.5 | MME-RealWorld | 3.44 | 2.02 | 8.03 | 9.74 | 2/4 |
| InternVL2.5 | TreeBench | 3.11 | 2.74 | 6.07 | 6.97 | 1/4 |
| InternVL2.5 | V * Bench | 4.19 | 2.62 | 9.63 | 11.52 | 2/4 |
| InternVL3 | HR-Bench | 0.44 | 0.00 | 1.23 | 1.76 | 3/4 |
| InternVL3 | MME-RealWorld | 2.16 | 0.68 | 5.51 | 7.30 | 2/4 |
| Benchmark | Easy | Scaling Bound | Ceiling Bound |
| HR-Bench | 33.5% | 46.5% | 20.0% |
| MME-RealWorld | 10.3% | 50.3% | 39.4% |
| TreeBench | 10.4% | 48.6% | 41.0% |
| V * Bench | 16.2% | 66.2% | 17.6% |
| Regime share | Dominant axis within Scaling Bound | ||||||
| Family | Benchmark | Easy | Scaling | Ceiling | -dominated | -dominated | Tie |
| InternVL | HR-Bench | 29.0% | 49.6% | 21.4% | 89.7% | 6.7% | 3.6% |
| InternVL | MME-RealWorld | 10.1% | 51.8% | 38.1% | 70.5% | 20.7% | 8.8% |
| InternVL | TreeBench | 8.6% | 49.4% | 42.0% | 88.5% | 7.0% | 4.5% |
| InternVL | V * Bench | 16.8% | 63.9% | 19.4% | 47.8% | 38.5% | 13.7% |
| QwenVL | HR-Bench | 40.3% | 41.8% | 17.9% | 63.0% | 31.6% | 5.5% |