Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 1022 FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Figures & tables
Figure 1: Compute-optimal loss frontiers for encoder-free and encoder-based models. Left: The text loss frontiers of the two architectures nearly overlap. Right: Encoder-free models have higher multimodal loss over the measured range, but their loss decreases faster with compute.
Figure 2: Architectural comparison of encoder-based and encoder-free MLLMs.
Figure 3: Estimating compute-optimal allocation with IsoFLOP profiles for the text objective (left) and the multimodal objective (right).
Attention
Text
Multimodal
Bidirectional
0.427
0.570
Causal
0.436
0.557
Table 1: Model allocation exponent a of encoder-free models under bidirectional and causal attention over visual tokens.
Figure 4: Efficiency gain analysis on the compute-optimal frontier for the text (top) and multimodal (bottom) objectives. (A) fitted loss–compute scaling laws, (B) compute efficiency gain ( EGC ), (C) model efficiency gain ( EGM ).
Figure 5: Efficiency gain analysis under overtraining for the text (top) and multimodal (bottom) objectives, with shades denoting overtraining factors k=1 – 5 from darkest to lightest. (A) loss–compute scaling laws, (B) compute efficiency gain ( EGC ), (C) model efficiency gain ( EGM ).
Figure 6: Compute efficiency gain on each multimodal topic. Curves from dark to light correspond to overtraining factors k=1 – 5 . Solid segments span the fitted budgets, and dashed segments extrapolate the fitted laws.
Figure 7: Visual learning across training tokens. (A, B) Multimodal validation loss during training. (C) Attention mass on visual tokens per decoder layer for the 8B models, before and after the encoder-free loss drop (shaded in A, B).
Figure 8: Comparison between causal and bidirectional attention on visual tokens. Values above 1.0 favor causal.
Figure 9: Layerwise representation evolution. Cosine similarity to the input representation at layer 0 across decoder layers for visual tokens (left) and text tokens (right).
Figure 10: Expert load imbalance over training. MaxVio is computed jointly over visual and text tokens (left), over visual tokens only (middle), and over text tokens only (right). Higher values indicate greater imbalance.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Objective
System
Estimator
Loss γ
Model a
Data b
Text
Encoder-free
IsoFLOP
0.0973
0.427
0.573
Text
Encoder-free
Envelope
0.0905
0.430
0.570
Text
Encoder-based
IsoFLOP
0.0979
0.422
0.578
Text
Encoder-based
Envelope
0.0915
0.435
0.565
Multimodal
Encoder-free
IsoFLOP
0.3778
0.570
0.430
Multimodal
Encoder-free
Envelope
0.3668
0.583
0.417
Appendix
Table 2: Scaling exponents estimated by IsoFLOP and validation curve envelope fitting.
Figure 11: Validation curve envelopes grouped by objective: (a) text and (b) multimodal. Within each group, the top row shows unsmoothed encoder-free and encoder-based validation trajectories. The bottom row shows representative estimates of compute-optimal Mopt and Dopt derived from the trajectories and binned in log compute, with fitted allocation laws.
Figure 12: Extrapolation error analysis on text loss. Laws fitted through 8×1019 FLOPs are evaluated against held-out IsoFLOP optima at 4×1020 FLOPs, a 5× extrapolation. Dashed segments indicate extrapolation.
Figure 13: Extrapolation error analysis on multimodal loss. Laws fitted through 4×1020 FLOPs are evaluated against held-out IsoFLOP optima at 1×1021 FLOPs, a 2.5× extrapolation. Dashed segments indicate extrapolation.
Figure 14: Sensitivity to the irreducible loss of each objective. Each row fixes Eo/E^o and jointly refits the encoder-free and encoder-based loss–compute scaling laws with a common Eo . (A) Fitted loss–compute exponent γ . (B) Implied crossover compute under compute-optimal allocation ( k=1 ) and 5× overtraining ( k=5 ). The dashed line marks the largest fitted budget. (C) Joint SSE on raw loss, normalized by its minimum. The red band marks ±10% around the fitted floor.
Emm/E^mm
Emm
γfree−γbased
Compute-optimal
5× overtraining
[0.90,1.10]
[0.363,0.444]
[0.077,0.078]
[4.8,9.0]×1021
[0.9,1.9]×1022
Appendix
Table 3: Multimodal sensitivity to perturbations of the fitted shared irreducible loss.
Figure 15: Conditional bootstrap uncertainty of the multimodal scaling fits and extrapolated crossover. Left: joint distribution of the encoder-free and encoder-based loss–compute exponents γ . Middle: joint distribution of their model allocation exponents a . Right: distributions of the extrapolated crossover compute under compute-optimal allocation and 5× overtraining. Horizontal bars span the 10 th to 90 th percentiles, and dashed lines mark point estimates.
Encoder-free
Encoder-based
Exponent
Estimate
80% interval
Estimate
80% interval
Loss γ
0.3778
[0.3686,0.3873]
0.2998
[0.2882,0.3112]
Model a
0.570
[0.546,0.595]
0.464
[0.458,0.472]
Data b
0.430
[0.405,0.454]
0.536
[0.528,0.543]
Appendix
Table 4: Conditional bootstrap uncertainty of the multimodal scaling exponents. Estimates are the fitted IsoFLOP exponents. Intervals are central 80% bootstrap percentile intervals ( 10 th to 90 th percentiles).
Accounting
System
Model a
Data b
Loss γ
Decoder only
Encoder-free
0.570
0.430
0.378
Decoder only
Encoder-based
0.464
0.536
0.300
Full
Encoder-free
0.570
0.430
0.378
Full
Encoder-based
0.273
0.727
0.362
Appendix
Table 5: Multimodal scaling exponents under decoder and full compute accounting.
Figure 16: Layerwise similarity to the decoder input across three scales. Encoder-free visual states diverge earlier at every scale, while text trajectories remain closely matched.
Figure 17: Compute-optimal allocation of encoder-free models under bidirectional and causal attention over visual tokens, for the text objective (left) and the multimodal objective (right). The top row shows IsoFLOP profiles, and the bottom row shows the fitted Mopt and Dopt laws. Diamonds mark fitted vertices and dashed segments indicate extrapolation.
Figure 18: IsoFLOP profiles and fitted compute-optimal allocation laws for three pure text sets. The encoder-free and encoder-based exponents remain close on every topic.
Figure 19: IsoFLOP profiles by topic and fitted compute-optimal allocation laws for five multimodal topics. On most topics, encoder-free models favor larger model scale.
Figure 20: Downstream benchmark performance against multimodal training compute. The y-axis is the unweighted average 3-shot score over the 11 benchmarks.
6.1B-A425M
8B-A577M
13B-A922M
16B-A1B
22B-A1.6B
33B-A2.2B
Benchmark
Based
Free
Based
Free
Based
Free
Based
Free
Based
Free
Based
Free
Perception
CV-Bench
42.5
46.9
48.2
43.5
46.1
50.4
49.5
44.1
51.5
48.7
56.2
53.8
POPE
66.8
60.3
64.9
61.5
75.2
57.0
69.9
53.0
67.2
66.7
71.6
61.8
MME
62.1
58.0
64.3
53.1
60.9
57.1
67.7
57.9
67.9
58.7
60.7
58.8
Document
Appendix
Table 6: 3-shot benchmark scores at about 100B training tokens. “Based” and “Free” denote encoder-based and encoder-free models, respectively.
Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain systematically uncharacterized. To address this gap, we investigate the optimal model size and token count for training a transformer-based vision-language model under a fixed computational budget. We demonstrate that minimal objective loss adheres to a predictable compute law, whereas compute-optimal model sizes and token counts scale as power laws. Notably, language and multimodal objectives manifest distinct scaling behaviors. The language allocation law is largely invariant to the composition of the data, indicating stable language learning regardless of the multimodal data ratio. Conversely, the multimodal allocation law is highly sensitive to this composition. Specifically, text-heavy mixtures become compute-efficient only at larger model scales, shifting the optimal resource allocation toward greater model capacity. Additionally, by modeling the influence of data composition on compute laws and allocation exponents, we derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture. Downstream evaluations further reveal that native multimodal pre-training induces positive cross-modal transfer, thereby enhancing pure-text spatial reasoning and enabling robust multimodal in-context learning. In summary, this empirical research establishes the essential groundwork for predictably scaling multimodal foundation models.
Haoyuan Wu, Aoqi Wu, Hai Wang +3
1The Chinese University of Hong Kong · 2LLM Department, Tencent
Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.
Jaehoon Lee, Harry Partridge, Mudith Jayasekara +3
Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.
Yilin Yang, Jun-Tao Tang, Kengyi Wang +3
School of Artificial Intelligence, Shanghai Jiao Tong University · Nanjing University · Fudan University +1