Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.
Figures & tables
Figure 1: Primary training setup. Only the 50M parameter MLP adapter is trained; the 743B pure language model remains frozen at its pretraining state.
Figure 2: Visual capability for growing language model sizes , with the same encoder and ≈ 50M-parameter adapter method throughout. MMMU-Pro rises steeply to 14B and plateaus by 32B, with MoE about 12 points above the dense plateau; blind floors stay flat within each family. BLINK single-image saturates at every scale. The break ( ∥ ) compresses the axis between 32B and the MoE.
GLM-5.2
Nemotron-3-Ultra
n
Chance
Blind
Vision
Blind
Vision
MMMU (standard, val)
805
26.4
63.6
71.4
60.9
67.3
MMMU-Pro (standard)
1730
12.2
42.3
56.8
39.7
56.9
BLINK single-image (6 tasks)
793
41.9
41.7
50.9
41.2
56.6
BLINK multi-image (8 tasks)
1108
37.1
37.8
39.6
38.1
46.9
Table 1: Benchmark results Accuracy (%); chance = random-guess rate. Artifacts served and scored with the official generative evaluator (decode configuration of Section 2 ); blind = same served system, no image. 95% binomial CIs span ±2.3 – 3.4 ; parse-fallback rates are in Appendix A .
reader
MMMU-Pro
MMMU
BLINK6
4B
31.9
49.1
44.4
8B
41.3
63.0
47.0
14B
45.3
62.6
49.8
32B
43.4
64.2
52.6
Table 2: Qwen3 models after full recipe (SFT + RL), accuracy %. Reasoning mode on ( /think ), greedy decode, 6,144-token cap, official parser; parse fallback ≤2% per cell. Pre-RL baselines under the identical protocol are in Appendix A .
Figure 3: Specific vision capabilities for GLM-5.2 adapted to vision. All 14 BLINK tasks for GLM-5.2 system, ordered by margin over chance. Squares indicate single-image tasks, and diamonds multi-image tasks. Decode configuration as in Section 2 , identical for vision and blind arms.
LIVR holdout
BLINK: trained task
BLINK: other 13 tasks
Capability
base
adapt
Δ
base
adapt
Δ
base
adapt
Δ
Visual corr. †
19.2
64.0
+44.8
26.2
56.4
+30.2
41.8
45.5
+3.6
Visual similarity
60.8
89.2
+28.4
57.8
84.4
+26.7
39.1
45.0
+5.9
Art style
51.2
81.6
+30.4
53.8
76.9
+23.1
39.5
42.1
+2.5
Semantic corr.
26.0
45.6
+19.6
23.7
40.3
+16.5
41.7
49.2
+7.4
Object localization
56.4
89.2
+32.8
47.5
63.1
+15.6
39.3
43.4
+4.1
Table 3: Task-specific projector training in GLM-5.2. Each row trains only the 49.5M-parameter projector on 1,000 LIVR-constructed examples, keeping GLM-5.2 and MoonViT frozen. Training uses 10 epochs, a learning rate of 10−4 , and one seed. We report final-checkpoint accuracy (%) on the LIVR holdout, the corresponding BLINK task, and the remaining 13 BLINK tasks pooled together ( n=1,778 ). Base and adapt denote scores before and after training. BLINK evaluation uses teacher-forced choice-likelihood scoring on images composited into a single input image. † : Evaluation on held-out HPatches and FunKPoint splits; see LIVR ( Li et al., 2025 ) .
Figure 4: The logit lens reads “grass” everywhere early whereas the Jacobian lens sees the dog from the first layers. Top rows: per-patch winner between “dog” and “grass” overlaid on the image at five depths (GLM, 743B) for each image token. Bottom: fraction of dog patches read as “dog” per layer (grass patches stay “grass” at 0.7–0.9 under the Jacobian lens throughout). Grass labels use a color heuristic; the mask-based control is Appendix Figure 10 .
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
global batch
128 conversations
optimizer
AdamW, β=(0.9,0.95) , weight decay 0.1
learning rate
3×10−4 , constant after 50-step linear warmup
gradient clip
1.0
precision / max sequence length
bf16 / 5,120 tokens (padded)
steps
2,048 (2 epochs)
shuffle
single fixed-seed shuffle, shared across runs
Appendix
Table 4: Supervised fine-tuning constants (identical at every scale).
prompts / group size per update
64 / 8 (512 responses)
sampling
temperature 1.0, top- p 0.95
response cap
6,144 tokens
clip / KL
εhigh=0.28 / β=10−3
learning rate
1×10−5 , projector only
updates / checkpointing
150 / every update
correctness judge
LLM judge against reference answers; judge failure aborts the update
Appendix
Table 5: Reinforcement-learning constants (identical at every scale).
reader
MMMU-Pro
MMMU
BLINK6
4B
26.5 (19)
47.6 (11)
43.4 (0)
8B
35.6 (26)
51.8 (18)
49.6 (8)
14B
44.5 (81)
60.5 (79)
49.4 (58)
32B
44.4 (62)
64.1 (72)
49.9 (80)
Appendix
Table 6: Pre-RL baselines: the SFT-only adapter, evaluated with the reader’s reasoning mode on ( /think ) under the protocol of Table 2 . Accuracy % (in parentheses: % of responses with a non-empty reasoning block).
GLM-5.2
Nemotron-3-Ultra
Blind
Vision
Blind
Vision
MMMU (standard, val)
≤3
10.1
≤3
10.8
MMMU-Pro (standard)
≤3
≤3
5.0
≤3
BLINK single-image (6 tasks)
≤3
≤3
≤3
16.8
BLINK multi-image (8 tasks)
≤3
44.5
≤3
30.0
Appendix
Table 7: Parse-fallback rates (%) for the arms of Table 1 : the share of responses with no parseable answer, scored by seeded random choice.
Figure 5: Training across reader scales. (a) SFT training loss (final losses 4B 0.230 , 8B 0.178 , 14B 0.163 , 32B 0.180 ; comparable within the shared-tokenizer Qwen3 family only). Every scale shows a grokking-like delayed transition: an extended plateau followed by an abrupt drop (Appendix D ). (b) RL reward on rendered math (Section 3.3 ), same recipe at every scale; all four runs complete 150 updates. Initial reward tracks reader scale (4B starts at zero, 32B at 0.72 ); the spread narrows from 0.72 at update 0 to 0.14 at update 150 (32B ends highest, ≈ 0.80).
Figure 6: MoE readers’ training curves (the runs that produced the released checkpoints). (a) SFT loss: Nemotron-3-Ultra descends without a visible plateau. GLM-5.2’s SFT log is too short to plot at this scale; its affine-initialized loss starts at 0.86 , already below the Qwen3 plateau levels, and descends smoothly. The plateau-then-collapse transition is a Qwen3-family observation. (b) RL reward on rendered math; dotted lines mark the released checkpoints (update 154 for GLM-5.2, 118 for Nemotron-3-Ultra). Loss and reward values are not comparable across readers or to the Qwen3 figures: the readers use different tokenizers, trainers, reward compositions, and judges. Only the curve shapes are comparable.
GLM-5.2
Qwen3-32B
Capability
base
adapt
Δ
base
adapt
Δ
Visual corr.
26.2
56.4
+30.2
26.2
53.5
+27.3
Visual similarity
57.8
84.4
+26.7
59.3
79.3
+20.0
Art style
53.8
76.9
+23.1
47.0
70.9
+23.9
Semantic corr.
23.7
40.3
+16.5
26.6
28.1
+1.4
Object localization
47.5
63.1
+15.6
45.9
59.0
+13.1
Appendix
Table 8: Capability fine-tuning: GLM-5.2 versus Qwen3-32B. Full version of Table 3 , to demonstrated the comparative utility of a larger language model.
benchmark
n
clean
no image
Δ (pp)
paired 95% CI
MMMU-Pro
1,730
62.89
41.85
+21.04
[18.50, 23.64]
Appendix
Table 9: GLM-5.3 accuracy (%) with clean vision and a matched no-image control. Intervals for clean minus no-image are 95% subject-stratified paired-bootstrap intervals based on 10,000 replicates.
Figure 7: Corruption families and levels. One representative MMMU-Pro image transformed by the evaluation code. Noise σ=0.10 and brightness 1.50× occur in both the moderate grid and severe screen and are shown once under the moderate grid; the right columns are the additional screened levels. Orange frames mark the two levels carried into holdout confirmation. Saturation was not part of the severe screen.
condition
accuracy (%)
Δ vs. clean (pp)
paired 95% CI
clean
62.89
—
—
noise σ=0.02
64.57
+1.68
[ −0.29 , 3.64]
noise σ=0.05
64.74
+1.85
[ −0.12 , 3.82]
noise σ=0.10
63.24
+0.35
[ −1.50 , 2.20]
brightness 1.10×
62.60
−0.29
[ −2.25 , 1.68]
brightness 1.25×
62.72
−0.17
[ −2.02 , 1.73]
Appendix
Table 10: GLM-5.3 on the full MMMU-Pro moderate corruption grid. Changes and pointwise paired 95% intervals are relative to the clean arm.
condition
accuracy (%)
Δ vs. clean (pp)
paired 95% CI
clean
61.11
—
—
noise σ=0.10
61.11
+0.00
[ −5.00 , 5.00]
noise σ=0.25
55.56
−5.56
[ −10.56 , −0.56 ]
noise σ=0.50
51.67
−9.44
[ −15.56 , −3.33 ]
noise σ=1.00 †
41.67
−19.44
[ −26.11 , −12.78 ]
brightness 1.5×
60.00
−1.11
[ −5.56 , 3.33]
Appendix
Table 11: Exploratory severe-corruption screen on the balanced 180-row MMMU-Pro panel. Accuracy and changes from the same-run clean arm count invalid generations as incorrect. Intervals are 10,000-replicate pointwise subject-stratified paired 95% bootstrap intervals and are not adjusted for multiple comparisons. † marks the levels carried into disjoint confirmation.
condition
accuracy (%)
Δ vs. clean (pp)
paired interval
clean
62.71
—
—
noise σ=1.0
41.29
−21.42
[ −24.52 , −18.39 ]
brightness 12×
45.23
−17.48
[ −20.32 , −14.58 ]
Appendix
Table 12: Severe-corruption confirmation on the 1,550 MMMU-Pro rows disjoint from the severity-selection panel. Intervals are simultaneous familywise 95% intervals.
Figure 8: The mask × placement grid. Rows: query position; columns: key position; dark cells: attention permitted. Text tokens are causal in all variants; the orange square marks the mutually visible image block. (c) matches the reference convention.
Figure 9: A shared plateau, entered by every cell; the cells differ in when they transition. Training loss (window-25 smoothed) for the 2×2 grid at the 14B reader, one run per cell. All four runs sit on a plateau between roughly 0.95 and 0.80 ; transitions are abrupt ( ∼ 40–50 steps); the reference convention transitions first.
cell (14B reader)
MMMU
MMMU-Pro
BLINK6
causal, in-message (reference)
55.3 (0.0)
33.2 (0.0)
49.7 (0.0)
bidirectional, in-message
53.4 (0.0)
33.6 (0.0)
48.6 (0.0)
bidirectional, prefix
51.8 (0.0)
32.8 (0.2)
48.8 (0.0)
causal, prefix
54.0 (0.8)
32.5 (0.2)
48.3 (0.0)
Appendix
Table 13: Eval grid for the 2×2 ablation, final SFT checkpoints (no RL), reasoning mode off, generative protocol, accuracy % (parse-fallback % in parentheses).
Figure 10: Bias-controlled reading: the adapter’s code is readable in gradient coordinates from the first layers; in embedding coordinates only late. Dog-vs-cat discrimination margin on masked patches (released 743B stitch), twelve COCO images, median. The Jacobian lens (solid) discriminates at every depth; the logit lens (dotted) tracks the word prior and is unreliable before ≈ layer 50.
Figure 11: Image and text tokens enter the frozen reader nearly orthogonal and merge gradually with depth. Top: cosine similarity between the mean image token and the mean text token per layer (unit-normalized tokens; 18 COCO images, released 743B stitch; x=−1 is the spliced input); half of the total rise is reached by layer 12. Bottom: joint per-layer PCA of image-token and text-token hidden states (31 images, unit-normalized residuals; same model and splice, but a separate image set from the curve, so the per-panel centroid cosines differ slightly from the curve at the same layers). The populations are disjoint at the input and begin to overlap by layer 77 while remaining largely separable.