cs.CVOct 5, 2026

Fitting Vision Adapters at Frontier Scales

Authors: Jaehoon Lee, Harry Partridge, Mudith Jayasekara, Charles O'Neill, Max Kirkby, Michael Psenka

Organizations: Base Labs

Abstract

Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

    Aug 4, 2026Yang Yang, Qinyu Zhao, Mouxiang Chen +5Multimodal Large Language ModelsVision Transformer