Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.
Figures & tables
Figure 1 : Three MLLM families distinguished by the representation supplied to the multimodal Transformer. Encoder-full MLLMs receive continuous features from a dedicated perceptual encoder. Discrete-token MLLMs receive discrete perceptual codes, while encoder-free MLLMs receive lightly projected patches or frames. We ask whether the latter two families form a Virtual Encoder inside the Transformer.
Figure 2 : When a multimodal transformer receives perceptual tokens that have not already been organized by a continuous modality encoder, it must construct a useful representation internally. In Gemma 4 12B, an early band of layers, which we call the Virtual Encoder, performs this encoding for both modalities up to a mid-depth readout boundary. There the vision representation aligns with the model’s language subspace, whereas the audio representation, though formed just as early, is not maintained downstream.
Figure 3 : Layerwise ∣C∣ -way linear-probe accuracy. Encoder-free and discrete-token MLLMs develop concept decodability inside the Transformer, whereas encoder-full controls start with highly decodable representations at layer 0.
Figure 4 : Image-level CKA. The horizontal axis is the target model’s hidden state; the vertical axis is DINOv2 hidden state for vision and AST hidden state for audio. The first three columns show vision against DINOv2; the last column shows audio against AST. Dashed cyan boxes mark the first 20% of MLLM states against DINOv2 layer 0 to 9 or the first 40% of AST states.
Figure 5 : Causal readout of modality information measured by single-layer corruption of the modality-token residual across depth. The top row shows classification accuracy, and the bottom row shows the mean absolute shift in the yes – no logit margin. For the image modality (left), accuracy collapses to the majority-class floor when the modality-token residual is corrupted at any layer up to the readout boundary (near layer 23; dashed vertical line), but fully recovers when corruption is applied beyond this boundary. The three image curves, corresponding to σ∈{250,500,1000} , all identify the same transition. The audio modality (right; σ=500 ) shows a similar boundary, although the effect of corruption is weaker.
Figure 6 : Converging evidence in Chameleon. (a) Concept decodability peaks at hidden state layer 8. (b) Image-level CKA against DINOv2 is high around the same early-to-middle transition. (c) Corrupting the VQ-image-token residual collapses the alive/not-alive task to its majority floor through layer 7 and recovers across layers 8 to 11 (shaded). (d) Image-text subspace overlap follows a different trajectory from Gemma 4 12B, peaking earlier and decreasing across the causal readout window.
Figure 7 : Modality subspace allocation in the shared residual stream. (a) Top- 20 overlap in Gemma 4 12B; the horizontal dashed line is the mean overlap of independent random 20 -dimensional subspaces in the original d -dimensional space. Image-text subspace alignment is maximal near the readout transition (vertical dashed line), whereas the audio pairs remain much weaker but above chance. (b) Image-audio overlap for Gemma 4 12B and the encoder-full Qwen3-Omni control. (c) The image-text peak and layer-mean audio overlaps for k∈{5,10,20,40} , with the corresponding random baseline.