cs.LGOct 2, 2026

Architecture-Dependent Fusion Pathways in MLLMs

Authors: Hebao Zhu, Dongxia Wu

Abstract

Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.

Explore similar work

CardsList
  1. Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation

    Jun 8, 2026Siyuan Liu, Jinyang WuEfficient Multimodal InferenceMultimodal Large Language Models

  2. The Alignment Illusion in Multimodal Large Language Models

    Sep 24, 2026Hong-Han Wang, Yuntao Wang, Hu DingRepresentational Similarity AnalysisMultimodal Large Language Models