Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
Figures & tables
Fig. 1: Mechanism overview of direction-dependent responses in multimodal geometric representations. A clean multimodal representation defines an operating point oc in relational geometry; a controlled degradation produces a displacement Δo , and the first-order volume response is governed jointly by the local operating-point gain ∥∇V(oc)∥ , the displacement magnitude D=∥Δo∥ , and the direction factor cosθ , whose product defines the Directional Geometric Response (DGR). Empirically, magnitude alone provides weak response predictability, the direction-free gain–magnitude product remains insufficient, and the direction-aware first-order quantity explains substantially more of the response heterogeneity. DGR is an explanatory quantity because it uses the observed degraded-state displacement.
Fig. 2: Local geometric decomposition of the first-order response. A degradation moves the relational geometry by Δo from the clean operating point oc ; the induced volume change is governed by the projection of Δo onto the local gradient ∇V(oc) , i.e., local gain gV× magnitude D×cosθ .
Fig. 3: Absolute volume response ∣ΔV∣ versus relational displacement magnitude D under single-axis degradation (top: MSR-VTT; bottom: DiDeMo; left: video blur; right: audio noise). Magnitude leaves most of the response variance unexplained, and magnitude-matched samples respond systematically differently.
Candidate
Information retained
OOS R2
Matched- D ranking
Decision
D
displacement magnitude
−0.004 – 0.146
0.509 – 0.523
Reject
S1 context
how compromised the sample state is
— †
0.447 / 0.420 †
Reject
S2 self-sensitivity
probe-induced displacement
— †
0.503 / 0.488 †
Reject
S3 self-consistency
two-view agreement
— †
0.505 / 0.460 †
Reject
∥∇V∥D
operating-point gain × magnitude (no direction)
0.002 – 0.224
0.517 – 0.574
Unstable
∣DGR∣
gain × magnitude × direction (projection onto ∇V(oc) )
0.838 – 0.969
0.864 – 0.963
Supported
TABLE I: Falsification ladder for explanations of geometric response. The table summarizes the falsification sequence, from scalar perturbation summaries to increasingly explicit geometric descriptions. R2 ranges span the four dataset–degradation cells; ranking accuracies are matched- D pairwise accuracies on ∣ΔV∣ at the primary bin width ( ε=0.01 ; chance 0.5 ). † Proxy statistics were computed only for the MSR-VTT cohort under the Phase-3A protocol (video/audio axes); “—” indicates that the metric was not evaluated under the four-cell ∣ΔV∣ protocol for that candidate. The supported quantity is an explanatory first-order term that uses the observed degraded-state displacement, not a deployment-time predictor.
OOS R2 for ∣ΔV∣
Matched- D ranking acc.
Sign accuracy
Cell
D
∥∇V∥D
∣DGR∣
D
∣DGR∣
DGR
MSR-VTT video ( N=878 )
0.146
0.224
0.969
0.523
0.963
0.989
MSR-VTT audio ( N=878 )
−0.004
0.002
0.925
0.515
0.903
0.950
DiDeMo video ( N=980 )
0.136
0.152
0.925
0.509
0.961
0.983
DiDeMo audio ( N=980 )
0.021
0.059
0.838
0.517
0.864
0.909
TABLE II: Main four-cell results. D is the relational displacement magnitude; ∥∇V∥D is the direction-free gain–magnitude product; ∣DGR∣ is the absolute first-order directional response, which uses the degraded-state displacement and is therefore explanatory only. Matched- D ranking is pairwise accuracy within narrow magnitude bins, with chance at 0.5 .
Fig. 4: Geometric explanation ladder: from magnitude to directional response. (a) Out-of-sample R2 for ∣ΔV∣ under three nested descriptions of the same perturbation: magnitude D ; the direction-free gain–magnitude product ∥∇V∥D ; and the directional first-order projection ∣DGR∣ . (b) The same comparison under matched- D pairwise ranking (chance 0.5 ). Direction, rather than the local gain alone, carries the missing structure.
Fig. 5: Pointwise validation of the Directional Geometric Response. Each panel compares the observed volume response ΔV with the first-order directional quantity DGR=∇V(oc)⊤Δo for one dataset–degradation cell. The dashed diagonal denotes the ideal first-order correspondence ΔV=DGR . The strong alignment across samples confirms that the directional first-order term captures the dominant response structure, while the visible deviations from the diagonal, particularly for the audio axes, quantify non-negligible higher-order effects.
Fig. 6: The gain-normalization intervention fails both pre-specified gates. (a) Predictability gate: the bootstrap CI of ΔR2 excludes zero in only one of four cells. (b) Clean-order gate: despite high median Spearman correlation ( ρmed ), the clean top-1 order keep rate falls below the 0.98 threshold on both datasets under the isotropic candidate, and collapses under the anisotropic variant.
Cell
ΔR2 CI
Clean ρmed
Top-1 keep
MSR-VTT video
[0.0014,0.0290]
0.9918 †
0.9715 †
MSR-VTT audio
CI includes 0
0.9918 †
0.9715 †
DiDeMo video
CI includes 0
0.9880 †
0.9296 †
DiDeMo audio
CI includes 0
0.9880 †
0.9296 †
Anisotropic variant
not rescued
—
0.485 / 0.386
TABLE III: Gain-normalization audit. ΔR2 is the bootstrap 95% CI of the out-of-sample R2 change under V=V/(gV+ε) ; clean-order statistics are computed at ε=0.1× median( gV ). The candidate passes neither gate.
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. We introduce credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit. We instantiate this principle with Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events, and CE-GRPO, which selects event boundaries using structural priors and type-normalized entropy, samples complete continuations from shared prefixes, and converts outcome differences into localized advantages. Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by 8.09 and 3.43 points, respectively. Its relative advantage increases with the number of intermediate events, demonstrating the value of representation--optimization co-design for long, dependency-heavy multimodal reasoning.
Jiani Guo, Junjie Wang, Jie Wu +5
1Tsinghua University · 3Zhejiang University · 2Microsoft Research
Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending models from pairwise to higher-order multimodal alignment can introduce optimization and representation challenges. We identify encoder Jacobian conditioning as a key factor in trimodal contrastive learning: poorly conditioned encoders exhibit collapsing or amplified singular-value spectra, leading to exploding Jacobian condition numbers and degraded multimodal alignment. We introduce geometry-preserving encoders (GPEs) by directly conditioning the Jacobian through regularization and demonstrating that simple modifications like LeakyReLU activations and residual paths recover these geometric benefits. Across a synthetic benchmark and four real-world datasets including missing modalities, improving Jacobian conditioning boosts retrieval and linear probe performance across multiple contrastive objectives, whereas expressive objectives yield little benefit in linear probes. More broadly, our results show that multimodal contrastive learning depends not only on objective expressivity, but also on the geometric and optimization properties of the underlying encoders.
Tillmann Rheude, Roland Eils, Benjamin Wild
Berlin Institute of Health, Charité - Universitätsmedizin Berlin · Department of Mathematics and Computer Science, Freie Universität Berlin · Intelligent Medicine Institute, Fudan University
Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of pretrained multimodal contrastive models can serve as a bridge, enabling models to perform multimodal training with unimodal data. However, the key premise of this paradigm remains insufficiently understood: can representations from different modalities be reliably interchanged? The core obstacle lies in the persistent Modality Gap in the shared space. In this work, we revisit the geometric nature of the modality gap. We find that modality representations already share compatible dominant semantic geometry. What truly hinders modality interchangeability is not a simple global shift, but an anisotropic residual structure concentrated along a small number of dominant directions. Based on this finding, we further propose the principle of anisotropic modality gap alignment: effective modality alignment should align with the target-modality distribution while preserving the semantic structure of the source modality. Guided by this principle, we propose an anisotropic geometric correction framework, AnisoAlign, for unpaired modality alignment. This framework leverages the internal geometric prior of the target modality and performs bounded correction on source-modality representations, thereby constructing substitute representations in the target modality. Experiments confirm its benefits in both geometric diagnostics and text-only MLLM training. Overall, this work recasts the modality gap from an empirical observation into a correctable, structured geometric phenomenon and provides a new representation alignment perspective for training multimodal models with unimodal data.