Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
Figures & tables
Fig. 1: Mechanism overview of direction-dependent responses in multimodal geometric representations. A clean multimodal representation defines an operating point oc in relational geometry; a controlled degradation produces a displacement Δo , and the first-order volume response is governed jointly by the local operating-point gain ∥∇V(oc)∥ , the displacement magnitude D=∥Δo∥ , and the direction factor cosθ , whose product defines the Directional Geometric Response (DGR). Empirically, magnitude alone provides weak response predictability, the direction-free gain–magnitude product remains insufficient, and the direction-aware first-order quantity explains substantially more of the response heterogeneity. DGR is an explanatory quantity because it uses the observed degraded-state displacement.
Fig. 2: Local geometric decomposition of the first-order response. A degradation moves the relational geometry by Δo from the clean operating point oc ; the induced volume change is governed by the projection of Δo onto the local gradient ∇V(oc) , i.e., local gain gV× magnitude D×cosθ .
Fig. 3: Absolute volume response ∣ΔV∣ versus relational displacement magnitude D under single-axis degradation (top: MSR-VTT; bottom: DiDeMo; left: video blur; right: audio noise). Magnitude leaves most of the response variance unexplained, and magnitude-matched samples respond systematically differently.
Candidate
Information retained
OOS R2
Matched- D ranking
Decision
D
displacement magnitude
−0.004 – 0.146
0.509 – 0.523
Reject
S1 context
how compromised the sample state is
— †
0.447 / 0.420 †
Reject
S2 self-sensitivity
probe-induced displacement
— †
0.503 / 0.488 †
Reject
S3 self-consistency
two-view agreement
— †
0.505 / 0.460 †
Reject
∥∇V∥D
operating-point gain × magnitude (no direction)
0.002 – 0.224
0.517 – 0.574
Unstable
∣DGR∣
gain × magnitude × direction (projection onto ∇V(oc) )
0.838 – 0.969
0.864 – 0.963
Supported
TABLE I: Falsification ladder for explanations of geometric response. The table summarizes the falsification sequence, from scalar perturbation summaries to increasingly explicit geometric descriptions. R2 ranges span the four dataset–degradation cells; ranking accuracies are matched- D pairwise accuracies on ∣ΔV∣ at the primary bin width ( ε=0.01 ; chance 0.5 ). † Proxy statistics were computed only for the MSR-VTT cohort under the Phase-3A protocol (video/audio axes); “—” indicates that the metric was not evaluated under the four-cell ∣ΔV∣ protocol for that candidate. The supported quantity is an explanatory first-order term that uses the observed degraded-state displacement, not a deployment-time predictor.
OOS R2 for ∣ΔV∣
Matched- D ranking acc.
Sign accuracy
Cell
D
∥∇V∥D
∣DGR∣
D
∣DGR∣
DGR
MSR-VTT video ( N=878 )
0.146
0.224
0.969
0.523
0.963
0.989
MSR-VTT audio ( N=878 )
−0.004
0.002
0.925
0.515
0.903
0.950
DiDeMo video ( N=980 )
0.136
0.152
0.925
0.509
0.961
0.983
DiDeMo audio ( N=980 )
0.021
0.059
0.838
0.517
0.864
0.909
TABLE II: Main four-cell results. D is the relational displacement magnitude; ∥∇V∥D is the direction-free gain–magnitude product; ∣DGR∣ is the absolute first-order directional response, which uses the degraded-state displacement and is therefore explanatory only. Matched- D ranking is pairwise accuracy within narrow magnitude bins, with chance at 0.5 .
Fig. 4: Geometric explanation ladder: from magnitude to directional response. (a) Out-of-sample R2 for ∣ΔV∣ under three nested descriptions of the same perturbation: magnitude D ; the direction-free gain–magnitude product ∥∇V∥D ; and the directional first-order projection ∣DGR∣ . (b) The same comparison under matched- D pairwise ranking (chance 0.5 ). Direction, rather than the local gain alone, carries the missing structure.
Fig. 5: Pointwise validation of the Directional Geometric Response. Each panel compares the observed volume response ΔV with the first-order directional quantity DGR=∇V(oc)⊤Δo for one dataset–degradation cell. The dashed diagonal denotes the ideal first-order correspondence ΔV=DGR . The strong alignment across samples confirms that the directional first-order term captures the dominant response structure, while the visible deviations from the diagonal, particularly for the audio axes, quantify non-negligible higher-order effects.
Fig. 6: The gain-normalization intervention fails both pre-specified gates. (a) Predictability gate: the bootstrap CI of ΔR2 excludes zero in only one of four cells. (b) Clean-order gate: despite high median Spearman correlation ( ρmed ), the clean top-1 order keep rate falls below the 0.98 threshold on both datasets under the isotropic candidate, and collapses under the anisotropic variant.
Cell
ΔR2 CI
Clean ρmed
Top-1 keep
MSR-VTT video
[0.0014,0.0290]
0.9918 †
0.9715 †
MSR-VTT audio
CI includes 0
0.9918 †
0.9715 †
DiDeMo video
CI includes 0
0.9880 †
0.9296 †
DiDeMo audio
CI includes 0
0.9880 †
0.9296 †
Anisotropic variant
not rescued
—
0.485 / 0.386
TABLE III: Gain-normalization audit. ΔR2 is the bootstrap 95% CI of the out-of-sample R2 change under V=V/(gV+ε) ; clean-order statistics are computed at ε=0.1× median( gV ). The candidate passes neither gate.
Berlin Institute of Health, Charité - Universitätsmedizin Berlin · Department of Mathematics and Computer Science, Freie Universität Berlin · Intelligent Medicine Institute, Fudan University