Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback
Authors: Ho-min Park, Byungkon Kang
Organizations: Data Science Center, Texas Children’s Hospital Baylor College of Medicine Houston, Texas 77030, USA · Department of Computer Science SUNY Korea Incheon, Republic of Korea
This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantitatively, our model outperforms or shows competitive performance on concatenation-based multimodal classification problems. Qualitatively, the proposed interactive mechanism allows the model to progressively refine the visual grounding when paired with complementary modality, thus demonstrating the power of mutual feedback under such settings.
Figures & tables
Figure 1: Diagrams for the main architecture. Left is the overall architecture, and the right is the schematic diagram for F ( G is symmetric).
Dataset
Metric
Concat
Best attn.
MEQ
Δ concat / attn.
Hateful Memes
AUROC
0.6873
0.7251
0.7021
+1.48 / − 2.31pp
VCR
Acc
56.38
62.09
62.95
+6.57 / +0.85pp
VQA v2
Acc
48.97
62.39
63.33
+14.36 / +0.94pp
SNLI-VE
Acc
71.61
74.40
75.15
+3.53 / +0.74pp
CMU-MOSEI
Acc-7
54.79
54.74
54.20
− 0.59 / − 0.54pp
Table 1: Accuracy results across multimodal benchmarks. All numbers are means over three seeds under an identical training recipe. Best attn. is the stronger of parameter-matched cross- and self-attention. VQA v2 and VCR use non-standard evaluation subsets; see Appendix C.1 .
Figure 2: Application of MEQ (trained on SNLI-VE) to camouflaged-object images from COD10k, where ground-truth masks allow direct verification of visual grounding. Border colors give the prediction at each step (green: entailment, red: contradiction, orange: neutral). Best viewed in color.
Figure 3: Plots of P(correct) vs. k .
Configuration
Acc (%)
Δ
Full MEQ
63.11
–
w/o Soft Gating
62.57
− 0.54
w/o Cross-Attention
58.90
− 4.20
w/o Self-Attention
62.33
− 0.78
Single Iteration
62.39
− 0.71
Table 2: Component ablation on VQA v2
Dataset
Metric
Both
Text only
Image only
SNLI-VE
Acc
75.15
70.42
33.85
Hateful Memes
AUROC
0.7080
0.6330
0.6526
CMU-MOSEI
Acc-7
54.70
53.97
–
Table 3: Single-modality controls. Each entry is a model retrained from scratch with one modality zeroed at the encoder output, means over three seeds. Metrics differ per benchmark, so only the verdicts are comparable across rows, not the sizes of the gaps. CMU-MOSEI has three modalities, so ‘text only’ there means audio and video both zeroed and an image-only arm does not apply.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Task
Train
Val
Test
Classes
Hateful Memes ( Kiela et al., 2020 )
Binary classification
8,500
500
1,000
2
VCR ( Zellers et al., 2019 )
4-way multiple choice
212K
26K
25K
4
VQA v2 ( Goyal et al., 2017 )
Open-ended QA
193K
21K
–
3,129
SNLI-VE ( Xie et al., 2019 )
Visual entailment
533K
9.6K
9.6K
3
CMU-MOSEI ( Zadeh et al., 2018 )
Sentiment (7-class)
16K
2K
5K
7
Appendix
Table 4: Dataset statistics.
SNLI-VE
Hateful
VCR
MOSEI
VQA v2
learning rate
1e-4
1e-4
1e-4
2e-5
1e-4
epochs
10
10
15
8
10
batch size
32
32
8
32
32
weight decay
0.01
0.01
0.01
1e-4
0.01
λj
0.5
0.5
0.5
0.5
0.5
λf
0.3
0.3
0.3
0.3
0
Appendix
Table 5: Training hyperparameters. σ(d) denotes a learned scalar passed through a sigmoid, initialised at d=0.5 so that β begins at 0.62 ; the other three benchmarks hold β fixed.
Dataset
Modality-1 encoder
Modality-2 encoder
SNLI-VE
CLIP image
BERT-base
Hateful Memes
CLIP image
BERT-base
VQA v2
CLIP image
BERT-base
VCR
CLIP image + RoI features
BERT-base
CMU-MOSEI
COVAREP (audio Degottex et al. (2014) ) FACET (visual Stöckli et al. (2018) )
BERT-base
Appendix
Table 6: Input encoders per benchmark (all frozen).
Perturbation
Concat.
MEQ
Image Perturbations
Noise ( σ=0.1 )
0.6900
0.7017
Noise ( σ=0.3 )
0.6603
0.6791
Blur ( r=2 )
0.6850
0.6923
Blur ( r=8 )
0.6582
0.6564
Text Perturbations
Appendix
Table 7: Robustness to input corruption on Hateful Memes (dev AUROC).
Figure 4: SNLI-VE dev accuracy read at each iteration k .
Concat
Cross-attn
Self-attn
LMF
MEQ
Accuracy
0.7495
0.7924
0.7849
0.7799
0.7960
Δ vs MEQ
− 4.65pp
− 0.35pp
− 1.11pp
− 1.61pp
–
Appendix
Table 8: COD10k results.
Residual
Metric
Dataset
λf=0.3
λf=0
λf=0.3
λf=0
SNLI-VE
1.0e-2
1.0e-1
75.15
75.06
Hateful Memes
4.2e-2
7.4e-2
70.21
70.37
CMU-MOSEI
2.3e-2
4.4e-2
54.20
54.97
VCR
1.1e-2
6.7e-2
62.95
60.49
VQA v2
9.7e-4
9.4e-2
54.29
63.33
Appendix
Table 9: Effect of the residual penalty.
Configuration
Acc (%)
Δ (pp)
called
gate
cross/self
λf=0.3 , no floor
54.29
–
–
–
λf=0.3 , gate floor 0.1
56.59
+ 2.30
yes
0.1010
0.0027
λf=0.3 , cross floor
53.35
− 0.94
no
0.0894
0.0323
λf=0
63.33
+ 9.04
yes
0.2339
0.0532
Appendix
Table 10: Structural safeguards against the identity collapse, on VQA v2. Means over three seeds. The last two columns are read at the epoch whose weights were kept, averaged over runs: the mean gate value on the image path, and that path’s cross-modal sensitivity divided by its self-modal sensitivity.
Dataset
Configuration
Score
Δ
called
VQA v2
MEQ (coupled states)
63.33
–
Single state, token level
62.98
− 0.35
yes
Feature-sum, one vector
51.94
− 11.39
yes
Hateful Memes
MEQ (coupled states)
0.7094
–
Single state, block-matched
0.7257
+ 1.63
no
Single state, model-matched
0.7262
+ 1.68
no
Appendix
Table 11: State-design comparison. Each arm replaces only the state layout, at matched block parameter count, under the recipe of Table 5 . Means over three seeds; a difference is called when it exceeds twice the sample standard deviation of the paired per-seed difference. Metrics differ per benchmark, so only the verdicts are comparable across blocks, not the sizes of the gaps.
Residual
Metric
Dataset
Seed
ρ at z(K)
Drift
at z(K)
at z(300)
at z(K)
at z(300)
n
CMU-MOSEI
42
0.965
0.38
3.2e-2
8.0e-6
53.9
51.8
1,861
VCR
42
1.011
0.60
1.6e-2
1.0e-3
62.8
55.5
5,000
Hateful Memes
43
1.083
1.05
3.3e-2
1.5e-3
70.7
56.6
500
VQA v2
42
1.340
1.42
7.4e-2
1.5e-3
62.8
27.4
21,029
SNLI-VE
43
1.689
4.51
1.1e-2
4.4e-3
75.0
53.2
9,602
Appendix
Table 12: Fixed-point behavior at the read depth K=10 , measured on the full evaluation set of each benchmark for a single seed. On CMU-MOSEI n is 1,861 of the 1,871 validation clips: the text and the acoustic-visual features are distributed separately and ten clips fail to align.
Model
ms / sample
GFLOPs / sample
peak MB
Concat MLP
1.465
19.27
1015
LMF
1.483
19.27
1003
Cross-attention
1.621
21.20
1004
Self-attention
1.745
21.96
1006
MEQ, K=1
1.627
20.89
1003
MEQ, K=5
2.304
27.64
1003
Appendix
Table 13: Inference cost at batch 32 on one RTX 3090, SNLI-VE architecture. Wall clock is the median of 30 timed passes after five warmups; FLOPs are counted for the whole forward including the frozen encoders; peak memory is measured per arm over an idle baseline. Random weights and random inputs, so this is the cost of a forward pass and not a statement about accuracy.
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.
Wanting Huang, Sanvesh Srivastava, Weiran Wang
Department of Computer Science University of Iowa · Department of Statistics University of Iowa
Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned paired multimodal datasets is often impractical, making end-to-end multimodal training difficult. Second, existing multimodal representations frequently entangle information shared across modalities with modality-specific information, hindering interpretability and control. We introduce MultiLoReFT, an efficient and scalable low-rank representation fine-tuning framework for multimodal learning with pretrained unimodal models. MultiLoReFT extends low-rank adaptation to the multimodal setting and learns interpretable projection subspaces that decouple shared and modality-specific information. Across simulated and real-world benchmarks, it produces representations that support multimodal prediction while explicitly revealing how shared and modality-specific information is distributed across modalities.
Sana Tonekaboni, Viktoria Schuster, Caroline Uhler
Massachusetts Institute of Technology, Cambridge, MA, USA · Eric and Wendy Schmidt Center, Broad Institute of MIT and Harvard, Cambridge, MA, USA · Vector Institute, Toronto, Canada +1
Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting the shared information between modalities to compensate for the impaired one. To this end, we analyze multimodal interactions -- redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information provided by the modalities -- to determine their impacts on model reliability. Specifically, amplifying redundant interactions would increase this exploitable shared information to resolve these issues; yet, modern instruction datasets often eliminate redundancies to prioritize visual grounding. We bridge this gap through a self-captioning workflow featuring a \textsc{Multimodal Interaction Gate}: a mechanism to convert unique interactions into redundant interactions. Our findings suggest that increasing redundancy can reduce visual induced errors by 38.3% and improve consistency by 16.8%.
Yuriel Ryan, Hei Man Ip, Adriel Kuek +2
Singapore University of Technology and Design · DSO National Laboratories · Massachusetts Institute of Technology