cs.LGSep 30, 2026

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

Authors: Ho-min Park, Byungkon Kang

Organizations: Data Science Center, Texas Children’s Hospital Baylor College of Medicine Houston, Texas 77030, USA · Department of Computer Science SUNY Korea Incheon, Republic of Korea

Abstract

This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantitatively, our model outperforms or shows competitive performance on concatenation-based multimodal classification problems. Qualitatively, the proposed interactive mechanism allows the model to progressively refine the visual grounding when paired with complementary modality, thus demonstrating the power of mutual feedback under such settings.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Structured Latent Modeling for Supervised Multimodal Information Decomposition

    Sep 28, 2026Wanting Huang, Sanvesh Srivastava, Weiran WangMultimodal LearningMultimodal Classification

  2. MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning

    Jul 18, 2026Sana Tonekaboni, Viktoria Schuster, Caroline UhlerMultimodal LearningModalities

  3. Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

    May 3, 2026Yuriel Ryan, Hei Man Ip, Adriel Kuek +2Recent Vision-Language ModelsRedundancy