cs.IRSep 4, 2026

Latent-Aligned Reasoning for Multimodal Recommendation

Authors: Jiarui Jin, Anyang Ji

Organizations: Xiaohongshu Inc. · Nanjing University

Abstract

Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding, yet a fundamental challenge persists when applying them to recommendation: as representations propagate through multi-step reasoning, both visual and textual signals progressively attenuate - a phenomenon we term cross-modal dilution. To address this, we propose LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM. In the first stage, learnable latent tokens are interleaved with multi-step chain-of-thought (CoT) reasoning and explicitly aligned with a frozen vision encoder, serving as visual checkpoints that preserve perceptual details throughout the reasoning chain. In the second stage, the latent representations are projected via a bridge MLP and trained with item-to-item contrastive learning; to prevent the reasoning semantics from fading, intermediate features are aligned with the CoT hidden states from the first stage, anchoring the final embeddings to the model's own reasoning output. Experiments on three public benchmarks and one industrial dataset show that LARK achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.

Explore similar work

CardsList
  1. Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

    Apr 8, 2026Mengdan Zhu, Senhao Cheng, Liang ZhaoVLM InterpretabilityLatent Visual Reasoning

  2. Distilling Visual Reasoning into Text Space

    Sep 28, 2026Wenhan Yang, Nilay Naharas, Ali Payani +1CoT DistillationVLM Distillation

  3. ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

    Jun 8, 2026Yihua Zhang, Mingfu Liang, Jiyan Yang +11Efficient Multimodal InferenceMultimodal CoT Reasoning