cs.IRSep 21, 2026

MM-VeriRec: Failure-Guided Fusion for Verifiable Agentic Multimodal Recommendation

Authors: Yufeng Wang

Organizations: Independent Research United States

Abstract

Images often carry recommendation constraints that text metadata only hints at. A movie may need to look dark, a product may need a minimal style, and a visually impossible request should be rejected. Agentic multimodal recommenders must reason over text-image evidence, decide when visual evidence is decisive, and abstain when no valid action exists. We introduce MM-VeriRec, a verifiable multimodal recommendation protocol and failure-guided fusion method for hidden visual constraints, image-text mismatch, and impossible-task abstention. MM-VeriRec builds tasks from real movie-poster and product-image datasets, verifies each recommendation with deterministic visual attributes, and converts failures into actionable labels: text-trap following, visual ignorance, and false acceptance. Fusion should not merely concatenate modalities, but should diagnose which modality failed and route to the appropriate repair. Across MM-ML 1M and Amazon Reviews datasets, stronger text and vision embeddings improve retrieval but do not remove these failure modes, whereas failure-guided fusion does. The adaptive attribute gate reads the same tags the verifier checks and its scores are verifier-aligned upper bounds testing whether the taxonomy routes to the correct repair. More informative is transfer under a non-aligned gate: an independently derived leave-one-out CLIP detector still reaches 0.7028 and 0.6111 visual-grounded success, above both a VBPR baseline and plain fusion. The text-versus-visual gap reproduces across two LLM families, and the repair that helps differs by domain. MM-VeriRec is both a benchmark and a practical diagnostic loop for trustworthy agentic multimodal recommendation.

Explore similar work

CardsList
  1. OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

    May 27, 2026Xinchen Zhang, Bowei Liu, Jiale Liu +7Multimodal Large Language ModelsReinforcement Learning with Verifiable Rewards

  2. ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

    Jun 8, 2026Yihua Zhang, Mingfu Liang, Jiyan Yang +11Efficient Multimodal InferenceMultimodal CoT Reasoning

  3. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

    Aug 28, 2026Zongkai Liu, Hui Zhang, Liqiang Niu +7Multimodal IRMultimodal Search Agents