cs.IROct 8, 2026

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Authors: Jianfei Zhao, Yifan Wang, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan, Yang Luo, Boyuan Pan, +2 more

Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Zhongguancun Academy · Xiaohongshu · Southeast Academy of Information Technology, Beijing Institute of Technology · Zhongguancun Laboratory

Abstract

Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 6, 2026cs.CV

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.
Sep 14, 2026cs.AI

Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing methods overlook the mismatch bettween candidate-aware retrieval supervision and input-only CoT generation: (1)trajectory-level rewards convey retrieval outcomes without explicitly identifying the input-supported evidence that distinguishes the positive from hard negatives; (2) input-only generation cannot directly assess whether further reasoning improves retrieval, potentially producing redundant CoTs with substantial latency. To bridge this gap, we propose Reason What Matters (ReWAM), a retrieval-grounded framework that aligns candidate-aware supervision with input-only generation. Specifically, we introduce Retrieval-Aware Self-Distillation (RASD), which extracts privileged guidance from input-supported facts and evidence distinguishing the positive from hard negatives. Conditioned on this guidance, an on-policy self-teacher provides token-level feedback to refine credit assignment, directing policy updates toward retrieval-relevant reasoning grounded in the input. We further propose Retrieval-Adaptive Inference (RAI), which learns a retrieval-aware stopping criterion from prefix-level retrieval feedback. It stops redundant reasoning without candidate access and uses speculative decoding to further reduce CoT latency. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. ReWAM thus enables high-quality retrieval through efficient input-only reasoning, making explicit CoT practical for corpus-scale multimodal retrieval. The code will be publicly available.
Jun 18, 2026cs.IR

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR). However, previous works have ignored the grain blindness when adapting the contrastive paradigm into retrieval tasks. Grain blindness refers to the tendency of the model to overlook grain-level information contained in the query, which is crucial for effectively handling complex queries. This stems from contrastive learning treating samples as a binary classification (positive/negative), while ignoring the different information carried by each negative sample. To address this, we argue that negatives should be treated differently according to their similarity to the positive sample, enabling the model to learn distinct grain information from each negative. In this paper, we introduce a simple but effective framework, called ELVA, a novel rule-based RL framework that mitigates grain blindness through ranking-driven MLLMs. 1) Instead of relying on reward models, we extend Reinforcement Learning with Verifiable Rewards (RLVR) to retrieval tasks, allowing the model to explore new ranking behaviors without explicit ranking labels. 2) By utilizing rule-based rewards, our approach jointly optimizes the ranking of negative samples while enlarging the similarity gap between positive and negative. To more precisely measure grain blindness, we further introduce MRBench, a new benchmark specifically designed for multi-grain query scenarios. ELVA achieves state-of-the-art results across standard retrieval benchmarks, and its notable 13.1% improvement on MRBench further demonstrates its effectiveness in alleviating grain blindness.