Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
Authors: Jianfei Zhao, Yifan Wang, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan, Yang Luo, Boyuan Pan, +2 more
Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Zhongguancun Academy · Xiaohongshu · Southeast Academy of Information Technology, Beijing Institute of Technology · Zhongguancun Laboratory
Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
Figures & tables
Figure 1: Autoregressive retrieval with ARR. Retrieved items provide context for successive query embeddings; the final embedding ranks the collection.
Figure 2: Overview of ARR training. (a) SFT uses offline feedback trajectories, stepwise contrastive supervision, and a degradation penalty. Only the initial-step loss backpropagates through standalone item embeddings. (b) RL freezes the SFT backbone and item index while training a query-side LoRA adapter. Terminal reciprocal-rank rewards supervise feedback selection through GRPO, complemented by contrastive supervision of the final embedding.
qt→di
qt→dt
qt→di,t
qi→dt
qi→di
qi,t→dt
qi,t→di
qi,t→di,t
Method
VN
CO
F200
WQ
ES
WQ
VN
CO
F200
NS
ON
InS
FIQ
CR
ON
InS
Avg.
R@5
R@5
R@10
R@5
R@5
R@5
R@5
R@5
R@10
R@5
R@5
R@5
R@10
R@5
R@5
R@5
2B models
LamRA-Ret ( Liu et al., 2025 )
30.8
78.8
23.1
82.5
54.3
77.8
31.2
88.5
27.1
28.7
51.1
44.2
28.9
47.7
72.3
60.8
51.6
ELVA ( Liu et al., 2026 )
35.6
80.3
25.0
88.0
56.1
80.5
33.4
90.2
25.9
29.3
52.0
47.4
30.9
50.0
72.8
61.3
53.8
ARR-SFT
32.6
84.4
27.2
93.3
60.4
84.1
32.1
94.3
27.6
32.2
55.2
42.9
33.1
59.7
74.2
62.8
56.0
Table 1: M-BEIR results in the local-pool setting (%). Superscripts t and i denote text and image inputs, respectively. ARR rows report the final state ( N=4 ).
Method
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT N=1
55.23
49.85
48.37
46.90
ARR-SFT N=4
55.63
55.95
56.01
56.01
ARR-RL N=1
55.44
55.57
55.54
55.47
ARR-RL N=2
56.40
56.59
56.58
56.64
ARR-RL N=3
56.40
56.57
56.62
56.63
ARR-RL N=4
56.36
56.68
56.66
56.74
Table 2: Effect of training and inference steps for ARR-2B on M-BEIR (Avg.%).
qt→di
qi→dt
qi,t→di
qdialog→di
qi⊕t→di
Method
Share4V
Urban*
Flickr
Share4V
Urban*
Flickr
CIRCO*
GeneCIS*
VisD*
MT-FIQ*
Avg.
R@1
R@1
R@1
R@1
R@1
R@1
MAP@5
R@1
R@1
R@5
LamRA-Ret ( Liu et al., 2025 )
93.3
95.1
82.8
88.1
94.3
92.7
33.2
18.9
62.8
60.9
72.21
TRACE ( Hao et al., 2026 )
94.9
94.8
84.5
89.1
94.1
94.5
34.8
20.5
65.4
63.2
73.58
ELVA ( Liu et al., 2026 )
96.6
96.1
84.4
92.0
95.5
95.2
34.5
20.2
65.3
61.2
74.10
ARR-SFT (Iter-1)
98.8
98.9
87.3
99.0
99.0
97.7
39.3
20.9
75.5
63.1
77.95
Table 3: Zero-shot retrieval results for 8B models (%). Iter-1 uses the initial query embedding without feedback; unqualified ARR rows use the final state ( N=4 ). An asterisk marks benchmarks with images sourced from COCO or FashionIQ.
Setting
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT
55.63
55.95
56.01
56.01
w/o sg(vd)
55.31
55.62
55.72
55.71
w/o Ldeg
55.69
55.61
55.66
55.68
ARR-RL
56.36
56.68
56.66
56.74
w/o Lret
56.20
56.40
56.39
56.43
w/o reward gate
56.39
56.55
56.52
56.39
Table 4: Training ablations for ARR-2B on M-BEIR (Avg.%). “w/o reward gate” applies final-step supervision to all trajectories.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Query → item
Train / dev / test queries
Pool
VisualNews (VN) ( Liu et al., 2021a )
t → i
99K / 20K / 20K
542K
MSCOCO (CO) ( Lin et al., 2014 )
t → i
100K / 24.8K / 24.8K
5K
Fashion200K (F200) ( Han et al., 2017 )
t → i
15K / 1.7K / 1.7K
201K
WebQA (WQ) ( Chang et al., 2022 )
t → t
16K / 1.7K / 2.4K
544K
EDIS (ES) ( Liu et al., 2023 )
t → i+t
26K / 3.2K / 3.2K
1M
WebQA (WQ)
t → i+t
17K / 1.7K / 2.5K
403K
Appendix
Table 5: M-BEIR tasks and rounded benchmark statistics ( Wei et al., 2024 ; Li et al., 2026b ) . Here, t denotes text and i denotes image; i+t denotes a joint image–text input. Split sizes count queries; “Pool” counts candidates in the local evaluation collection. These are benchmark statistics, not counts after ARR-specific preprocessing.
Benchmark
Query → item
Queries / candidates
ShareGPT4V (Share4V) ( Chen et al., 2024 )
t → i; i → t
1K / 1K (both directions)
Urban-1K (Urban) ( Zhang et al., 2024 )
t → i; i → t
1K / 1K (both directions)
Flickr30K (Flickr) ( Plummer et al., 2015 )
t → i; i → t
t → i: 5K / 1K; i → t: 1K / 5K
CIRCO ( Baldrati et al., 2023 )
i+t → i
800 / 120K
GeneCIS ( Vaze et al., 2023 )
i+t → i
8K / 10–15
Visual Dialog (VisD) ( Das et al., 2017 )
dialogue → i
2K / 2K
Appendix
Table 6: Zero-shot evaluation tasks and metrics. Counts are listed separately when retrieval directions differ. Candidate counts refer to the pool searched by each query; GeneCIS uses query-specific candidate sets.
Setting
Value
Optimizer
AdamW
Weight decay
0.01
Learning rate
5×10−5
Learning-rate schedule
Cosine
Warmup steps
150
Training epochs
2
Appendix
Table 7: SFT training settings shared by ARR-2B and ARR-8B.
Setting
Value
Learning rate
5×10−6
Learning-rate schedule
Cosine
Warmup steps
50
Training epochs
2
Effective global batch size (trajectories)
256
LoRA rank
128
Appendix
Table 8: RL training settings shared by ARR-2B and ARR-8B. All other settings follow those used for SFT.
Setting
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT
λdeg=0
55.69
55.61
55.66
55.68
λdeg=0.1 (default)
55.63
55.95
56.01
56.01
λdeg=1.0
55.60
55.76
55.86
55.89
ARR-RL
λret=0
56.20
56.40
56.39
56.43
Appendix
Table 9: Sensitivity to auxiliary loss weights for ARR-2B on M-BEIR (average recall, %). Bold marks the best result in each column within each training stage.
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.
Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing methods overlook the mismatch bettween candidate-aware retrieval supervision and input-only CoT generation: (1)trajectory-level rewards convey retrieval outcomes without explicitly identifying the input-supported evidence that distinguishes the positive from hard negatives; (2) input-only generation cannot directly assess whether further reasoning improves retrieval, potentially producing redundant CoTs with substantial latency. To bridge this gap, we propose Reason What Matters (ReWAM), a retrieval-grounded framework that aligns candidate-aware supervision with input-only generation. Specifically, we introduce Retrieval-Aware Self-Distillation (RASD), which extracts privileged guidance from input-supported facts and evidence distinguishing the positive from hard negatives. Conditioned on this guidance, an on-policy self-teacher provides token-level feedback to refine credit assignment, directing policy updates toward retrieval-relevant reasoning grounded in the input. We further propose Retrieval-Adaptive Inference (RAI), which learns a retrieval-aware stopping criterion from prefix-level retrieval feedback. It stops redundant reasoning without candidate access and uses speculative decoding to further reduce CoT latency. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. ReWAM thus enables high-quality retrieval through efficient input-only reasoning, making explicit CoT practical for corpus-scale multimodal retrieval. The code will be publicly available.
Mingzhou Jiang, Peixi Wu, Hang Cheng +7
Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Artificial Intelligence and Data Science, USTC · Kuaishou Technology +1
Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR). However, previous works have ignored the grain blindness when adapting the contrastive paradigm into retrieval tasks. Grain blindness refers to the tendency of the model to overlook grain-level information contained in the query, which is crucial for effectively handling complex queries. This stems from contrastive learning treating samples as a binary classification (positive/negative), while ignoring the different information carried by each negative sample. To address this, we argue that negatives should be treated differently according to their similarity to the positive sample, enabling the model to learn distinct grain information from each negative. In this paper, we introduce a simple but effective framework, called ELVA, a novel rule-based RL framework that mitigates grain blindness through ranking-driven MLLMs. 1) Instead of relying on reward models, we extend Reinforcement Learning with Verifiable Rewards (RLVR) to retrieval tasks, allowing the model to explore new ranking behaviors without explicit ranking labels. 2) By utilizing rule-based rewards, our approach jointly optimizes the ranking of negative samples while enlarging the similarity gap between positive and negative. To more precisely measure grain blindness, we further introduce MRBench, a new benchmark specifically designed for multi-grain query scenarios. ELVA achieves state-of-the-art results across standard retrieval benchmarks, and its notable 13.1% improvement on MRBench further demonstrates its effectiveness in alleviating grain blindness.
Yuhan Liu, Pei Fu, Hang Li +8
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University · MiLM Plus, Xiaomi Inc · Zhongguancun Academy, Beijing, China