Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
Authors: Jianfei Zhao, Yifan Wang, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan, Yang Luo, Boyuan Pan, +2 more
Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Zhongguancun Academy · Xiaohongshu · Southeast Academy of Information Technology, Beijing Institute of Technology · Zhongguancun Laboratory
Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
Figures & tables
Figure 1: Autoregressive retrieval with ARR. Retrieved items provide context for successive query embeddings; the final embedding ranks the collection.
Figure 2: Overview of ARR training. (a) SFT uses offline feedback trajectories, stepwise contrastive supervision, and a degradation penalty. Only the initial-step loss backpropagates through standalone item embeddings. (b) RL freezes the SFT backbone and item index while training a query-side LoRA adapter. Terminal reciprocal-rank rewards supervise feedback selection through GRPO, complemented by contrastive supervision of the final embedding.
qt→di
qt→dt
qt→di,t
qi→dt
qi→di
qi,t→dt
qi,t→di
qi,t→di,t
Method
VN
CO
F200
WQ
ES
WQ
VN
CO
F200
NS
ON
InS
FIQ
CR
ON
InS
Avg.
R@5
R@5
R@10
R@5
R@5
R@5
R@5
R@5
R@10
R@5
R@5
R@5
R@10
R@5
R@5
R@5
2B models
LamRA-Ret ( Liu et al., 2025 )
30.8
78.8
23.1
82.5
54.3
77.8
31.2
88.5
27.1
28.7
51.1
44.2
28.9
47.7
72.3
60.8
51.6
ELVA ( Liu et al., 2026 )
35.6
80.3
25.0
88.0
56.1
80.5
33.4
90.2
25.9
29.3
52.0
47.4
30.9
50.0
72.8
61.3
53.8
ARR-SFT
32.6
84.4
27.2
93.3
60.4
84.1
32.1
94.3
27.6
32.2
55.2
42.9
33.1
59.7
74.2
62.8
56.0
Table 1: M-BEIR results in the local-pool setting (%). Superscripts t and i denote text and image inputs, respectively. ARR rows report the final state ( N=4 ).
Method
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT N=1
55.23
49.85
48.37
46.90
ARR-SFT N=4
55.63
55.95
56.01
56.01
ARR-RL N=1
55.44
55.57
55.54
55.47
ARR-RL N=2
56.40
56.59
56.58
56.64
ARR-RL N=3
56.40
56.57
56.62
56.63
ARR-RL N=4
56.36
56.68
56.66
56.74
Table 2: Effect of training and inference steps for ARR-2B on M-BEIR (Avg.%).
qt→di
qi→dt
qi,t→di
qdialog→di
qi⊕t→di
Method
Share4V
Urban*
Flickr
Share4V
Urban*
Flickr
CIRCO*
GeneCIS*
VisD*
MT-FIQ*
Avg.
R@1
R@1
R@1
R@1
R@1
R@1
MAP@5
R@1
R@1
R@5
LamRA-Ret ( Liu et al., 2025 )
93.3
95.1
82.8
88.1
94.3
92.7
33.2
18.9
62.8
60.9
72.21
TRACE ( Hao et al., 2026 )
94.9
94.8
84.5
89.1
94.1
94.5
34.8
20.5
65.4
63.2
73.58
ELVA ( Liu et al., 2026 )
96.6
96.1
84.4
92.0
95.5
95.2
34.5
20.2
65.3
61.2
74.10
ARR-SFT (Iter-1)
98.8
98.9
87.3
99.0
99.0
97.7
39.3
20.9
75.5
63.1
77.95
Table 3: Zero-shot retrieval results for 8B models (%). Iter-1 uses the initial query embedding without feedback; unqualified ARR rows use the final state ( N=4 ). An asterisk marks benchmarks with images sourced from COCO or FashionIQ.
Setting
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT
55.63
55.95
56.01
56.01
w/o sg(vd)
55.31
55.62
55.72
55.71
w/o Ldeg
55.69
55.61
55.66
55.68
ARR-RL
56.36
56.68
56.66
56.74
w/o Lret
56.20
56.40
56.39
56.43
w/o reward gate
56.39
56.55
56.52
56.39
Table 4: Training ablations for ARR-2B on M-BEIR (Avg.%). “w/o reward gate” applies final-step supervision to all trajectories.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Query → item
Train / dev / test queries
Pool
VisualNews (VN) ( Liu et al., 2021a )
t → i
99K / 20K / 20K
542K
MSCOCO (CO) ( Lin et al., 2014 )
t → i
100K / 24.8K / 24.8K
5K
Fashion200K (F200) ( Han et al., 2017 )
t → i
15K / 1.7K / 1.7K
201K
WebQA (WQ) ( Chang et al., 2022 )
t → t
16K / 1.7K / 2.4K
544K
EDIS (ES) ( Liu et al., 2023 )
t → i+t
26K / 3.2K / 3.2K
1M
WebQA (WQ)
t → i+t
17K / 1.7K / 2.5K
403K
Appendix
Table 5: M-BEIR tasks and rounded benchmark statistics ( Wei et al., 2024 ; Li et al., 2026b ) . Here, t denotes text and i denotes image; i+t denotes a joint image–text input. Split sizes count queries; “Pool” counts candidates in the local evaluation collection. These are benchmark statistics, not counts after ARR-specific preprocessing.
Benchmark
Query → item
Queries / candidates
ShareGPT4V (Share4V) ( Chen et al., 2024 )
t → i; i → t
1K / 1K (both directions)
Urban-1K (Urban) ( Zhang et al., 2024 )
t → i; i → t
1K / 1K (both directions)
Flickr30K (Flickr) ( Plummer et al., 2015 )
t → i; i → t
t → i: 5K / 1K; i → t: 1K / 5K
CIRCO ( Baldrati et al., 2023 )
i+t → i
800 / 120K
GeneCIS ( Vaze et al., 2023 )
i+t → i
8K / 10–15
Visual Dialog (VisD) ( Das et al., 2017 )
dialogue → i
2K / 2K
Appendix
Table 6: Zero-shot evaluation tasks and metrics. Counts are listed separately when retrieval directions differ. Candidate counts refer to the pool searched by each query; GeneCIS uses query-specific candidate sets.
Setting
Value
Optimizer
AdamW
Weight decay
0.01
Learning rate
5×10−5
Learning-rate schedule
Cosine
Warmup steps
150
Training epochs
2
Appendix
Table 7: SFT training settings shared by ARR-2B and ARR-8B.
Setting
Value
Learning rate
5×10−6
Learning-rate schedule
Cosine
Warmup steps
50
Training epochs
2
Effective global batch size (trajectories)
256
LoRA rank
128
Appendix
Table 8: RL training settings shared by ARR-2B and ARR-8B. All other settings follow those used for SFT.
Setting
Iter-1
Iter-2
Iter-3
Iter-4
ARR-SFT
λdeg=0
55.69
55.61
55.66
55.68
λdeg=0.1 (default)
55.63
55.95
56.01
56.01
λdeg=1.0
55.60
55.76
55.86
55.89
ARR-RL
λret=0
56.20
56.40
56.39
56.43
Appendix
Table 9: Sensitivity to auxiliary loss weights for ARR-2B on M-BEIR (average recall, %). Bold marks the best result in each column within each training stage.
Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Artificial Intelligence and Data Science, USTC · Kuaishou Technology +1
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University · MiLM Plus, Xiaomi Inc · Zhongguancun Academy, Beijing, China