Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
Figures & tables
Figure 1: RELER workflow. Top: embedding rollouts yield rewards for RLOO+CMP updates. Bottom left: deterministic cosine retrieval at inference. Bottom right: query-only adaptation with a frozen document index and generator.
StackExchange
Coding
Theorem-based
Method
Bio.
Earth.
Econ.
Psy.
Rob.
Stack.
Sus.
Leet.
Pony
AoPS
TheoQ.
TheoT.
Avg.
Original query
BGE-M3
9.5
15.3
11.9
13.2
12.1
10.6
10.2
15.8
15.2
1.3
8.4
4.3
10.66
LambdaLoss
17.2
14.1
12.4
16.2
12.8
14.0
16.6
15.0
3.5
2.7
8.2
4.2
11.40 ± 0.08
InfoNCE
18.0
17.0
14.7
20.2
15.2
14.3
16.7
10.0
8.5
2.2
10.1
4.7
12.63 ± 0.11
RELER
22.9
25.8
19.4
24.7
17.0
17.0
18.4
4.9
6.6
1.7
10.2
5.4
14.49 ± 0.04
Table 1: BRIGHT nDCG@10 ( × 100). Subsets show three-seed means; Avg. is the macro-average ± SD. Pretrained models are evaluated once. Bold/underline mark best/second-best scores per model and query setting; shading marks RELER. TheoQ./TheoT.: TheoremQA questions/theorems.
Reward
Rollout
CMP
StackEx.
Coding
Theorem
Avg.
(a)
Graded + pairwise
Product
Yes
31.25
5.79
16.73
23.38 ± 0.17
(b)
Graded + pairwise
Product
No
28.09
5.73
18.27
21.91 ± 0.70
(c)
Graded nDCG
Product
Yes
29.80
6.21
17.24
22.72 ± 0.18
(d)
Graded nDCG
Product
No
26.22
5.79
18.36
20.85 ± 0.68
(e)
Graded nDCG
Paired
No
22.55
6.68
17.78
18.71 ± 0.53
(f)
Graded nDCG
Query only
No
17.92
8.64
15.85
15.85 ± 0.26
Table 2: BRIGHT ablations (Q3E-0.6B, original queries). Avg. is nDCG@10 ( × 100), mean ± SD over three seeds; StackEx., Coding, and Theorem average 7, 2, and 3 subsets, respectively. Shading marks full RELER. Query-only and document-only sample one embedding side while updating the shared encoder. Full results are in Appendix C.3 .
Figure 2: Training dynamics and sensitivity. (a) CMP reduces gradient noise before fine-tuning (inset: zoom). (b) Pairwise feedback improves BRIGHT despite similar training nDCG reward curves. (Table 2 , rows (a,c) ). (c,d) BRIGHT nDCG@10 ( × 100) is highest at λ=0.5 and stable across ρ=0.60 – 0.90 . Filled points mark defaults; unsmoothed curves and error bars show three-seed means ± SD. Sweep details are in Appendix C.4 .
In-domain
Out-of-domain
Avg.
Method
EM
F1
Hit@10
EM
F1
Hit@10
EM
F1
Hit@10
Q3E-0.6B
33.79
44.17
65.67
31.71
39.40
51.94
32.30
40.76
55.87
InfoNCE
33.63
44.07
64.30
30.19
37.76
46.72
31.18
39.56
51.74
RLOO, graded nDCG
34.66
45.08
67.25
31.38
39.08
51.97
32.32
40.80
56.34
RLOO, answer F1
34.71
45.16
66.77
31.38
39.14
52.61
32.33
40.86
56.66
RLOO, nDCG + F1
34.82
45.14
66.87
32.21
39.48
52.19
32.96
41.10
56.38
Table 3: Fixed-index QA ( × 100), seed 42. In-domain/out-of-domain average 2/5 datasets; Avg. covers all seven. Bold marks column bests. Per-dataset results are in Appendix D.5 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Q3E-0.6B BRIGHT adaptation
Trainable parameters
Full shared encoder
Optimizer / learning rate
AdamW / 5×10−6
Global / per-device batch
128 / 16 (8 devices)
Training steps
113
Schedule / warm-up fraction
Linear / 0.03
Weight decay
0.01
Appendix
Table 4: Training settings for Q3E-0.6B BRIGHT adaptation. All 4B BRIGHT methods use LoRA (rank 16) with a learning rate of 2×10−4 .
Original query
GPT-4 reasoning query
Model
Method
42
3407
2026
42
3407
2026
Q3E-0.6B
Q3E-0.6B
15.10
–
–
20.38
–
–
InfoNCE
22.50
22.56
22.87
26.40
26.59
26.56
LambdaLoss
19.13
19.30
19.67
23.61
23.96
24.34
RELER
23.27
23.57
23.29
27.15
27.19
27.10
Q3E-4B
Q3E-4B
18.70
–
–
22.01
–
–
Appendix
Table 5: Per-seed BRIGHT nDCG@10 ( ×100 ), averaged over 12 subsets. The pretrained encoders are evaluated once. GPT-4 reasoning queries use the released gpt4_reason.query inputs.
StackExchange
Coding
Theorem-based
Method
Bio.
Earth.
Econ.
Psy.
Rob.
Stack.
Sus.
Leet.
Pony
AoPS
TheoQ.
TheoT.
Avg.
General-purpose methods
BM25
18.9
27.2
14.9
12.5
13.6
18.4
15.0
24.4
7.9
6.2
10.4
4.9
14.5
OpenAI-3-Large
23.3
26.7
19.5
27.6
12.8
14.3
20.5
23.6
2.4
8.5
23.5
11.7
17.9
Google-Gecko-1B-768
22.7
34.8
19.6
27.8
15.7
20.1
17.1
29.6
3.6
9.3
23.8
15.9
20.0
GritLM-7B
24.8
32.3
18.9
19.8
17.1
13.6
17.8
29.9
22.0
8.8
25.2
21.2
21.0
Appendix
Table 6: Original-query BRIGHT nDCG@10 ( ×100 ). Literature rows are from Table 2 of Chen et al. (2026) ; MS MARCO fine-tuning controls and pretrained Qwen3-Embedding rows are omitted. ReasonEmbed-Qwen3-4B uses Redapter. Our 4B rows repeat Table 1 ; Avg. reports three-seed mean ± sample SD for trained recipes. Literature and our results use different training and evaluation protocols. Abbreviations follow the main table.
Table 8: Complete BRIGHT subset results for the pairwise-weight and alignment sweeps in Figure 2 (c,d). The first column is λ in the upper block and ρ in the lower block.
SNR
CMP / RLOO
Encoder state
Microbatch
RLOO
CMP
V
t
Vt
Before fine-tuning
Stack Overflow
0.25
5.35
0.00226
1.14
0.00258
Math-theorem
0.27
6.00
0.00182
1.11
0.00203
Biology
0.20
4.85
0.00169
1.13
0.00191
Step 50
Stack Overflow
0.23
4.07
0.00322
1.14
0.00367
Math-theorem
0.22
3.81
0.00321
1.11
0.00358
Appendix
Table 9: Complete-encoder gradient measurements for the combined reward. Each row uses 64 paired rollout draws at one fixed encoder state and microbatch. V is the sample gradient variance and t is mean forward/backward time per draw. Ratios compare CMP to RLOO; SNR uses the bias-corrected signal estimate.
Setting
Value
Training budget / seed
One epoch / 42
Optimizer / learning rate
AdamW / 5×10−6
Global / per-device batch
128 / 16
Schedule / warm-up / weight decay
Cosine / 0.03 / 0.01
Training query limit
128 tokens
InfoNCE temperature / anchor coefficient
0.03 / 0.5
Appendix
Table 10: Fixed-index RAG training and generation settings. The query encoder is fully fine-tuned; the document index and generator remain frozen.
Method
NQ
HotpotQA
PopQA
TriviaQA
2Wiki
MuSiQue
Bamboogle
Queries ( n )
3,610
7,405
14,267
11,313
12,576
2,417
125
Answer EM
Q3E-0.6B
34.99
32.59
42.22
61.00
28.76
7.36
19.20
InfoNCE
37.01
30.26
42.07
60.26
23.81
5.63
19.20
RLOO, graded nDCG
36.01
33.32
41.82
60.54
29.03
6.48
19.02
RLOO, answer F1
35.94
33.49
42.40
61.43
28.98
7.27
16.80
Appendix
Table 11: Per-dataset fixed-index QA results ( ×100 ), seed 42. n is the number of evaluation queries; 2Wiki denotes 2WikiMultihopQA. The same document index, generator, and query counts are used for all recipes.
Objective
Classif.
Clust.
Pair class.
Rerank.
Retrieval
STS
Summ.
Task avg.
Type avg.
InfoNCE
72.84
44.19
82.11
44.80
51.52
75.94
29.49
60.99
57.27
RELER
73.41
45.81
79.76
45.13
50.04
76.81
26.61
61.01
56.80
Appendix
Table 12: Direct embedding training from Qwen3-0.6B on MTEB English v2 ( ×100 ), seed 42. Each type column averages its tasks; Task avg. weights all 41 tasks equally, while Type avg. weights the seven task-type means equally. Classif.: classification; Clust.: clustering; Pair class.: pair classification; Rerank.: reranking; Summ.: summarization. Bold marks the higher score in each column.
Dense retrieval embedding models are a fundamental component of modern retrieval-based AI systems. Most dense retrievers are trained with contrastive objectives, which require labeled positive and negative document pairs that are often costly and difficult to obtain. In this work, we investigate whether the autoregressive next-token prediction objective of a large language model (LLM) can provide supervision for dense retrieval. The intuition is simple: if a document contains information relevant to a query, conditioning on that document should make the target output easier for the LLM to predict. A key challenge is that the next-token prediction loss is computed inside the LLM, while the retriever is a separate embedding model. To address this challenge, we propose DREAM (Dense Retrieval Embeddings via Autoregressive Modeling), which injects retriever-generated query-document similarity scores into selected attention heads of a frozen LLM. During training, these scores determine how much attention each candidate document receives while the LLM predicts the target output. The resulting prediction loss provides gradients for retriever training through the attention mechanism. We evaluate DREAM on retrieval benchmarks BEIR and RTEB using embedding backbones ranging from 0.5B to 3B parameters. DREAM consistently outperforms existing baselines across different model scales. These results demonstrate that DREAM provides a promising approach for training dense retrievers through autoregressive modeling.
Yixuan Tang, Yi Yang
The Hong Kong University of Science and Technology
Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a growing gap between how users express their needs and how retrievers interpret them. In this paper, we present GEM, a generative embedding model that augments retrieval through its own knowledge by explicitly reasoning about user intent and relevance criteria. GEM unifies generation and embedding within a single model: it first reasons over the query, then appends an embedding token to encode the enriched context for retrieval. \zhili{Evaluated on reasoning-intensive and instruction-following retrieval tasks, GEM demonstrates the effectiveness of its reasoning-augmented retrieval, outperforming its non-reasoning variant and matching baselines using substantially larger models.} Furthermore, GEM's generative nature allows test-time compute scaling via prompting to further enhance retrieval performance. Our code is available at: https://anonymous.4open.science/r/GEM.
Dense vector retrieval is the practical backbone of Retrieval- Augmented Generation (RAG), but similarity search can suffer from precision limitations. Conversely, utility-based approaches leveraging LLM re-ranking often achieve superior performance but are computationally prohibitive and prone to noise inherent in perplexity estimation. We propose Utility-Aligned Embeddings (UAE), a framework designed to merge these advantages into a practical, high-performance retrieval method. We formulate retrieval as a distribution matching problem, training a bi-encoder to imitate a utility distribution derived from perplexity reduction using a Utility-Modulated InfoNCE objective. This approach injects graded utility signals directly into the embedding space without requiring test-time LLM inference. On the QASPER benchmark, UAE improves retrieval Recall@1 by 30.59%, MAP by 30.16% and Token F1 by 17.3% over the strong semantic baseline BGE-Base. Crucially, UAE is over 180x faster than the efficient LLM re-ranking methods preserving competitive performance, demonstrating that aligning retrieval with generative utility yields reliable contexts at scale.
Rajinder Sandhu, Di Mu, Cheng Chang +4
Layer 6 AI · Toronto, ON, Canada · Dalhousie University +1