Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
Figures & tables
Figure 1: RELER workflow. Top: embedding rollouts yield rewards for RLOO+CMP updates. Bottom left: deterministic cosine retrieval at inference. Bottom right: query-only adaptation with a frozen document index and generator.
StackExchange
Coding
Theorem-based
Method
Bio.
Earth.
Econ.
Psy.
Rob.
Stack.
Sus.
Leet.
Pony
AoPS
TheoQ.
TheoT.
Avg.
Original query
BGE-M3
9.5
15.3
11.9
13.2
12.1
10.6
10.2
15.8
15.2
1.3
8.4
4.3
10.66
LambdaLoss
17.2
14.1
12.4
16.2
12.8
14.0
16.6
15.0
3.5
2.7
8.2
4.2
11.40 ± 0.08
InfoNCE
18.0
17.0
14.7
20.2
15.2
14.3
16.7
10.0
8.5
2.2
10.1
4.7
12.63 ± 0.11
RELER
22.9
25.8
19.4
24.7
17.0
17.0
18.4
4.9
6.6
1.7
10.2
5.4
14.49 ± 0.04
Table 1: BRIGHT nDCG@10 ( × 100). Subsets show three-seed means; Avg. is the macro-average ± SD. Pretrained models are evaluated once. Bold/underline mark best/second-best scores per model and query setting; shading marks RELER. TheoQ./TheoT.: TheoremQA questions/theorems.
Reward
Rollout
CMP
StackEx.
Coding
Theorem
Avg.
(a)
Graded + pairwise
Product
Yes
31.25
5.79
16.73
23.38 ± 0.17
(b)
Graded + pairwise
Product
No
28.09
5.73
18.27
21.91 ± 0.70
(c)
Graded nDCG
Product
Yes
29.80
6.21
17.24
22.72 ± 0.18
(d)
Graded nDCG
Product
No
26.22
5.79
18.36
20.85 ± 0.68
(e)
Graded nDCG
Paired
No
22.55
6.68
17.78
18.71 ± 0.53
(f)
Graded nDCG
Query only
No
17.92
8.64
15.85
15.85 ± 0.26
Table 2: BRIGHT ablations (Q3E-0.6B, original queries). Avg. is nDCG@10 ( × 100), mean ± SD over three seeds; StackEx., Coding, and Theorem average 7, 2, and 3 subsets, respectively. Shading marks full RELER. Query-only and document-only sample one embedding side while updating the shared encoder. Full results are in Appendix C.3 .
Figure 2: Training dynamics and sensitivity. (a) CMP reduces gradient noise before fine-tuning (inset: zoom). (b) Pairwise feedback improves BRIGHT despite similar training nDCG reward curves. (Table 2 , rows (a,c) ). (c,d) BRIGHT nDCG@10 ( × 100) is highest at λ=0.5 and stable across ρ=0.60 – 0.90 . Filled points mark defaults; unsmoothed curves and error bars show three-seed means ± SD. Sweep details are in Appendix C.4 .
In-domain
Out-of-domain
Avg.
Method
EM
F1
Hit@10
EM
F1
Hit@10
EM
F1
Hit@10
Q3E-0.6B
33.79
44.17
65.67
31.71
39.40
51.94
32.30
40.76
55.87
InfoNCE
33.63
44.07
64.30
30.19
37.76
46.72
31.18
39.56
51.74
RLOO, graded nDCG
34.66
45.08
67.25
31.38
39.08
51.97
32.32
40.80
56.34
RLOO, answer F1
34.71
45.16
66.77
31.38
39.14
52.61
32.33
40.86
56.66
RLOO, nDCG + F1
34.82
45.14
66.87
32.21
39.48
52.19
32.96
41.10
56.38
Table 3: Fixed-index QA ( × 100), seed 42. In-domain/out-of-domain average 2/5 datasets; Avg. covers all seven. Bold marks column bests. Per-dataset results are in Appendix D.5 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Q3E-0.6B BRIGHT adaptation
Trainable parameters
Full shared encoder
Optimizer / learning rate
AdamW / 5×10−6
Global / per-device batch
128 / 16 (8 devices)
Training steps
113
Schedule / warm-up fraction
Linear / 0.03
Weight decay
0.01
Appendix
Table 4: Training settings for Q3E-0.6B BRIGHT adaptation. All 4B BRIGHT methods use LoRA (rank 16) with a learning rate of 2×10−4 .
Original query
GPT-4 reasoning query
Model
Method
42
3407
2026
42
3407
2026
Q3E-0.6B
Q3E-0.6B
15.10
–
–
20.38
–
–
InfoNCE
22.50
22.56
22.87
26.40
26.59
26.56
LambdaLoss
19.13
19.30
19.67
23.61
23.96
24.34
RELER
23.27
23.57
23.29
27.15
27.19
27.10
Q3E-4B
Q3E-4B
18.70
–
–
22.01
–
–
Appendix
Table 5: Per-seed BRIGHT nDCG@10 ( ×100 ), averaged over 12 subsets. The pretrained encoders are evaluated once. GPT-4 reasoning queries use the released gpt4_reason.query inputs.
StackExchange
Coding
Theorem-based
Method
Bio.
Earth.
Econ.
Psy.
Rob.
Stack.
Sus.
Leet.
Pony
AoPS
TheoQ.
TheoT.
Avg.
General-purpose methods
BM25
18.9
27.2
14.9
12.5
13.6
18.4
15.0
24.4
7.9
6.2
10.4
4.9
14.5
OpenAI-3-Large
23.3
26.7
19.5
27.6
12.8
14.3
20.5
23.6
2.4
8.5
23.5
11.7
17.9
Google-Gecko-1B-768
22.7
34.8
19.6
27.8
15.7
20.1
17.1
29.6
3.6
9.3
23.8
15.9
20.0
GritLM-7B
24.8
32.3
18.9
19.8
17.1
13.6
17.8
29.9
22.0
8.8
25.2
21.2
21.0
Appendix
Table 6: Original-query BRIGHT nDCG@10 ( ×100 ). Literature rows are from Table 2 of Chen et al. (2026) ; MS MARCO fine-tuning controls and pretrained Qwen3-Embedding rows are omitted. ReasonEmbed-Qwen3-4B uses Redapter. Our 4B rows repeat Table 1 ; Avg. reports three-seed mean ± sample SD for trained recipes. Literature and our results use different training and evaluation protocols. Abbreviations follow the main table.
Table 8: Complete BRIGHT subset results for the pairwise-weight and alignment sweeps in Figure 2 (c,d). The first column is λ in the upper block and ρ in the lower block.
SNR
CMP / RLOO
Encoder state
Microbatch
RLOO
CMP
V
t
Vt
Before fine-tuning
Stack Overflow
0.25
5.35
0.00226
1.14
0.00258
Math-theorem
0.27
6.00
0.00182
1.11
0.00203
Biology
0.20
4.85
0.00169
1.13
0.00191
Step 50
Stack Overflow
0.23
4.07
0.00322
1.14
0.00367
Math-theorem
0.22
3.81
0.00321
1.11
0.00358
Appendix
Table 9: Complete-encoder gradient measurements for the combined reward. Each row uses 64 paired rollout draws at one fixed encoder state and microbatch. V is the sample gradient variance and t is mean forward/backward time per draw. Ratios compare CMP to RLOO; SNR uses the bias-corrected signal estimate.
Setting
Value
Training budget / seed
One epoch / 42
Optimizer / learning rate
AdamW / 5×10−6
Global / per-device batch
128 / 16
Schedule / warm-up / weight decay
Cosine / 0.03 / 0.01
Training query limit
128 tokens
InfoNCE temperature / anchor coefficient
0.03 / 0.5
Appendix
Table 10: Fixed-index RAG training and generation settings. The query encoder is fully fine-tuned; the document index and generator remain frozen.
Method
NQ
HotpotQA
PopQA
TriviaQA
2Wiki
MuSiQue
Bamboogle
Queries ( n )
3,610
7,405
14,267
11,313
12,576
2,417
125
Answer EM
Q3E-0.6B
34.99
32.59
42.22
61.00
28.76
7.36
19.20
InfoNCE
37.01
30.26
42.07
60.26
23.81
5.63
19.20
RLOO, graded nDCG
36.01
33.32
41.82
60.54
29.03
6.48
19.02
RLOO, answer F1
35.94
33.49
42.40
61.43
28.98
7.27
16.80
Appendix
Table 11: Per-dataset fixed-index QA results ( ×100 ), seed 42. n is the number of evaluation queries; 2Wiki denotes 2WikiMultihopQA. The same document index, generator, and query counts are used for all recipes.
Objective
Classif.
Clust.
Pair class.
Rerank.
Retrieval
STS
Summ.
Task avg.
Type avg.
InfoNCE
72.84
44.19
82.11
44.80
51.52
75.94
29.49
60.99
57.27
RELER
73.41
45.81
79.76
45.13
50.04
76.81
26.61
61.01
56.80
Appendix
Table 12: Direct embedding training from Qwen3-0.6B on MTEB English v2 ( ×100 ), seed 42. Each type column averages its tasks; Task avg. weights all 41 tasks equally, while Type avg. weights the seven task-type means equally. Classif.: classification; Clust.: clustering; Pair class.: pair classification; Rerank.: reranking; Summ.: summarization. Bold marks the higher score in each column.