cs.IROct 6, 2026

Learning to Retrieve via Reinforcement Learning in Embedding Space

Authors: Qi Liu, Fengming Liang, Yiqun Chen, Erhan Zhang, Jiaxin Mao

Organizations: Renmin University of China

Abstract

Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DREAM: Dense Retrieval Embeddings via Autoregressive Modeling

    Jun 23, 2026Yixuan Tang, Yi YangRetrieval LayerAutoregressive Token Prediction

  2. GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

    Aug 13, 2026Zhili Shen, Craig MacdonaldMultimodal Embeddings

  3. Aligning Dense Retrievers with LLM Utility via Distillation

    Apr 24, 2026Rajinder Sandhu, Di Mu, Cheng Chang +4Retrieval LayerRetrievers