Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at https://github.com/nourj98/RICE.
Figures & tables
Retriever
LLM
ArguAna
FiQA
News
NFCorpus
NQ
Robust04
SCIDOCS
SciFact
Signal-1M
COVID
Avg.
BM25
–
.932
.539
.447
.246
.751
.375
.348
.925
.370
.109
.504
Dense Retrieval w/ Training
BGE-base-en-v1.5
–
.992
.742
.499
.337
.942
.351
.496
.967
.311
.141
.578
Qwen3-Embedding-8B
Qwen3-8B
.996
.929
.570
.388
.977
.494
.636
.973
.292
.194
.645
LLM2Vec-Gen
Qwen3-8B
.988
.670
.484
.310
.871
.356
.430
.948
.245
.122
.542
Promptodile
Llama-3.1-8B
.994
.709
–
.316
–
–
.409
.964
–
–
–
Table 1: Main results (Recall@100) across BEIR datasets. Dense Retrieval w/ Training covers human labels, synthetic labels, and self-supervised objectives. Bold denotes the best result within each model family under Dense Retrieval w/o Training.
Figure 1: Recall@100 across varying numbers of in-context examples for generating RICE (Qwen3.5-9B) query encodings.
Dynamic Examples
FiQA
NFCorpus
SciFact
Fixed
.700
.343
.955
Query-Specific
.710
.360
.977
Document-Specific
.672
.333
.933
Query & Document-Specific
.719
.348
.973
Table 2: Fixed and dynamic exemplar selection with RICE using Qwen3.5-9B.
Example Pairs
NFCorpus
SCIDOCS
None
.305
.457
Relevant
.343
.478
Non-Relevant
.345
.473
Mixed
.345
.484
Random
.346
.497
Relevant (cross-corpus)
.312
.457
Table 3: RICE (Qwen3.5-9B) using different in-context example compositions. None denotes RICE (Zero-Shot). Relevant (cross-corpus) uses examples from Robust04.
Representative Word
NFCorpus
SciFact
Original
.343
.955
Shuffled
.355
.950
Blank
.343
.957
Fixed: “word”
.337
.958
Fixed: “orange”
.326
.928
Fixed: “feather”
.345
.948
Table 4: Representative word ablations for RICE with Qwen3.5-9B. Random sets w1,…,w10 to feather, basket, curtain, candle, window, compass, marble, lantern, violin, and saddle.
Figure 2: Prompt templates used for PromptReps and RICE. The first two panels show the query- and document-encoding prompts used by PromptReps. The next two panels show the query-side in-context example and target prompt used by RICE. The final two panels show the corresponding document-side prompts. RICE (Zero-Shot) only utilizes the RICE target query-encoding and document-encoding prompts, without in-context examples.
Dense retrieval embedding models are a fundamental component of modern retrieval-based AI systems. Most dense retrievers are trained with contrastive objectives, which require labeled positive and negative document pairs that are often costly and difficult to obtain. In this work, we investigate whether the autoregressive next-token prediction objective of a large language model (LLM) can provide supervision for dense retrieval. The intuition is simple: if a document contains information relevant to a query, conditioning on that document should make the target output easier for the LLM to predict. A key challenge is that the next-token prediction loss is computed inside the LLM, while the retriever is a separate embedding model. To address this challenge, we propose DREAM (Dense Retrieval Embeddings via Autoregressive Modeling), which injects retriever-generated query-document similarity scores into selected attention heads of a frozen LLM. During training, these scores determine how much attention each candidate document receives while the LLM predicts the target output. The resulting prediction loss provides gradients for retriever training through the attention mechanism. We evaluate DREAM on retrieval benchmarks BEIR and RTEB using embedding backbones ranging from 0.5B to 3B parameters. DREAM consistently outperforms existing baselines across different model scales. These results demonstrate that DREAM provides a promising approach for training dense retrievers through autoregressive modeling.
Yixuan Tang, Yi Yang
The Hong Kong University of Science and Technology
Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a growing gap between how users express their needs and how retrievers interpret them. In this paper, we present GEM, a generative embedding model that augments retrieval through its own knowledge by explicitly reasoning about user intent and relevance criteria. GEM unifies generation and embedding within a single model: it first reasons over the query, then appends an embedding token to encode the enriched context for retrieval. \zhili{Evaluated on reasoning-intensive and instruction-following retrieval tasks, GEM demonstrates the effectiveness of its reasoning-augmented retrieval, outperforming its non-reasoning variant and matching baselines using substantially larger models.} Furthermore, GEM's generative nature allows test-time compute scaling via prompting to further enhance retrieval performance. Our code is available at: https://anonymous.4open.science/r/GEM.
Decoder-only large language models (LLMs) are increasingly replacing BERT-style architectures as the backbone for dense retrieval, achieving substantial performance gains and broad adoption. However, the robustness of these LLM-based retrievers remains underexplored. In this paper, we present the first systematic study of the robustness of state-of-the-art open-source LLM-based dense retrievers from two complementary perspectives: generalizability and stability. For generalizability, we evaluate retrieval effectiveness across four benchmarks spanning 30 datasets, using linear mixed-effects models to estimate marginal mean performance and disentangle intrinsic model capability from dataset heterogeneity. Our analysis reveals that while instruction-tuned models generally excel, those optimized for complex reasoning often suffer a ``specialization tax,'' exhibiting limited generalizability in broader contexts. For stability, we assess model resilience against both unintentional query variations~(e.g., paraphrasing, typos) and malicious adversarial attacks~(e.g., corpus poisoning). We find that LLM-based retrievers show improved robustness against typos and corpus poisoning compared to encoder-only baselines, yet remain vulnerable to semantic perturbations like synonymizing. Further analysis shows that embedding geometry (e.g., angular uniformity) provides predictive signals for lexical stability and suggests that scaling model size generally improves robustness. These findings inform future robustness-aware retriever design and principled benchmarking. Our code is publicly available at https://github.com/liyongkang123/Robust_LLM_Retriever_Eval.
Yongkang Li, Panagiotis Eustratiadis, Yixing Fan +1
University of Amsterdam, The Netherlands · Chinese Academy of Sciences, China