Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at https://github.com/nourj98/RICE.
Figures & tables
Retriever
LLM
ArguAna
FiQA
News
NFCorpus
NQ
Robust04
SCIDOCS
SciFact
Signal-1M
COVID
Avg.
BM25
–
.932
.539
.447
.246
.751
.375
.348
.925
.370
.109
.504
Dense Retrieval w/ Training
BGE-base-en-v1.5
–
.992
.742
.499
.337
.942
.351
.496
.967
.311
.141
.578
Qwen3-Embedding-8B
Qwen3-8B
.996
.929
.570
.388
.977
.494
.636
.973
.292
.194
.645
LLM2Vec-Gen
Qwen3-8B
.988
.670
.484
.310
.871
.356
.430
.948
.245
.122
.542
Promptodile
Llama-3.1-8B
.994
.709
–
.316
–
–
.409
.964
–
–
–
Table 1: Main results (Recall@100) across BEIR datasets. Dense Retrieval w/ Training covers human labels, synthetic labels, and self-supervised objectives. Bold denotes the best result within each model family under Dense Retrieval w/o Training.
Figure 1: Recall@100 across varying numbers of in-context examples for generating RICE (Qwen3.5-9B) query encodings.
Dynamic Examples
FiQA
NFCorpus
SciFact
Fixed
.700
.343
.955
Query-Specific
.710
.360
.977
Document-Specific
.672
.333
.933
Query & Document-Specific
.719
.348
.973
Table 2: Fixed and dynamic exemplar selection with RICE using Qwen3.5-9B.
Example Pairs
NFCorpus
SCIDOCS
None
.305
.457
Relevant
.343
.478
Non-Relevant
.345
.473
Mixed
.345
.484
Random
.346
.497
Relevant (cross-corpus)
.312
.457
Table 3: RICE (Qwen3.5-9B) using different in-context example compositions. None denotes RICE (Zero-Shot). Relevant (cross-corpus) uses examples from Robust04.
Representative Word
NFCorpus
SciFact
Original
.343
.955
Shuffled
.355
.950
Blank
.343
.957
Fixed: “word”
.337
.958
Fixed: “orange”
.326
.928
Fixed: “feather”
.345
.948
Table 4: Representative word ablations for RICE with Qwen3.5-9B. Random sets w1,…,w10 to feather, basket, curtain, candle, window, compass, marble, lantern, violin, and saddle.
Figure 2: Prompt templates used for PromptReps and RICE. The first two panels show the query- and document-encoding prompts used by PromptReps. The next two panels show the query-side in-context example and target prompt used by RICE. The final two panels show the corresponding document-side prompts. RICE (Zero-Shot) only utilizes the RICE target query-encoding and document-encoding prompts, without in-context examples.