Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.
Figures & tables
Retrieve parallel sentences.
Retrieve the corresponding translation in { lang } .
Retrieve text based on user query.
Retrieve semantically similar text.
Group documents related to the same topic.
asdfjkl qpwoeiru zxcvbnm
Table 1: Examples of instructions in our experiments, with {lang} replaced with target language in Tatoeba evaluation. Instructions are intentionally varied.
ARCChallenge(ndcg@10)
Tatoeba(recall@1)
Model
top- k
R2
MV
top- k
R2
MV
BAAI/bge-m3
0.286
0.059
0.457*
0.350
0.065
0.554*
codefuse-ai/F2LLM-v2-4B
0.571
0.272
0.814 *
0.538
0.014
0.750*
google/embeddinggemma-300m
0.143
0.000
-0.021
0.325
0.011
0.252*
intfloat/multilingual-e5-large-instruct
0.429
0.082
0.571*
0.575
0.052
0.787*
microsoft/harrier-oss-v1-0.6b
0.429
0.103
0.543*
0.588
0.056
0.81*
Table 2: Evaluation of relevant against other instructions. Star ( ∗ ) indicates significance p<0.05 .
Figure 1: ARCChallenge (up) and Tatoeba:fin-eng (down), displacement, sim. improvement, and angulation. Each marker is a separate instruction. On Tatoeba, relevant instructions are highly distinguishable, unlike ARCChallenge, where especially displacement is low for relevant instructions.
ARCChallenge (ndcg@10)
Tatoeba (recall@1)
Model
displ.
sim.
angl.
displ.
sim.
angl.
BAAI/bge-m3
0.701
0.696
0.700
0.838
0.827
0.870
codefuse-ai/F2LLM-v2-4B
0.916
0.910
0.918
0.890
0.976
0.950
google/embeddinggemma-300m
0.520
0.184
0.448
0.631
0.655
0.681
intfloat/multilingual-e5-large-instruct
0.784
0.800
0.769
0.877
0.999
1.000
microsoft/harrier-oss-v1-0.6b
0.780
0.836
0.877
0.896
0.877
0.948
Table 3: Cross-validated ROC-AUC of logistic regression using the evaluation score and each structural metric (displacement (displ.), similarity improvement (sim.) and angulation (angl.)) as the predictors. High values indicate greater separability between relevant and other instructions, with 0.5 corresponding to chance level. Relevant instructions are more separable on Tatoeba.
ARCChallenge(ndcg@10)
Tatoeba(Recall@1)
Model
displ.
sim.
angl.
displ.
sim.
angl.
BAAI/bge-m3
- 0.823 *
0.555 *
- 0.763 *
- 0.837 *
0.839 *
- 0.510 *
codefuse-ai/F2LLM-v2-4B
0.207
0.884 *
0.522 *
- 0.826 *
0.848 *
0.168*
google/embeddinggemma-300m
- 0.772 *
0.028
- 0.542 *
- 0.658 *
0.516 *
- 0.503 *
intfloat/multilingual-e5-large-instruct
- 0.639 *
-0.132
-0.369*
- 0.699 *
0.864 *
0.118*
microsoft/harrier-oss-v1-0.6b
- 0.812 *
0.049
- 0.514 *
-0.323*
0.54 *
0.164*
Table 4: Spearman correlations between evaluation score and displacement (displ.), similarity improvement (sim.) and angulation (angl.). Values summarize the monotonicity seen in the results: the relationship between low displacement and high score is consistent in both datasets, while similarity improvement and angulation show more variation. Star ( ∗ ) indicates significance p<0.05 .
Figure 2: Taoteba:deu for multilingual-e5-large-instruct (left) and harrier-oss-v1-0.6b (right): Evaluation without (up) and with distractors (down). Each circle corresponds to a relevant instruction-query, and its position indicates its relative normalized displacement between Q=query, D=distractor, and A=target. The more movement towards the answer, the better the score is retained.
ARCChallenge
Tatoeba (max)
Model
normal
distr.
rel.
normal
distr.
rel.
BAAI/bge-m3
0.032
0.001
0.000
0.987
0.471
–
codefuse-ai/F2LLM-v2-4B
0.214
0.062
–
0.998
0.626
0.55
google/embeddinggemma-300m
0.019
0.005
0.002
0.725
0.25
–
intfloat/multilingual-e5-large-instruct
0.074
0.007
0.0
0.994
0.954
–
microsoft/harrier-oss-v1-0.6b
0.106
0.007
–
0.988
0.956
–
Table 5: Distractor results for ARCChallenge and Tatoeba, both with maximal recall@1 over all instructions (and all languages in Tatoeba): normal column refers to normal evaluation setup with no distractors, distr. column to setup where distractors are introduced to the target corpus, and rel. showing the best performance of relevant instructions, if lower than the distr. column.
Figure 8
SQuAD
Tatoeba (mean)
Metric
Model
normal
random
paraphr.
normal
random
paraphr.
Recall@1
baseline
0.724
0.685
0.014
0.844
0.769
0.102
fine-tuned
0.706
0.702
0.470
0.837
0.833
0.569
Recall@5
baseline
0.918
0.892
0.778
0.926
0.878
0.832
fine-tuned
0.912
0.910
0.890
0.923
0.921
0.908
Table 6: Recall@1 and Recall@5 test results on SQuAD and Tatoeba for the baseline Qwen/Qwen3-Embedding-0.6B compared to fine-tuned models. Tatoeba results are reported as a mean across 8 languages (per language results in Appendix G ).
Figure 5: Performance on the MMTEB benchmark categories for fine-tuned models. The absolute differences are small, meaning our fine-tuning setup does not affect the overall performance.
Figure 6: SQuAD: before and after for relevant instructions, with and without distractors.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Prompt
SQuAD
Paraphrase the following question while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased question and nothing else. Question: {input}
Tatoeba
Paraphrase the following text while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased text and nothing else. Text: {input}
ARCChallenge
Paraphrase the following text while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased text and nothing else. Text: {input}
Appendix
Table 7: Prompts used to generate synthetic distractors.
Model
Parameters
Dimensions
Citation
Octen/Octen-Embedding-8B
7.6B
4096
Octen Team (2025)
codefuse-ai/F2LLM-v2-4B
4B
2560
Zhang et al. (2026)
Qwen/Qwen3-Embedding-4B
4B
2560
Zhang et al. (2025)
Qwen/Qwen3-Embedding-0.6B
596M
1024
Zhang et al. (2025)
microsoft/harrier-oss-v1-0.6b
596M
1024
Microsoft (2026)
BAAI/bge-m3
568M
1024
Chen et al. (2024)
Appendix
Table 8: Models used in this study.
Instruction
Source
Motivation
Relevant
”Identify categories in user passages.”
MTEB, clustering
MTEB abstask
”Classify user passages.”
MTEB, classification
MTEB abstask
”Retrieve text that are semantically similar to the given text.”
MTEB, pair classification
MTEB abstask
”Retrieve text based on user query.”
MTEB, retrieval
MTEB abstask
Retrieval
”Retrieve semantically similar text.”
MTEB, STS
MTEB abstask
”Given a news summary, retrieve other semantically similar summaries.”
MTEB, summarization
MTEB abstask
Appendix
Table 9: Instructions used in our experiments, their sources, reasoning for their selection, and which task they are considered to be relevant.
Figure 8: Displacement, similarity improvement, and angulation on outlier codefuse-ai/F2LLM-v2-4B against ndcg@10. Best performing and in all metrics clearly distinguishable instruction is “Retrieve the answer to the question.”
Figure 9: Relative distance plot for codefuse-ai/F2LLM-v2-4B on recall@1: While the best-performing instruction displays large displacement, high similarity increase, and angulation toward the correct target, distractors still inhibit performance greatly.
Dataset
Most improved instructions
Impr.
Tatoeba:cmn
Retrieve the corresponding translation in Mandarin Chinese.
0.49
Given an English sentence, find its translation in Mandarin Chinese.
0.38
Retrieve the corresponding translation.
0.30
Tatoeba:fin
Retrieve the corresponding translation in Finnish.
0.23
Given an English sentence, find its translation in Finnish.
0.21
Translate to Finnish.
0.16
Appendix
Table 10: Top 3 instructions that improved their performance (Recall @ 1) the most during fine-tuning.
Recall @ 1
Recall @ 5
NDCG @ 10
Model
Normal
Paraphr.
Normal
Paraphr.
Normal
Paraphr.
baseline
0.059
0.000
0.183
0.138
0.139
0.097
SQuAD-finetuned
0.051
0.003( ↑ )
0.173
0.142( ↑ )
0.130
0.094 ( ↓ )
Appendix
Table 11: Fine-tuned performance for ARCChallenge: The baseline performance for ARCChallenge is low, but recall performance with paraphrase distractors increases while rank-based NDCG goes down. Note that the rank-based approach does not reveal distractor performance as strongly as recall@1.
Figure 10: Language-specific results for Tatoeba fine-tuning checkpoints on the devevelopment data.
Figure 11: Development set performance during SQuAD and Tatoeba fine-tuning without query-side negatives. Continued fine-tuning without inserted distractors does not improve performance in the paraphrase distractor setting.
Figure 12: MTEB English v2 benchmark results for the fine-tuned models, compared with the base model. We report the absolute difference, where positive values indicate better performance after fine-tuning.
Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruction. We present an empirical study of prompt sensitivity across 6 embedding models, 11 datasets, and 15 task-specific prompts per dataset, a total of 990. We show that reported scores misrepresent the distribution of scores over plausible prompts. The default prompt can both systematically understate or overstate performance. Furthermore, we show that the leaderboard ranking is not robust to prompt selection: by choosing prompts favorably, any model in our study can be promoted to first place. Our findings suggest that single-prompt evaluation is insufficient for instruction-tuned embedding models and that benchmarks should incorporate prompt robustness, either by evaluating over multiple prompts or by reporting sensitivity alongside point estimates.
Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a growing gap between how users express their needs and how retrievers interpret them. In this paper, we present GEM, a generative embedding model that augments retrieval through its own knowledge by explicitly reasoning about user intent and relevance criteria. GEM unifies generation and embedding within a single model: it first reasons over the query, then appends an embedding token to encode the enriched context for retrieval. \zhili{Evaluated on reasoning-intensive and instruction-following retrieval tasks, GEM demonstrates the effectiveness of its reasoning-augmented retrieval, outperforming its non-reasoning variant and matching baselines using substantially larger models.} Furthermore, GEM's generative nature allows test-time compute scaling via prompting to further enhance retrieval performance. Our code is available at: https://anonymous.4open.science/r/GEM.
Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.