Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models
Organizations: TurkuNLP University of Turku, Finland · Ellis Institute Finland
Abstract
Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.
Figures & tables
| Retrieve parallel sentences. | Retrieve the corresponding translation in { lang } . |
|---|---|
| Retrieve text based on user query. | Retrieve semantically similar text. |
| Group documents related to the same topic. | asdfjkl qpwoeiru zxcvbnm |
| ARCChallenge(ndcg@10) | Tatoeba(recall@1) | |||||
|---|---|---|---|---|---|---|
| Model | top- | MV | top- | MV | ||
| BAAI/bge-m3 | 0.286 | 0.059 | 0.457* | 0.350 | 0.065 | 0.554* |
| codefuse-ai/F2LLM-v2-4B | 0.571 | 0.272 | 0.814 * | 0.538 | 0.014 | 0.750* |
| google/embeddinggemma-300m | 0.143 | 0.000 | -0.021 | 0.325 | 0.011 | 0.252* |
| intfloat/multilingual-e5-large-instruct | 0.429 | 0.082 | 0.571* | 0.575 | 0.052 | 0.787* |
| microsoft/harrier-oss-v1-0.6b | 0.429 | 0.103 | 0.543* | 0.588 | 0.056 | 0.81* |
| ARCChallenge (ndcg@10) | Tatoeba (recall@1) | |||||
|---|---|---|---|---|---|---|
| Model | displ. | sim. | angl. | displ. | sim. | angl. |
| BAAI/bge-m3 | 0.701 | 0.696 | 0.700 | 0.838 | 0.827 | 0.870 |
| codefuse-ai/F2LLM-v2-4B | 0.916 | 0.910 | 0.918 | 0.890 | 0.976 | 0.950 |
| google/embeddinggemma-300m | 0.520 | 0.184 | 0.448 | 0.631 | 0.655 | 0.681 |
| intfloat/multilingual-e5-large-instruct | 0.784 | 0.800 | 0.769 | 0.877 | 0.999 | 1.000 |
| microsoft/harrier-oss-v1-0.6b | 0.780 | 0.836 | 0.877 | 0.896 | 0.877 | 0.948 |
| ARCChallenge(ndcg@10) | Tatoeba(Recall@1) | |||||
|---|---|---|---|---|---|---|
| Model | displ. | sim. | angl. | displ. | sim. | angl. |
| BAAI/bge-m3 | - 0.823 * | 0.555 * | - 0.763 * | - 0.837 * | 0.839 * | - 0.510 * |
| codefuse-ai/F2LLM-v2-4B | 0.207 | 0.884 * | 0.522 * | - 0.826 * | 0.848 * | 0.168* |
| google/embeddinggemma-300m | - 0.772 * | 0.028 | - 0.542 * | - 0.658 * | 0.516 * | - 0.503 * |
| intfloat/multilingual-e5-large-instruct | - 0.639 * | -0.132 | -0.369* | - 0.699 * | 0.864 * | 0.118* |
| microsoft/harrier-oss-v1-0.6b | - 0.812 * | 0.049 | - 0.514 * | -0.323* | 0.54 * | 0.164* |
| ARCChallenge | Tatoeba (max) | ||||||
|---|---|---|---|---|---|---|---|
| Model | normal | distr. | rel. | normal | distr. | rel. | |
| BAAI/bge-m3 | 0.032 | 0.001 | 0.000 | 0.987 | 0.471 | – | |
| codefuse-ai/F2LLM-v2-4B | 0.214 | 0.062 | – | 0.998 | 0.626 | 0.55 | |
| google/embeddinggemma-300m | 0.019 | 0.005 | 0.002 | 0.725 | 0.25 | – | |
| intfloat/multilingual-e5-large-instruct | 0.074 | 0.007 | 0.0 | 0.994 | 0.954 | – | |
| microsoft/harrier-oss-v1-0.6b | 0.106 | 0.007 | – | 0.988 | 0.956 | – | |
| SQuAD | Tatoeba (mean) | |||||||
|---|---|---|---|---|---|---|---|---|
| Metric | Model | normal | random | paraphr. | normal | random | paraphr. | |
| Recall@1 | baseline | 0.724 | 0.685 | 0.014 | 0.844 | 0.769 | 0.102 | |
| fine-tuned | 0.706 | 0.702 | 0.470 | 0.837 | 0.833 | 0.569 | ||
| Recall@5 | baseline | 0.918 | 0.892 | 0.778 | 0.926 | 0.878 | 0.832 | |
| fine-tuned | 0.912 | 0.910 | 0.890 | 0.923 | 0.921 | 0.908 | ||
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Prompt |
|---|---|
| SQuAD | Paraphrase the following question while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased question and nothing else. Question: {input} |
| Tatoeba | Paraphrase the following text while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased text and nothing else. Text: {input} |
| ARCChallenge | Paraphrase the following text while preserving its meaning exactly. Use different wording and avoid producing a near-identical phrasing. Output only the paraphrased text and nothing else. Text: {input} |
| Model | Parameters | Dimensions | Citation |
|---|---|---|---|
| Octen/Octen-Embedding-8B | 7.6B | 4096 | Octen Team (2025) |
| codefuse-ai/F2LLM-v2-4B | 4B | 2560 | Zhang et al. (2026) |
| Qwen/Qwen3-Embedding-4B | 4B | 2560 | Zhang et al. (2025) |
| Qwen/Qwen3-Embedding-0.6B | 596M | 1024 | Zhang et al. (2025) |
| microsoft/harrier-oss-v1-0.6b | 596M | 1024 | Microsoft (2026) |
| BAAI/bge-m3 | 568M | 1024 | Chen et al. (2024) |
| Instruction | Source | Motivation | Relevant |
|---|---|---|---|
| ”Identify categories in user passages.” | MTEB, clustering | MTEB abstask | |
| ”Classify user passages.” | MTEB, classification | MTEB abstask | |
| ”Retrieve text that are semantically similar to the given text.” | MTEB, pair classification | MTEB abstask | |
| ”Retrieve text based on user query.” | MTEB, retrieval | MTEB abstask | Retrieval |
| ”Retrieve semantically similar text.” | MTEB, STS | MTEB abstask | |
| ”Given a news summary, retrieve other semantically similar summaries.” | MTEB, summarization | MTEB abstask |
| Dataset | Most improved instructions | Impr. |
|---|---|---|
| Tatoeba:cmn | Retrieve the corresponding translation in Mandarin Chinese. | 0.49 |
| Given an English sentence, find its translation in Mandarin Chinese. | 0.38 | |
| Retrieve the corresponding translation. | 0.30 | |
| Tatoeba:fin | Retrieve the corresponding translation in Finnish. | 0.23 |
| Given an English sentence, find its translation in Finnish. | 0.21 | |
| Translate to Finnish. | 0.16 |
| Recall 1 | Recall 5 | NDCG 10 | ||||
|---|---|---|---|---|---|---|
| Model | Normal | Paraphr. | Normal | Paraphr. | Normal | Paraphr. |
| baseline | 0.059 | 0.000 | 0.183 | 0.138 | 0.139 | 0.097 |
| SQuAD-finetuned | 0.051 | 0.003( ) | 0.173 | 0.142( ) | 0.130 | 0.094 ( ) |