Lend Me Your Eyes: Instruction-Aware Text Embeddings via Attention Relay
Organizations: University of California, Santa Cruz
Abstract
Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.
Figures & tables
| Clustering, V-measure | STS, Spearman | Triplets | ||||
| Embedder | NYT | FewRel | FewNerd | FewEvent | InstructSTSB | IE |
| Qwen3-Embedding-0.6B (document mode) | 55.9 | 54.4 | 42.7 | 59.1 | 0.0 | 100.0 |
| + relay | 64.6 | 57.6 | 49.5 | 63.5 | 18.2 | 142.6 |
| Gain | +8.7 | +3.2 | +6.8 | +4.5 | +18.2 | +42.6 |
| bge-large-en-v1.5 | 56.7 | 55.4 | 43.3 | 59.0 | 0.0 | 100.0 |
| + relay | 67.2 | 58.8 | 48.0 | 64.3 | 20.3 | 157.8 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Hugging Face |
| Instruction-tuned LLMs | |
| Qwen3-0.6B | Qwen/Qwen3-0.6B |
| Qwen3-1.7B | Qwen/Qwen3-1.7B |
| Qwen3-4B | Qwen/Qwen3-4B |
| Qwen3-8B | Qwen/Qwen3-8B |
| Llama-3.1-8B-Instruct | meta-llama/Llama-3.1-8B-Instruct |
| Dataset | Hugging Face |
| Used throughout | |
| NYT | BrandonZYW/NYTClustering |
| IntentEmotion (IE) | BrandonZYW/IntentEmotion |
| Used in § 4 | |
| FewRel | BrandonZYW/FewRelClustering |
| FewNerd | BrandonZYW/FewNerdClustering |
| LLM | Blocks | Block read |
|---|---|---|
| Qwen3-0.6B | 28 | 24 |
| Qwen3-1.7B | 28 | 24 |
| Qwen3-4B | 36 | 29 |
| Qwen3-8B | 36 | 29 |
| Llama-3.1-8B-Instruct | 32 | 24 |
| OLMo-3-7B-Instruct | 32 | 23 |
| Embedder | Pooling | Relayed into | Max. tokens |
|---|---|---|---|
| Qwen3-Embedding-0.6B | last token | attention | – |
| Qwen3-Embedding-4B | last token | attention | – |
| Qwen3-Embedding-8B | last token | attention | – |
| bge-large-en-v1.5 | [CLS] | attention | 512 |
| bge-small-en-v1.5 | [CLS] | attention | 512 |
| bge-base-en-v1.5 | [CLS] | attention | 512 |
| Dataset | Instruction |
|---|---|
| NYT, topic | What is the topic of news? |
| NYT, location | Where did the news happen? |
| IntentEmotion, emotion | How does the customer feel? |
| IntentEmotion, intent | What does the customer need? |
| FewRel | Here is a sentence. Please tell me the relation type between two specified entities appended after the sentence. |
| FewNerd | Here is a sentence. Please tell me the type of the specified entity appended after the sentence. |
| Embedder | Condition | Input |
|---|---|---|
| Every embedder but e5-large-v2 | alone and relayed | {text} |
| e5-large-v2 | alone and relayed | query: {text} |
| all-mpnet-base-v2, bge-large, gte-large | the question in the input, before or after the text (Figure 3 ) | {question} {text} {text} {question} |
| e5-large-v2 | the same | query: {question} {text} query: {text} {question} |
| bge-large, gte-large, all-mpnet-base-v2 | the set’s instruction in the input (Table 10 ) | {question} {text} |
| e5-large-v2 | the same | query: {question} {text} |
| Question | Task description |
|---|---|
| NYT, topic | Identify the topic of the given news article |
| NYT, location | Identify the location where the given news article happened |
| IE, emotion | Identify the emotion the customer expresses in the given message |
| IE, intent | Identify what the customer needs in the given message |
| Embedder | NYT | FewRel | FewNerd | FewEvent | InstructSTSB | IE |
|---|---|---|---|---|---|---|
| Qwen3-Embedding-0.6B | +8.7 | +3.2 | +6.8 | +4.5 | +18.2 | +42.6 |
| [2.9, 12.2] | [1.5, 5.0] | [4.4, 8.2] | [2.4, 6.4] | [16.6, 20.0] | [40.2, 45.1] | |
| bge-large-en-v1.5 | +10.6 | +3.4 | +4.7 | +5.3 | +20.3 | +57.8 |
| [7.5, 16.2] | [1.8, 5.0] | [2.1, 6.0] | [3.1, 6.7] | [18.5, 22.1] | [54.9, 60.6] | |
| gte-large-en-v1.5 | +12.0 | +4.3 | +3.5 | +4.3 | +17.9 | +42.3 |
| [6.7, 16.0] | [2.4, 5.8] | [1.3, 4.9] | [2.3, 5.7] | [16.3, 19.5] | [39.7, 44.9] |
| Embedder | FewRel | FewNerd | FewEvent | InstructSTSB |
|---|---|---|---|---|
| Qwen3-1.7B | ||||
| Qwen3-Embedding-0.6B | 54.4 57.6 | 42.7 49.5 | 59.1 63.5 | 0.0 18.2 |
| bge-large-en-v1.5 | 55.4 58.8 | 43.3 48.0 | 59.0 64.3 | 0.0 20.3 |
| gte-large-en-v1.5 | 54.1 58.4 | 42.3 45.8 | 56.8 61.1 | 0.0 17.9 |
| all-mpnet-base-v2 | 52.1 55.8 | 40.9 45.9 | 55.5 61.7 | 0.0 11.6 |
| e5-large-v2 | 56.4 57.7 | 41.4 43.2 | 61.1 64.0 | 0.0 7.4 |
| Model | alone | relayed |
|---|---|---|
| Qwen3-1.7B, the full pass | 4.58 | |
| Qwen3-Embedding-0.6B | 2.28 | 2.61 |
| bge-large-en-v1.5 | 0.70 | 0.98 |
| gte-large-en-v1.5 | 1.13 | 1.40 |
| all-mpnet-base-v2 | 2.58 | 2.56 |
| e5-large-v2 | 5.88 | 5.99 |