Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.
Figures & tables
Figure 1: One customer message read by Qwen3-1.7B under two questions, each token shaded by its reading weight.
Figure 2
Figure 4: Attention Relay: the LLM’s reading weights, carried through characters to the embedder’s tokens, pool the embedder’s own states.
Clustering, V-measure ↑
STS, Spearman ↑
Triplets ↑
Embedder
NYT
FewRel
FewNerd
FewEvent
InstructSTSB
IE
Qwen3-Embedding-0.6B (document mode)
55.9
54.4
42.7
59.1
0.0
100.0
+ relay
64.6
57.6
49.5
63.5
18.2
142.6
Gain
+8.7
+3.2
+6.8
+4.5
+18.2
+42.6
bge-large-en-v1.5
56.7
55.4
43.3
59.0
0.0
100.0
+ relay
67.2
58.8
48.0
64.3
20.3
157.8
Table 1: Five embedders that take no instruction, alone and relayed from Qwen3-1.7B (+ relay), and Qwen3-Embedding-0.6B trained with instructions (bottom row). IE is 100 when the question is ignored and 200 when it is always followed.
Figure 5: The gain of relaying for every pair of an instruction-tuned LLM and an embedder.
Figure 6: The gain of relaying from each instruction-tuned LLM and from its base checkpoint.
Figure 7: The gain of relaying from each block of each LLM; dots mark the blocks we read.
Figure 8
Figure 10: Probe accuracy and V-measure of the asked aspect, from each embedder alone to relayed.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Hugging Face
Instruction-tuned LLMs
Qwen3-0.6B
Qwen/Qwen3-0.6B
Qwen3-1.7B
Qwen/Qwen3-1.7B
Qwen3-4B
Qwen/Qwen3-4B
Qwen3-8B
Qwen/Qwen3-8B
Llama-3.1-8B-Instruct
meta-llama/Llama-3.1-8B-Instruct
Appendix
Table 2: Every model in the paper.
Dataset
Hugging Face
Used throughout
NYT
BrandonZYW/NYTClustering
IntentEmotion (IE)
BrandonZYW/IntentEmotion
Used in § 4
FewRel
BrandonZYW/FewRelClustering
FewNerd
BrandonZYW/FewNerdClustering
Appendix
Table 3: Every dataset, from InBedder’s evaluation ( Peng et al., 2024 ) .
LLM
Blocks
Block read
Qwen3-0.6B
28
24
Qwen3-1.7B
28
24
Qwen3-4B
36
29
Qwen3-8B
36
29
Llama-3.1-8B-Instruct
32
24
OLMo-3-7B-Instruct
32
23
Appendix
Table 4: Each LLM, the block its reading weights are taken at.
Embedder
Pooling
Relayed into
Max. tokens
Qwen3-Embedding-0.6B
last token
attention
–
Qwen3-Embedding-4B
last token
attention
–
Qwen3-Embedding-8B
last token
attention
–
bge-large-en-v1.5
[CLS]
attention
512
bge-small-en-v1.5
[CLS]
attention
512
bge-base-en-v1.5
[CLS]
attention
512
Appendix
Table 5: The setting of each embedder, the same with every LLM.
Dataset
Instruction
NYT, topic
What is the topic of news?
NYT, location
Where did the news happen?
IntentEmotion, emotion
How does the customer feel?
IntentEmotion, intent
What does the customer need?
FewRel
Here is a sentence. Please tell me the relation type between two specified entities appended after the sentence.
FewNerd
Here is a sentence. Please tell me the type of the specified entity appended after the sentence.
Appendix
Table 6: The instruction each dataset puts in the template’s instruction slot.
Embedder
Condition
Input
Every embedder but e5-large-v2
alone and relayed
{text}
e5-large-v2
alone and relayed
query: {text}
all-mpnet-base-v2, bge-large, gte-large
the question in the input, before or after the text (Figure 3 )
{question} {text} {text} {question}
e5-large-v2
the same
query: {question} {text} query: {text} {question}
bge-large, gte-large, all-mpnet-base-v2
the set’s instruction in the input (Table 10 )
{question} {text}
e5-large-v2
the same
query: {question} {text}
Appendix
Table 7: What each embedder reads around the text; \n is a line break.
Question
Task description
NYT, topic
Identify the topic of the given news article
NYT, location
Identify the location where the given news article happened
IE, emotion
Identify the emotion the customer expresses in the given message
IE, intent
Identify what the customer needs in the given message
Appendix
Table 8: Task descriptions Qwen3-Embedding-0.6B reads in its instruction mode for NYT and IE.
Embedder
NYT
FewRel
FewNerd
FewEvent
InstructSTSB
IE
Qwen3-Embedding-0.6B
+8.7
+3.2
+6.8
+4.5
+18.2
+42.6
[2.9, 12.2]
[1.5, 5.0]
[4.4, 8.2]
[2.4, 6.4]
[16.6, 20.0]
[40.2, 45.1]
bge-large-en-v1.5
+10.6
+3.4
+4.7
+5.3
+20.3
+57.8
[7.5, 16.2]
[1.8, 5.0]
[2.1, 6.0]
[3.1, 6.7]
[18.5, 22.1]
[54.9, 60.6]
gte-large-en-v1.5
+12.0
+4.3
+3.5
+4.3
+17.9
+42.3
[6.7, 16.0]
[2.4, 5.8]
[1.3, 4.9]
[2.3, 5.7]
[16.3, 19.5]
[39.7, 44.9]
Appendix
Table 9: The gains of Table 1 , with their 95% paired bootstrap intervals below them.
Embedder
FewRel
FewNerd
FewEvent
InstructSTSB
Qwen3-1.7B
Qwen3-Embedding-0.6B
54.4 → 57.6
42.7 → 49.5
59.1 → 63.5
0.0 → 18.2
bge-large-en-v1.5
55.4 → 58.8
43.3 → 48.0
59.0 → 64.3
0.0 → 20.3
gte-large-en-v1.5
54.1 → 58.4
42.3 → 45.8
56.8 → 61.1
0.0 → 17.9
all-mpnet-base-v2
52.1 → 55.8
40.9 → 45.9
55.5 → 61.7
0.0 → 11.6
e5-large-v2
56.4 → 57.7
41.4 → 43.2
61.1 → 64.0
0.0 → 7.4
Appendix
Table 10: InBedder’s other sets with the weights of Qwen3-1.7B, Llama-3.1-8B-Instruct and OLMo-3-7B-Instruct: each embedder alone, relayed, and given the instruction in its input.
Figure 19
Model
alone
relayed
Qwen3-1.7B, the full pass
4.58
Qwen3-Embedding-0.6B
2.28
2.61
bge-large-en-v1.5
0.70
0.98
gte-large-en-v1.5
1.13
1.40
all-mpnet-base-v2
2.58
2.56
e5-large-v2
5.88
5.99
Appendix
Table 11: Milliseconds per text and question on one H100: the full pass of Qwen3-1.7B, and each embedder alone and relayed.
In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We provide the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, across three families (Llama-3.2-3B, Qwen-2.5-3B, Gemma-3-4B). Five independent methods, node and edge attribution, superposition role analysis, causal ablation, and path patching, converge on gating, with the same heads, in the same late-layers, are found to be reweighted rather than replaced with a high node overlap (0.60-0.82). Behaviorally, tuning shifts models toward parametric memory, making instruct models reject a terse counterfactual context far more than base ones, the opposite of a naive user-following expectation. Yet this added skepticism is a factor of framing since it disappears when the same false claim is delivered as a coherent, evidential passage. The robustness that instruction tuning buys against terse injection is therefore real but narrow. More broadly, we believe that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.
Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottlenecks in computational efficiency and cross-architecture transferability. Whenever a new backbone emerges, existing approaches require costly retraining from scratch. To address this, we propose PromptEmbedder, a novel dual-LLM framework that decouples embedding knowledge from specific backbone weights. PromptEmbedder utilizes a Prompting LLM to generate instruction-aware soft prompts for a frozen Embedding LLM via a differentiable generation process with continuous relaxation, ensuring full gradient flow during contrastive training. By localizing task-specific knowledge within the Prompting LLM, adapting to new architectures requires only retraining a lightweight linear alignment matrix. Evaluations on the MTEB benchmark show that PromptEmbedder achieves comparable performance with LoRA finetuning while reducing GPU memory by 40% and accelerating training by 3.7x. Our approach establishes a scalable, architecture-agnostic paradigm for efficient LLM-based representation learning.
Yu-Che Tsai, Kuan-Yu Chen, Yuan-Hao Chen +4
Department of Computer Science and Information Engineering, National Taiwan University · National Taiwan University AI Center of Research Excellence Taipei, Taiwan
Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruction. We present an empirical study of prompt sensitivity across 6 embedding models, 11 datasets, and 15 task-specific prompts per dataset, a total of 990. We show that reported scores misrepresent the distribution of scores over plausible prompts. The default prompt can both systematically understate or overstate performance. Furthermore, we show that the leaderboard ranking is not robust to prompt selection: by choosing prompts favorably, any model in our study can be promoted to first place. Our findings suggest that single-prompt evaluation is insufficient for instruction-tuned embedding models and that benchmarks should incorporate prompt robustness, either by evaluating over multiple prompts or by reporting sensitivity alongside point estimates.