Draft-KV: Learning Useful Latent Communication Between Language Models
Authors: Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, +2 more
Organizations: City University of Hong Kong · University of Science and Technology of China · University of Electronic Science and Technology of China · AIPD, Tencent · Theory Lab, Huawei · Hong Kong Metropolitan University
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
Figures & tables
Evaluation setting
Score (%)
Model Pair
Benchmark
Protocol
Receiver
Sharer
Text-to-Text
Cache-to-Cache
Draft-KV
Qwen3-0.6B → Qwen2.5-0.5B
MMLU-Redux
37.45
45.19
42.56
33.43 /33.40
46.11 /36.59
ARC-Easy
63.47
69.15
66.71
56.31 /53.11
74.28 /61.91
ARC-Challenge
40.10
51.71
49.49
39.51 /35.32
54.52 /39.76
OpenBookQA
43.40
44.60
45.60
41.00 /40.00
52.00 /41.20
C-EVAL
Public
41.75
41.38
41.46
37.74 /36.11
47.40 /40.56
Table 1: Main results across model pairs, benchmarks, and evaluation protocols. Public settings are scored by option accuracy, Private settings by exact match. Cache-to-Cache and Draft-KV entries show Matched/Deranged scores, where the Deranged message comes from a different question. Bold and underlined values denote the best and second-best score in each row. Model names omit the -Instruct suffix of Qwen2.5-0.5B-Instruct and Llama-3.2-3B-Instruct; Qwen3 names are as released (Appendix C.1 ).
ARC-Challenge
MMLU-Redux
Variant
Matched
Pairing gain
Matched
Pairing gain
Full Draft-KV
83.53
45.90
65.07
28.23
Prompt KV
49.83
0.17
43.59
0.09
w/o reconstruction pretraining
76.71
38.74
52.40
16.58
w/o answer alignment
80.03
41.72
55.40
19.25
Table 2: Message-source and training-stage ablations for Qwen3-1.7B → Qwen2.5-0.5B-Instruct under the Public protocol.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Mean gain
Mismatch harm AR−AD
Protocol
n
G
P
Mean
Median
Min
Max
Public
35
24.24
25.65
1.41
1.19
−3.48
11.65
Private
14
2.51
13.49
10.98
11.30
−5.51
19.91
Appendix
Table 3: Distribution of Draft-KV intervention contrasts across the settings of Table 1 . All quantities are in percentage points.
C2C
Draft-KV
Sharer → Receiver
AR
AM
Aadapter
ρadapter
AM
Aadapter
ρadapter
MMLU-Redux
Q3-0.6B → Q2.5-0.5B
37.45
33.43
31.92
95.48%
46.11
40.07
86.90%
Q3-1.7B → Q2.5-0.5B
37.45
34.16
34.43
100.79%
65.07
33.22
51.05%
Q3-4B → Q2.5-0.5B
37.45
34.57
37.07
107.23%
75.98
36.17
47.60%
Q2.5-0.5B → Q3-0.6B
30.63
42.92
41.23
96.06%
45.01
37.64
83.63%
Appendix
Table 4: Adapter-only comparison across three benchmarks and four model pairs. All accuracies and retained accuracies are percentages. Receiver-only and Matched follow Table 1 ; only adapter-only accuracies are added from the new evaluations. Q3 and Q2.5 denote Qwen3 and Qwen2.5-Instruct. Retained accuracy ρadapter=Aadapter/AM . Values above 100% mean adapter-only exceeds Matched.
Sharer → Receiver
Receiver-only
Adapter-only
Self-donor
Q3-0.6B → Q2.5-0.5B
36.13
39.65
37.89
Q3-1.7B → Q2.5-0.5B
36.13
33.59
60.55
Q3-4B → Q2.5-0.5B
36.13
34.96
64.45
Q2.5-0.5B → Q3-0.6B
37.50
41.02
42.19
Appendix
Table 5: Draft-KV self-donor diagnostic on the first 512 MMLU-Redux examples. All entries are accuracy (%). Self-donor and adapter-only use the same one-shot queries. All comparisons use this subset.
Sharer → Receiver
Layers ( ℓ←π(ℓ) )
CS
CR
HR
Projections
∣θ∣
Qwen3-0.6B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-1.7B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-4B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-8B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Llama-3.2-3B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen2.5-0.5B → Qwen3-0.6B
18,20,22,24 ← 14,16,18,20
128
1024
8
1,048,576
1,048,608
Appendix
Table 6: Interface configuration per model pair. Layer assignments are written receiver ← sharer. CS and CR are flattened KV widths, HR the number of receiver KV heads and hence of gates. Model names omit -Instruct suffixes as in Table 1 .
Stage 1
Stage 2
Sharer → Receiver
micro × accum
seed
micro × accum
seed
Qwen3-0.6B → Qwen2.5-0.5B
16×4
20260826
4×8
20260828
Qwen3-1.7B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Qwen3-4B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Qwen3-8B → Qwen2.5-0.5B
4×16
91827
4×8
91827
Llama-3.2-3B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Appendix
Table 7: Per-pair microbatch × gradient accumulation and data seeds for Stages 1 and 2. Effective batches are 64 and 32 respectively for every pair. Model names omit -Instruct suffixes as in Table 1 .
Latency by stage (ms)
Total latency
Arithmetic
Condition
Sharer prefill
Draft decode
Extrac- tion
Projec- tion
Recv. prefill
Recv. decode
ms
×
TFLOPs
×
Interface GFLOPs
Public protocol, MMLU-Redux
Receiver-only
—
—
—
—
16.7
—
16.7
0.001
0.108
0.016
—
Sharer-only
33.6
11,868.1
—
—
—
—
11,907.2
1.000
6.801
1.000
—
Text-to-Text
33.6
11,868.1
—
—
18.0
—
11,925.2
1.002
7.187
1.057
—
Cache-to-Cache
45.1
—
—
35.6
20.7
—
101.4
0.009
2.256
0.332
118.3
Appendix
Table 8: Inference cost for one question with the Qwen3-8B sharer and the Qwen2.5-0.5B-Instruct receiver. Stage latencies are medians over 200 questions at batch size one after 20 warmup questions, timed with CUDA events. Floating-point counts are analytic; the interface column covers the projections of Equation 2 and the external attention of Equation 10 , and is reported in GFLOPs. Columns marked × give the ratio to Sharer-only on the same benchmark. Dashes mark stages a condition does not run.
Variant
AM
AD
AR
G
P
ARC-Challenge test
Full Draft-KV
83.53
37.63
40.10
43.43
45.90
Prompt KV
49.83
49.66
40.10
9.73
0.17
w/o reconstruction pretraining
76.71
37.97
40.10
36.61
38.74
w/o answer alignment
80.03
38.31
40.10
39.93
41.72
MMLU-Redux test
Appendix
Table 9: Complete accuracy contrasts for the core ablations on the two reported benchmarks. Accuracies are percentages; G and P are percentage points. Differences are computed from the reported two-decimal accuracies.
Variant
Benchmark
AM
AD
G
P
AR−AD
hˉ
ν
Full
ARC-C
83.53
37.63
43.43
45.90
2.47
0.038
19.7 (231)
w/o Guard
ARC-C
82.94
35.92
42.84
47.02
4.18
0.091
37.9 (444)
Full
MMLU-R
65.07
36.84
27.62
28.23
0.61
0.052
24.6 (1386)
w/o Guard
MMLU-R
64.68
36.24
27.23
28.44
1.21
0.104
42.8 (2411)
Appendix
Table 10: Guard ablation. Accuracies are percentages; G , P , and AR−AD are in pp. hˉ is the mean hinge and ν the threshold-exceedance rate, with the number of affected examples in parentheses. Receiver-only baselines are 40.10% (ARC-C) and 37.45% (MMLU-R).
Sharer
MMLU-R
ARC-E
ARC-C
OBQA
C-EVAL
Mean
G (pp)
Qwen3-0.6B
85.3±88.1
49.5±33.8
58.5±41.9
24.4±27.0
73.2±93.0
58.2
9.63
Qwen3-1.7B
293.1±112.9
219.9±73.0
243.4±81.7
220.7±66.3
332.1±120.0
261.8
30.57
Qwen3-4B
263.5±108.4
198.2±59.5
212.9±69.4
199.8±58.7
242.2±128.0
223.3
39.14
Qwen3-8B
358.9±102.3
263.6±67.0
285.3±76.7
251.5±60.4
355.8±113.3
303.0
42.25
Appendix
Table 11: Mean draft length in decoded tokens (mean ± SD) for the four scaling points, all decoded greedily under a 512-token cap. The last column is the unweighted mean over the five benchmarks, shown against the mean system gain of Figure 5 .
Sharer
Benchmark
Reported
Retrained
Qwen2.5-0.5B-Instruct
MMLU-Redux
42.92
42.95
OpenBookQA
52.60
52.80
ARC-Challenge
54.52
54.18
C-EVAL
41.77
41.68
Llama-3.2-1B
MMLU-Redux
44.42
43.98
OpenBookQA
47.80
47.40
Appendix
Table 12: C2C accuracies reported in the original paper versus our retraining under their released recipe, with Qwen3-0.6B as the receiver.
Benchmark
step 2054
step 4109
final
Llama → 0.5B
MMLU-R
Δ
−0.43
+0.11
−0.47
+1.07
CI
[−.80,−.07]
[−.34,+.55]
[−.91,−.03]
[+.50,+1.65]
ARC-E
Δ
−0.31
+0.35
−0.09
+2.28
CI
[−.88,+.26]
[−.22,+.92]
[−.70,+.52]
[+1.19,+3.38]
ARC-C
Δ
+0.26
−0.35
+0.42
+1.13
CI
[−.61,+1.13]
[−1.22,+.52]
[−.53,+1.38]
[−.17,+2.52]
Appendix
Table 13: Matched–Deranged differences (percentage points) with 95% per-question paired bootstrap intervals, for four configurations of C2C fusers we trained ourselves under the authors’ recipe. The first three columns are training checkpoints (step 2054, step 4109, final) of a Qwen2.5-0.5B-Instruct → Qwen3-0.6B fuser; the fourth is a Llama-3.2-3B-Instruct → Qwen2.5-0.5B-Instruct fuser. The Qwen2.5-0.5B-Instruct → Qwen3-0.6B row of Table 1 and Figure 1 (a) instead report the authors’ released fuser. Intervals that exclude zero are bolded; thirteen of the twenty contain zero.