Draft-KV: Learning Useful Latent Communication Between Language Models
Authors: Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, +2 more
Organizations: City University of Hong Kong · University of Science and Technology of China · University of Electronic Science and Technology of China · AIPD, Tencent · Theory Lab, Huawei · Hong Kong Metropolitan University
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
Figures & tables
Evaluation setting
Score (%)
Model Pair
Benchmark
Protocol
Receiver
Sharer
Text-to-Text
Cache-to-Cache
Draft-KV
Qwen3-0.6B → Qwen2.5-0.5B
MMLU-Redux
37.45
45.19
42.56
33.43 /33.40
46.11 /36.59
ARC-Easy
63.47
69.15
66.71
56.31 /53.11
74.28 /61.91
ARC-Challenge
40.10
51.71
49.49
39.51 /35.32
54.52 /39.76
OpenBookQA
43.40
44.60
45.60
41.00 /40.00
52.00 /41.20
C-EVAL
Public
41.75
41.38
41.46
37.74 /36.11
47.40 /40.56
Table 1: Main results across model pairs, benchmarks, and evaluation protocols. Public settings are scored by option accuracy, Private settings by exact match. Cache-to-Cache and Draft-KV entries show Matched/Deranged scores, where the Deranged message comes from a different question. Bold and underlined values denote the best and second-best score in each row. Model names omit the -Instruct suffix of Qwen2.5-0.5B-Instruct and Llama-3.2-3B-Instruct; Qwen3 names are as released (Appendix C.1 ).
ARC-Challenge
MMLU-Redux
Variant
Matched
Pairing gain
Matched
Pairing gain
Full Draft-KV
83.53
45.90
65.07
28.23
Prompt KV
49.83
0.17
43.59
0.09
w/o reconstruction pretraining
76.71
38.74
52.40
16.58
w/o answer alignment
80.03
41.72
55.40
19.25
Table 2: Message-source and training-stage ablations for Qwen3-1.7B → Qwen2.5-0.5B-Instruct under the Public protocol.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Mean gain
Mismatch harm AR−AD
Protocol
n
G
P
Mean
Median
Min
Max
Public
35
24.24
25.65
1.41
1.19
−3.48
11.65
Private
14
2.51
13.49
10.98
11.30
−5.51
19.91
Appendix
Table 3: Distribution of Draft-KV intervention contrasts across the settings of Table 1 . All quantities are in percentage points.
C2C
Draft-KV
Sharer → Receiver
AR
AM
Aadapter
ρadapter
AM
Aadapter
ρadapter
MMLU-Redux
Q3-0.6B → Q2.5-0.5B
37.45
33.43
31.92
95.48%
46.11
40.07
86.90%
Q3-1.7B → Q2.5-0.5B
37.45
34.16
34.43
100.79%
65.07
33.22
51.05%
Q3-4B → Q2.5-0.5B
37.45
34.57
37.07
107.23%
75.98
36.17
47.60%
Q2.5-0.5B → Q3-0.6B
30.63
42.92
41.23
96.06%
45.01
37.64
83.63%
Appendix
Table 4: Adapter-only comparison across three benchmarks and four model pairs. All accuracies and retained accuracies are percentages. Receiver-only and Matched follow Table 1 ; only adapter-only accuracies are added from the new evaluations. Q3 and Q2.5 denote Qwen3 and Qwen2.5-Instruct. Retained accuracy ρadapter=Aadapter/AM . Values above 100% mean adapter-only exceeds Matched.
Sharer → Receiver
Receiver-only
Adapter-only
Self-donor
Q3-0.6B → Q2.5-0.5B
36.13
39.65
37.89
Q3-1.7B → Q2.5-0.5B
36.13
33.59
60.55
Q3-4B → Q2.5-0.5B
36.13
34.96
64.45
Q2.5-0.5B → Q3-0.6B
37.50
41.02
42.19
Appendix
Table 5: Draft-KV self-donor diagnostic on the first 512 MMLU-Redux examples. All entries are accuracy (%). Self-donor and adapter-only use the same one-shot queries. All comparisons use this subset.
Sharer → Receiver
Layers ( ℓ←π(ℓ) )
CS
CR
HR
Projections
∣θ∣
Qwen3-0.6B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-1.7B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-4B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen3-8B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Llama-3.2-3B → Qwen2.5-0.5B
14,16,18,20 ← 18,20,22,24
1024
128
2
1,048,576
1,048,584
Qwen2.5-0.5B → Qwen3-0.6B
18,20,22,24 ← 14,16,18,20
128
1024
8
1,048,576
1,048,608
Appendix
Table 6: Interface configuration per model pair. Layer assignments are written receiver ← sharer. CS and CR are flattened KV widths, HR the number of receiver KV heads and hence of gates. Model names omit -Instruct suffixes as in Table 1 .
Stage 1
Stage 2
Sharer → Receiver
micro × accum
seed
micro × accum
seed
Qwen3-0.6B → Qwen2.5-0.5B
16×4
20260826
4×8
20260828
Qwen3-1.7B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Qwen3-4B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Qwen3-8B → Qwen2.5-0.5B
4×16
91827
4×8
91827
Llama-3.2-3B → Qwen2.5-0.5B
8×8
91827
4×8
91827
Appendix
Table 7: Per-pair microbatch × gradient accumulation and data seeds for Stages 1 and 2. Effective batches are 64 and 32 respectively for every pair. Model names omit -Instruct suffixes as in Table 1 .
Latency by stage (ms)
Total latency
Arithmetic
Condition
Sharer prefill
Draft decode
Extrac- tion
Projec- tion
Recv. prefill
Recv. decode
ms
×
TFLOPs
×
Interface GFLOPs
Public protocol, MMLU-Redux
Receiver-only
—
—
—
—
16.7
—
16.7
0.001
0.108
0.016
—
Sharer-only
33.6
11,868.1
—
—
—
—
11,907.2
1.000
6.801
1.000
—
Text-to-Text
33.6
11,868.1
—
—
18.0
—
11,925.2
1.002
7.187
1.057
—
Cache-to-Cache
45.1
—
—
35.6
20.7
—
101.4
0.009
2.256
0.332
118.3
Appendix
Table 8: Inference cost for one question with the Qwen3-8B sharer and the Qwen2.5-0.5B-Instruct receiver. Stage latencies are medians over 200 questions at batch size one after 20 warmup questions, timed with CUDA events. Floating-point counts are analytic; the interface column covers the projections of Equation 2 and the external attention of Equation 10 , and is reported in GFLOPs. Columns marked × give the ratio to Sharer-only on the same benchmark. Dashes mark stages a condition does not run.
Variant
AM
AD
AR
G
P
ARC-Challenge test
Full Draft-KV
83.53
37.63
40.10
43.43
45.90
Prompt KV
49.83
49.66
40.10
9.73
0.17
w/o reconstruction pretraining
76.71
37.97
40.10
36.61
38.74
w/o answer alignment
80.03
38.31
40.10
39.93
41.72
MMLU-Redux test
Appendix
Table 9: Complete accuracy contrasts for the core ablations on the two reported benchmarks. Accuracies are percentages; G and P are percentage points. Differences are computed from the reported two-decimal accuracies.
Variant
Benchmark
AM
AD
G
P
AR−AD
hˉ
ν
Full
ARC-C
83.53
37.63
43.43
45.90
2.47
0.038
19.7 (231)
w/o Guard
ARC-C
82.94
35.92
42.84
47.02
4.18
0.091
37.9 (444)
Full
MMLU-R
65.07
36.84
27.62
28.23
0.61
0.052
24.6 (1386)
w/o Guard
MMLU-R
64.68
36.24
27.23
28.44
1.21
0.104
42.8 (2411)
Appendix
Table 10: Guard ablation. Accuracies are percentages; G , P , and AR−AD are in pp. hˉ is the mean hinge and ν the threshold-exceedance rate, with the number of affected examples in parentheses. Receiver-only baselines are 40.10% (ARC-C) and 37.45% (MMLU-R).
Sharer
MMLU-R
ARC-E
ARC-C
OBQA
C-EVAL
Mean
G (pp)
Qwen3-0.6B
85.3±88.1
49.5±33.8
58.5±41.9
24.4±27.0
73.2±93.0
58.2
9.63
Qwen3-1.7B
293.1±112.9
219.9±73.0
243.4±81.7
220.7±66.3
332.1±120.0
261.8
30.57
Qwen3-4B
263.5±108.4
198.2±59.5
212.9±69.4
199.8±58.7
242.2±128.0
223.3
39.14
Qwen3-8B
358.9±102.3
263.6±67.0
285.3±76.7
251.5±60.4
355.8±113.3
303.0
42.25
Appendix
Table 11: Mean draft length in decoded tokens (mean ± SD) for the four scaling points, all decoded greedily under a 512-token cap. The last column is the unweighted mean over the five benchmarks, shown against the mean system gain of Figure 5 .
Sharer
Benchmark
Reported
Retrained
Qwen2.5-0.5B-Instruct
MMLU-Redux
42.92
42.95
OpenBookQA
52.60
52.80
ARC-Challenge
54.52
54.18
C-EVAL
41.77
41.68
Llama-3.2-1B
MMLU-Redux
44.42
43.98
OpenBookQA
47.80
47.40
Appendix
Table 12: C2C accuracies reported in the original paper versus our retraining under their released recipe, with Qwen3-0.6B as the receiver.
Benchmark
step 2054
step 4109
final
Llama → 0.5B
MMLU-R
Δ
−0.43
+0.11
−0.47
+1.07
CI
[−.80,−.07]
[−.34,+.55]
[−.91,−.03]
[+.50,+1.65]
ARC-E
Δ
−0.31
+0.35
−0.09
+2.28
CI
[−.88,+.26]
[−.22,+.92]
[−.70,+.52]
[+1.19,+3.38]
ARC-C
Δ
+0.26
−0.35
+0.42
+1.13
CI
[−.61,+1.13]
[−1.22,+.52]
[−.53,+1.38]
[−.17,+2.52]
Appendix
Table 13: Matched–Deranged differences (percentage points) with 95% per-question paired bootstrap intervals, for four configurations of C2C fusers we trained ourselves under the authors’ recipe. The first three columns are training checkpoints (step 2054, step 4109, final) of a Qwen2.5-0.5B-Instruct → Qwen3-0.6B fuser; the fourth is a Llama-3.2-3B-Instruct → Qwen2.5-0.5B-Instruct fuser. The Qwen2.5-0.5B-Instruct → Qwen3-0.6B row of Table 1 and Figure 1 (a) instead report the authors’ released fuser. Intervals that exclude zero are bolded; thirteen of the twenty contain zero.
LLM agents today communicate via text, which incurs considerable latency and information loss due to the need to autoregressively decode the sharer model's state and encode at the receiver model. Recent work such as Cache-to-Cache (C2C; Fu et al., 2026) seeks to exchange KV caches by learning adapters that translate sharer KV matrices to the receiver model. However, the adapters are large and expensive to train, and translate individual tokens, which requires the target context to be identical. This is unsuitable for agent communication, where the LLMs have differing context. We introduce Latent Cache Flow (LCF). To address efficiency, we observe that keys and values can be jointly translated and compressed, reducing the adapter to about 4% of C2C's size. To address differing context, we design the adapter to transmit a summary of new information that the target model does not have. Our early experiments show that a pruned 13 MB LCF adapter can be more accurate than C2C at 956 MB in shared-context settings; for different contexts, LCF improves F1 by 7.5% and Exact Match by 23% while 8.5 times faster than text-based communication.
Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
SiYuan Ma, Yiqin Luo, Zhangji +8
1Nanyang Technological University · 2Southern University of Science and Technology · 3Tianjin University +7
In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interface state helps or harms later learning. First, a leakage-controlled causal interoperability audit over all 30 ordered pairs of six independently trained restricted societies -- under sealed held-out structure and a preregistered raw/orthogonal/linear/nonlinear alignment ladder -- shows the six semantically similar interfaces do not form one raw language: one same-initialization pair is exactly interoperable in both directions, a second shows asymmetric partial compatibility, and all 26 cross-initialization directions fail every frozen alignment rung. Second, within the tested decomposition and a single sealed source formulation, a source-span control localizes strict zero-shot failure to interpretation and execution of the new operator instructions. Third, in a matched adaptation factorial, the globally trained communication interface acts as a severe negative-transfer prior: reinitializing only the packet reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857. Fourth, across two restricted checkpoints and two independently frozen target streams each, inherited interfaces never exceeded fresh-interface controls by the preregistered 0.10 margin. All primary conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial concerns one globally visible parent-cohort checkpoint, while an appendix adds a post hoc tagged-global twin case study.