Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.
Figures & tables
Figure 1: Framework overview. Top: receivers read one message as different tasks (1), PIR separates interpretation from capability risk (2), and the sender infers the receiver and queries only if worthwhile (3). Bottom: offline probes train the PIR model, which picks the verified rewrite to send.
Figure 2: Interpretation failure is separate from execution failure and depends on the receiver (test split). (a) YA after a correct read ( ∘ ) and a misread ( ∙ ) of the probe. (b) YI per receiver. Bands span the lowest to the highest receiver.
Per receiver ↑
Across receivers ↑
Calibration ↓
Method
AUROC
AUPRC
AUROC
AUPRC
Brier
NLL
ECE
A0 Global prior
.500
.045
.500
.045
.0428
.180
.029
A0 Receiver prior
.500
.045
.697
.086
.0415
.168
.002
A1 Message judge
.552
.057
.540
.055
.0428
.180
.031
A2 Equiv. judge
.500
.045
.500
.045
.0428
.180
.029
A4 YA predictor
.511
.048
.530
.050
.4229
1.329
.540
Table 1: Prospective PIR prediction (test split, six-dataset means, seed 0). Ours is best on all seven metrics (bold, second best underlined, A0 excluded).
YI↓
ΔYI (ours − row)
YA↓
Condition
Gen.
Sel.
(%)
points [95% CI]
(%)
Original message
–
–
4.16
−1.85 [ −2.09 , −1.62 ] ‡
59.33
Generic rewrite
–
–
3.85
−1.54 [ −1.78 , −1.30 ] ‡
58.15
Selection only
–
✓
3.47
−1.15 [ −1.41 , −0.93 ] ‡
58.20
Posterior only
πt
✓
3.43
−1.11 [ −1.33 , −0.90 ] ‡
58.59
Guidance only
✓
–
2.29
+0.03 [ +0.01 , +0.05 ]
58.33
Table 2: Message revision on 1,579 held-out questions: ours beats every pre-registered reference on YI ( ‡ Holm p<0.001 ). Gen.: learned risk guides generation ( πt : posterior only). Sel.: it selects the sent message.
Primary cohort
Replication
YI
Gap
Net
Ident.
YI
Gap
Policy (query rate)
(%)
(%)
utility
(%)
(%)
(%)
No query (0%)
3.84
0
.9616
24.6
3.93
0
Exact 20% quota
Random (3 seeds)
3.84
3
.9614
25.7
3.93
4
IG
3.83
5
.9615
28.5
3.94
<0
Table 3: Query policies (interpretation objective, c=0.001 ). Gap: share of the no-query to true-identity YI gap that is closed. Ident.: posterior mode is correct. IG: information gain. † Secondary comparator. Bold: best deployable value.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Defined in
General
E , P (or P ), Var
Expectation, probability, and variance
–
1{⋅}
Indicator: 1 if the condition holds and 0 otherwise
–
H(⋅) , I(⋅;⋅∣⋅)
Shannon entropy and conditional mutual information, in nats
–
DKL(⋅∥⋅)
Kullback–Leibler divergence
–
Ber(p)
Bernoulli distribution with mean p
–
Appendix
Table 4: Notation, grouped by the part of the framework that introduces each symbol. The last column gives where the symbol is defined, and – marks standard notation.
Receiver
First
Second
Last
All
First or second [95% CI]
Qwen3.6
1.28
1.61
2.33
1.74
1.45 [1.26, 1.65]
Gemma
1.59
2.49
2.57
2.22
2.04 [1.81, 2.28]
gpt-oss
3.27
2.08
1.63
2.33
2.68 [2.42, 2.95]
Nemotron
3.14
4.51
24.49
10.72
3.83 [3.57, 4.09]
Mistral
5.26
5.74
4.45
5.16
5.50 [5.15, 5.85]
Highest/lowest
4.1
3.6
15.0
6.2
3.8 [3.4, 4.3]
Appendix
Table 5: Probe error rate (%) by the position of the intended task (test split, parsed probes). Last column: probes with the intended task first or second, with question-cluster bootstrap 95% CIs.
ID
Method
Definition
A0
Base-rate priors
Global P(YI=1) and receiver-specific P(YI=1∣R) estimated on training data.
A1
Message-only ambiguity
Fixed strong judge sees m only and scores whether materially different task interpretations are possible.
A2
Semantic equivalence
Judge sees (m,z⋆,xS) and predicts whether m faithfully specifies the intended task, with no receiver information.
A3
Receiver-agnostic PIR
Same learned PIR architecture but remove eR / receiver posterior information.
A4
Action-failure predictor
Same architecture/input budget trained against YA rather than YI , evaluated against both targets.
A5
Listener/ToM
Predict P(ZR=z∣m,x,eR) over probe task identities and use 1−P(ZR=z⋆) as risk.
Appendix
Table 6: Baseline definitions.
Figure 3: Interpretation failure depends on the receiver and the dataset (test split). (a) YI (%) per receiver and dataset on a log colour scale. Boxes mark the least risky receiver in each dataset, and the bottom row gives the ratio of the highest to the lowest receiver. (b) Misread pairs split by the answer check. Labels give the share of misreads that pass.
Marginal rates (%)
Off-diagonal cells (%)
Dataset
Pairs
YI [95% CI]
YA [95% CI]
misread, pass
read, fail
FreebaseQA
6,550
2.5 [1.7, 3.4]
24.0 [21.7, 26.4]
1.8
23.3
WebQSP
1,095
5.5 [2.5, 8.6]
54.1 [47.5, 60.7]
1.8
50.3
GrailQA
6,430
7.9 [6.5, 9.4]
89.6 [87.9, 91.3]
0.8
82.4
SQ-WD
7,090
5.2 [4.1, 6.4]
61.9 [59.3, 64.4]
2.2
58.8
PopQA
7,255
3.3 [2.4, 4.2]
75.7 [73.5, 77.9]
0.8
73.2
Appendix
Table 7: Interpretation ( YI ) and execution ( YA ) failures per dataset (test split, pooled over five receivers, 95% question-level normal intervals).
Per receiver ↑
Across receivers ↑
Calibration ↓
Method
AUROC
AUPRC
AUROC
AUPRC
Brier
NLL
ECE
A3 (no receiver)
.729
.128
.701
.115
.0414
.168
.032
A3 + recv. intercept
.729
.128
.792
.158
.0401
.156
.015
A3 + recv. calib.
.728
.128
.796
.158
.0399
.155
.012
A3 + recv. map ∗
.728
.128
.797
.160
.0398
.154
.010
Ours
.732
.135
.800
.164
.0398
.154
.010
Appendix
Table 8: Receiver levels or interaction (Experiment 2, test split, six-dataset means, seed 0). Maps are refitted on the validation questions used by the full model ( ∗ on all validation data). Difference: ours minus A3 + recv. calib., with paired question-group bootstrap 95% CIs.
Figure 4: Receiver information calibrates the predicted risk (test split, seed 0 in (a, b), three-seed means in (c)). (a) Within-receiver ECE per dataset. (b) Mean predicted versus measured YI in each of the 30 receiver–dataset cells. (c) AUROC across versus within receivers, with arrows from A3 to ours.
Dataset
Method
AUROC [95% CI]
AUPRC
Brier ↓
NLL ↓
ECE ↓
FreebaseQA
A0 global base rate
0.500
0.025
0.0246
0.118
0.024
A0 per-receiver base rate
0.500
0.025
0.0238
0.105
0.003
A1 message-only judge
0.529 [0.481, 0.578]
0.029
0.0246
0.118
0.024
A2 equivalence judge
0.500
0.025
0.0246
0.118
0.024
A3 receiver-agnostic PIR
0.690 [0.645, 0.730]
0.085
0.0243
0.113
0.024
A4 action-failure predictor
0.558 [0.518, 0.599]
0.032
0.115
0.377
0.234
Appendix
Table 9: Experiment 2 results per dataset (test split, seed 0, with metrics computed within each receiver and macro-averaged). AUROC has a paired 95% bootstrap CI. Bold: best value per dataset among learned methods.
Variant
AUROC (pooled) ↑
AUROC (macro) ↑
Brier ↓
NLL ↓
ECE (macro) ↓
FreebaseQA
Full model (with eR )
0.810
0.698
0.0234
0.100
0.005
Receiver removed (A3)
0.663
0.691
0.0243
0.113
0.024
Receiver shuffled
0.661
0.691
0.0243
0.113
0.024
Δ full − A3
+ 0.147
+ 0.008
− 0.0009
− 0.013
95% CI
[0.124, 0.170]
[ − 0.0012, − 0.0007]
[ − 0.015, − 0.011]
Appendix
Table 10: Receiver-information ablation (test split, three-seed means). Δ : full minus receiver-agnostic, with paired 95% CIs in the row below. Last row per dataset: A4 scored on YI , with its AUROC on YA in parentheses.
Figure 5: Behavioural history sharpens the receiver belief (960 audit histories per length). (a) Accuracy and (b) calibrated NLL by history length for genuine and shuffled history, with binomial 95% CIs in (a). (c) One-vs-rest AUROC at ∣Ht∣=5 against each receiver’s pooled YI , with 95% bootstrap CIs.
Calibrated, genuine history
Raw
Shuffled history
∣Ht∣
Accuracy ↑
NLL ↓
Brier ↓
ECE ↓
NLL ↓
Accuracy
NLL
0
20.0%
1.609
0.800
0.000
–
–
–
1
22.6%
1.634
0.811
0.056
1.626
20.9%
1.625
3
23.5%
1.604
0.798
0.013
1.603
18.4%
1.633
5
24.9%
1.593
0.794
0.014
1.594
21.3%
1.617
10
26.6%
1.591
0.793
0.031
1.593
19.9%
1.628
Appendix
Table 11: Receiver-type inference by history length ∣Ht∣ (960 audit histories per length). Length 0 is the uniform prior over five receivers, and “calibrated” uses temperature scaling fitted on the calibration split.
Dataset
Accuracy ↑
NLL ↓
Brier ↓
ECE ↓
FreebaseQA
21.5%
1.604
0.798
0.025
WebQSP
24.6%
1.598
0.796
0.028
GrailQA
20.5%
1.608
0.799
0.040
SQ-WD
18.3%
1.610
0.801
0.051
PopQA
25.9%
1.583
0.790
0.008
Mintaka
25.9%
1.588
0.792
0.017
Appendix
Table 12: Receiver-type inference at ∣Ht∣=5 : by dataset of the calibration tasks (top, 960 histories each) and row-normalised confusion matrix (bottom, %, true receiver in rows).
Figure 6: Risk-guided revision lowers measured interpretation failure on the 1,579 held-out questions. Paired differences in YI , ours minus each reference (points, 95% CIs, negative favours ours). (b) Per dataset against the original (filled) and the generic rewrite (hollow), post hoc.
Figure 7: One risk feature is shared by all receivers, and guidance removes it. (a) Learned scores srk and intercepts br (points of YI ). Boxed: the score and its 2.5% bootstrap quantile are both positive, the guidance rule for a known receiver. (b) YI per receiver. (c) Form of all sent messages (unchanged originals included) and YI per condition, with guided conditions in blue.
Condition
YI
YA
None
Distr.
Changed
Val. YI
Original message
4.16
59.33
0.93
3.20
0.0
5.01
Generic rewrite
3.85
58.15
0.88
2.95
97.2
4.97
Selection only
3.47
58.20
0.86
2.59
86.7
4.60
Posterior only
3.43
58.59
0.81
2.60
86.1
4.23
Ours
2.31
58.33
0.72
1.56
90.0
2.83
Guidance only
2.29
58.33
0.70
1.54
91.5
2.66
Appendix
Table 13: Experiment 4 results by condition. YI , YA , None , and Distr. (probes that choose the contrast task) are percentages on the 1,579 held-out questions, and Changed is the share of messages rewritten. Val. YI is on the 300 validation questions of the check before the main run.
YI (%)
Ours minus
Generic minus
Receiver
Original
Ours
original [95% CI]
generic
sel. only
post. only
original
Value
Nemotron
8.73
3.76
−4.96 [ −5.34 , −4.59 ]
−3.54
−2.70
−2.66
−1.42
−5.60
Gemma
2.84
2.16
−0.69 [ −1.04 , −0.35 ]
−0.88
−0.55
−0.42
+0.20
+0.07
gpt-oss
1.91
1.51
−0.40 [ −0.62 , −0.17 ]
−0.56
−0.63
−0.52
+0.16
+0.01
Mistral
5.27
2.65
−2.62 [ −3.00 , −2.25 ]
−2.04
−1.50
−1.60
−0.58
−4.09
Qwen
2.07
1.50
−0.57 [ −0.85 , −0.31 ]
−0.66
−0.40
−0.36
+0.09
−0.38
Appendix
Table 14: Experiment 4 results per receiver on the held-out questions. Ours minus each reference and the generic rewrite minus the original are paired differences in YI (points), and Value is ours minus the original on the 275 value-versus-attribute questions.
Figure 8: NetVoII allocates queries by decision value (primary cohort, interpretation objective, c=0.001 ). (a) YI against the share of episodes queried. (b) Paired differences at matched quotas with 95% CIs, equal to the net-utility contrasts with the sign reversed. (c) Share of episodes whose posterior mode is the true receiver. (d) Strict rule at six query costs c : measured YI (top) and share of episodes queried (bottom).
Figure 9: Information is not value (primary cohort). (a) Expected IG of the most informative query (nats) and share of episodes with positive predicted VoII, per dataset. (b) Share of episodes whose best query is worth more than a cost c , with pooled values marked at four of the tested costs.
Policy
10%
20%
30%
50%
Always
Primary cohort (test split). No query: 3.84 (0.9616), true identity: 3.77 (0.9623), adapted VoI † without queries: 4.42 (0.9558).
Random (3-seed mean)
3.84 (0.9615)
3.84 (0.9614)
3.83 (0.9614)
3.83 (0.9612)
3.84 (0.9606)
IG
3.83 (0.9616)
3.83 (0.9615)
3.83 (0.9614)
3.82 (0.9613)
3.80 (0.9610)
Adapted VoI †
4.42 (0.9557)
4.42 (0.9556)
4.42 (0.9555)
4.42 (0.9553)
–
NetVoII (ours)
3.80 (0.9619)
3.79 (0.9619)
3.78 (0.9619)
3.78 (0.9617)
–
Replication cohort. No query: 3.93 (0.9607), true identity: 3.91 (0.9609), adapted VoI † without queries: 4.27 (0.9573).
Appendix
Table 15: Experiment 5 budget grid (interpretation objective): measured YI (%) by exact query quota, with net utility at c=0.001 in parentheses. Bold: net-utility contrasts against both IG and random querying survive Holm correction. † Secondary comparator: its estimates never let a query change the message.
Cost c
Queried (%)
YI (%)
1−YA (%)
Net utility
Primary cohort (test split)
0
32.9
3.78
35.66
0.9622
0.0001
30.2
3.78
35.68
0.9622
0.0005
24.0
3.79
35.62
0.9620
0.001 (main)
18.2
3.79
35.68
0.9619
0.005
0.8
3.83
35.72
0.9617
Appendix
Table 16: Strict rule of Experiment 5 across query costs c (interpretation objective, net utility at each row’s cost). The last two rows of each block re-run IG and random querying at the query count of the c=0.001 rule.
YI at the 20% quota
Queried by NetVoII (%)
Unit
No query
NetVoII
Inf. gain
Random
20% quota
Strict
YI strict
True identity
FreebaseQA
2.16
2.16
2.15
2.16
1.5
1.1
2.16
2.10
WebQSP
5.07
5.07
5.07
5.07
0.0
0.0
5.07
5.14
GrailQA
6.96
6.96
6.96
6.96
0.0
0.0
6.96
6.96
SQ-WD
4.53
4.51
4.50
4.53
4.9
4.1
4.51
4.45
PopQA
2.61
2.42
2.62
2.60
49.8
45.1
2.42
2.52
Appendix
Table 17: Experiment 5 results per dataset and per receiver (primary cohort, interpretation objective): YI (%) at the pooled 20% quota and under the strict rule, with the share of episodes that NetVoII queries.
Figure 10: NetVoII spends a pooled quota where predicted value is high (primary cohort, interpretation objective). (a) Share of each dataset’s episodes queried at pooled quotas of 10–50% and under the strict rule. (b) Change in YI against no query at the 20% quota.
Faculty of Dental Medicine and Oral Health Sciences, McGill University, Montréal, QC, Canada · Chandar Research Lab, Polytechnique Montréal, Montréal, H3T 1J4, Canada · Mila – Québec Artificial Intelligence Institute, Montréal, H2S 3H1, Canada +2