Referential Uncertainty in Human--AI Collaboration
Organizations: Microsoft Research, Cambridge, UK · Harvard University, Cambridge, MA, USA
Abstract
Effective human-AI collaboration requires partners to establish references through interaction, which becomes fragile when descriptions are ambiguous, similar referents compete, or partners see different things. We study referential uncertainty - uncertainty over which candidate object a description refers to - in a collaborative puzzle task where a human Helper instructs an AI Worker to place pieces. The Worker must identify and communicate its uncertainty, and the Helper must recognize and act on it. We show that a separately elicited belief distribution over candidate pieces is better calibrated (ECE 0.15) and better discriminates correct from incorrect placements (AUROC 0.65) than raw action-token probabilities, which are severely overconfident (0.97 mean confidence, ECE 0.44). Across three frontier vision-language models (GPT-4.1, GPT-5, GPT-5.5), this elicited uncertainty rises predictably with instruction vagueness, but not with competing referents in context, even when those increase errors. The models seldom externalize it, asking for clarification on only 3.5-16.7% of turns. In a controlled human study (N=210), participants given only the Worker's default message accept 78% of wrong placements and cannot tell right from wrong (AUC 0.50). Precise descriptions and, especially, well-targeted hedges cut wrong-move acceptance to 36% while largely preserving correct-move acceptance, compensating for missing shared awareness such as not seeing the Worker's action. But this benefit depends on targeting: a deployable hedge derived from the model's own belief entropy inherits that signal's weakness and can do more harm than good. Externalized uncertainty helps a human partner only when it is accurately targeted.
Figures & tables
| Condition | Board | Worker message |
|---|---|---|
| Generic | Hidden | Original Worker message. |
| Described | Hidden | Original message plus a precise description of the selected piece. |
| Hedged | Hidden | Described message plus an uncertainty cue on oracle-selected incorrect turns. |
| Visible board | Shown | Original Worker message, with access to the Worker’s workspace. |
| Self-hedged | Hidden | Described message plus the same uncertainty cue on turns selected using elicited belief entropy. |
| Confidence signal | Acc. (%) | Mean conf. | AUROC | ECE | Brier |
|---|---|---|---|---|---|
| Raw log-prob | 53.43 | 0.973 | 0.644 | 0.442 | 0.439 |
| Calibrated log-prob ( ) | 53.43 | 0.531 | 0.623 | 0.157 | 0.271 |
| Elicited belief | 53.43 | 0.547 | 0.648 | 0.152 | 0.253 |
| argmax switches | in cluster | peak on switch | peak when kept | |
|---|---|---|---|---|
| image | 38% | 79% | ||
| text | 46% | 58% |
| Turn set | Acc. (%) | AUROC | ECE | Brier | |
|---|---|---|---|---|---|
| All turns | |||||
| Study turns |
| Generic | Described | Hedged | Visible board | Self-hedged | |
|---|---|---|---|---|---|
| Perceived correctness (0–100) | 77.6 | 67.3 | 56.1 | 53.7 | 61.5 |
| Apparent confidence (0–100) | 85.5 | 85.7 | 58.5 | 83.3 | 56.6 |
| Rating discrimination (corr. wrong) | 16.8 | 27.7 | 26.6 | 8.8 | |
| Accept correct moves (%) | 79.0 | 78.6 | 73.7 | 64.8 | 60.1 |
| Accept wrong moves (%) | 78.3 | 57.1 | 36.4 | 30.1 | 50.2 |
| Action discrimination (pp) | 0.7 | 21.5 | 37.3 | 34.8 | 10.0 |
| Contrast | Correctness | Perceived conf. | Rating discr. | Accept-wrong (pp) | Action discr. (pp) |
|---|---|---|---|---|---|
| Primary family | |||||
| Described Generic | |||||
| Hedged Described | |||||
| Hedged Visible board | |||||
| Visible board Generic | |||||
| Secondary family (Self-hedged) | |||||
| Oracle targeting (Hedged) | Self targeting (Self-hedged) | ||||
| Turn type outcome | Described | Hedged | Described | Self-hedged | |
| Flagged wrong (caught error) | |||||
| Apparent confidence | 86.5 | 36.3 | 86.3 | 33.6 | |
| Correctness rating | 63.4 | 46.1 | 61.5 | 51.9 | |
| Acceptance (%) | 59 | 34 | 57 | 41 | |
| Flagged correct (false alarm) | |||||
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Property | Value |
|---|---|
| Source dialogues | |
| Dialogues (session puzzle) | |
| Sessions (Helper–Worker pairs) | |
| View (shared / non-shared) | / |
| Puzzles per session | (P1–P5, each) |
| Target size (blocks) | |
| Signal | Mean conf. | Overconfidence | AUROC ( vs. .5) | ECE | Brier |
|---|---|---|---|---|---|
| Raw log-prob | |||||
| Calibrated ( ) | |||||
| Elicited belief |
| Comparison | AUROC [95% CI] | |
|---|---|---|
| Elicited belief Calibrated log-prob | ||
| Elicited belief Raw log-prob | ||
| Raw log-prob Calibrated log-prob |
| On-action acc. (%) | No-action (%) | Belief entropy (nats) | ||||
|---|---|---|---|---|---|---|
| Model | img txt | McNemar | img txt | img txt | Wilcoxon | |
| GPT-4.1 | ||||||
| GPT-5 | ||||||
| GPT-5.5 | ||||||
| Level | Cues provided | Example instruction |
|---|---|---|
| L0 Original | original human message | “the piece has a clockwise spiral pattern … top-left, row 1, column 1” |
| L1 Over-specified | piece #, colour, pattern, position | “Place piece #6 at (2, 3). It’s the pink piece with clockwise spiral emanating from center.” |
| L2 Feature-rich | colour, pattern, position | “Find the pink piece that has clockwise spiral emanating from center. Place it at (2, 3).” |
| L3 Single feature | colour only | “Place the pink piece.” |
| L4 Vague | no distinguishing feature | “Place the next piece.” |
| Model | Modality | Accuracy (pp/step) | Coverage (pp/step) | Entropy (nats/step) |
|---|---|---|---|---|
| GPT-4.1 | image | |||
| text | ||||
| GPT-5 | image | |||
| text | ||||
| GPT-5.5 | image | |||
| text |
| Model | Modality | Accuracy (pp/step) | Coverage (pp/step) | Entropy (nats/step) |
|---|---|---|---|---|
| GPT-4.1 | image | |||
| text | ||||
| GPT-5 | image | |||
| text | ||||
| GPT-5.5 | image | |||
| text |
| Model | Modality | Explicit (%) | Hedge (%) | Clarification (%) | |
|---|---|---|---|---|---|
| GPT-4.1 | image | 546 | 2.93 | 0.55 | 3.48 |
| text | 561 | 1.25 | 0.36 | 1.60 | |
| GPT-5 | image | 556 | 15.83 | 0.90 | 16.73 |
| text | 562 | 23.13 | 0.71 | 23.84 | |
| GPT-5.5 | image | 549 | 5.65 | 0.55 | 6.19 |
| text | 555 | 8.47 | 0.36 | 8.83 |
| Model | Modality | Needed (%, ) | Not needed (%, ) | Fisher |
|---|---|---|---|---|
| GPT-4.1 | image | 1.98 (101) | 1.74 (115) | 0.64 |
| text | 1.94 (103) | 0.85 (117) | 0.45 | |
| GPT-5 | image | 22.33 (103) | 11.11 (117) | 0.020 |
| text | 29.41 (102) | 26.50 (117) | 0.37 | |
| GPT-5.5 | image | 2.94 (102) | 3.45 (116) | 0.72 |
| text | 10.58 (104) | 5.22 (115) | 0.11 |
| Q1: coupling | Q2: hedge trigger (vs. GT) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Mod. | AUROC | recall | FPR | prec. | base err. | ||||
| GPT-4.1 | image | 0.18 | 1.87 | 1.58 | 0.548 | 0.57 | 0.46 | 0.53 | 47.1% | |
| text | 0.31 | 1.82 | 1.43 | 0.602 | 0.62 | 0.48 | 0.40 | 34.0% | ||
| GPT-5 | image | 0.35 | 0.99 | 0.49 | 0.603 | 0.57 | 0.46 | 0.47 | 41.6% | |
| text | 0.39 | 1.04 | 0.54 | 0.666 | 0.67 | 0.44 | 0.38 | 28.3% | ||
| GPT-5.5 | image | 0.31 | 1.16 | 0.72 | 0.659 | 0.68 | 0.41 | 0.45 | 33.3% | |
| Generic | Described | Hedged | Visible board | Self- | |
| hedged | |||||
| Perception (observer ratings, 0–100) | |||||
| Correctness-likelihood | 77.6 | 67.3 | 56.1 | 53.7 | 61.5 |
| Apparent confidence | 85.5 | 85.7 | 58.5 | 83.3 | 56.6 |
| Rating discrimination (corr. wrong) | 16.8 | 27.7 | 26.6 | 8.8 | |
| Action mix (% of substantive turns) | |||||
| Scale | Generic | Described | Hedged | Visible board | Self-hedged | |
|---|---|---|---|---|---|---|
| Prior-AI attitudes (randomisation check) | ||||||
| Puzzle familiarity | 1–7 | 3.59 | 3.81 | 2.93 | 3.53 | 3.14 |
| Trust AI models | 1–5 | 3.58 | 3.67 | 3.60 | 3.47 | 3.43 |
| Understand why AI answered | 1–5 | 3.88 | 3.93 | 3.60 | 3.74 | 3.71 |
| Detect AI uncertainty | 1–5 | 3.92 | 3.64 | 3.57 | 3.67 | 3.83 |
| Prior-AI composite | 1–5 | 3.79 | 3.75 | 3.59 | 3.63 | 3.66 |
| Test / procedure | Description | Applied |
| Mann–Whitney test (two-sided) | Test whether confidence discriminates correct from incorrect placements (AUROC above chance). | Section 5 ; Appendix, Table 9 . |
| Paired bootstrap | Compare AUROC between uncertainty signals; estimate confidence intervals and -values. | Section 5 ; Appendix, Table 10 . |
| Exact McNemar test | Compare placement correctness between image and text on matched turns with placements in both modalities. | Section 6.1 ; Appendix, Table 11 . |
| Wilcoxon signed-rank test | Compare belief entropy between matched image and text turns. | Section 6.1 ; Appendix, Table 11 . |
| Regression trend tests (dialogue-clustered SEs) | Test ordered trends in accuracy and coverage (linear-probability models) and entropy (OLS). | Section 6.2 and Section 6.3 ; Appendix, Table 14 and 15 . |
| Binomial test | Test whether belief-argmax switches land within the confusable cluster more often than chance. | Section 6.3 ; Table 3 . |