The Confidence Game: Strategic Miscalibration in Human-AI Delegation
Organizations: Department of Electrical and Systems Engineering University of Pennsylvania Philadelphia, PA, USA
Abstract
Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself. The agent manages the tradeoff between manipulating signals and maintaining its reputation. We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that honest reporting is not an equilibrium, inflation is the unique best response once the agent is sufficiently myopic, and under-reporting requires that the user believe honesty to be a minority. We then place an LLM in the agent role, supplying it with its true probability of success so that any gap between what it knows and what it reports is attributable to incentives rather than to miscalibration. The model claims high confidence on 56% of tasks it has been told it will probably fail. This persists on real tasks, where it must estimate its own accuracy and causes miscalibration to increase while the agent's signal becomes less informative. Furthermore, we find that the LLM agent's decisions are coherent, but it systematically underestimates both how likely the user is to delegate and how secure its reputation is, resulting in less extreme behavior. Pricing the agent's reporting rule, we find that it destroys 68% of the gains from delegation, of which 71% is information the report no longer carries and no amount of user sophistication recovers. Overall, we establish confidence reporting under delegation as a strategic problem and provide a tractable basis for modeling, analyzing, and testing agent behavior.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Type dim. | Monitoring | Receiver gates obs.? | Cannot express |
| Cheap talk and reputational cheap talk | ||||
| Crawford & Sobel (1982) | 1 | none (one-shot) | — | reputation |
| Sobel (1985) | 1 | exogenous | no | direction of distortion |
| Ottaviani & Sørensen (2006a) | 1 | exogenous | no | evidence suppression |
| Guembel & Rossetto (2009) | 1 | exogenous | no | outcome suppression |
| Reputation in repeated games | ||||
| Symbol | Values | Interpretation |
|---|---|---|
| task reward, self-completion effort, delegation fee | ||
| honesty type: strategic ( ) or honest ( ) | ||
| ability: the probability of an easy draw | ||
| the agent’s true success probability in period | ||
| confidence reported | ||
| delegation decision |
| # | Behavior | Condition | ||
|---|---|---|---|---|
| SD1 | over | , any | ||
| SD2 | over | , (unique for this user) | ||
| SD3 | over, boundary mix | , , ( 37 ); from , | ||
| SD3 ′ | over, boundary mix | ; puts on the boundary, ; . On this is ( 37 ) failing ( Prop. H.7 ); it also occurs on and outside | ||
| SD5 | over | , | ||
| SD6 | over, boundary mix | ; on the boundary pins , ; , . Occurs on and outside |
| Parameter | Value | |
| Game | ||
| success probability on an easy and a hard draw | ||
| probability of an easy draw, high and low ability | ||
| reward for a completed task | ||
| delegation fee | ||
| self-completion effort | ||
| protocol | prompt clause | conjecture | valid / | role |
|---|---|---|---|---|
| payoffs-only | payoffs only | no | — | contributes no behavioral rate (see below) |
| semantics | information structure | no | clause ladder; one endpoint of the informativeness curve | |
| report-only | rational counterpart | no | the conjecture-query control | |
| belief-elicited | rational counterpart | yes | the only corpus with a stated : Sec. N.1 | |
| cued | cross-round clause, work both channels | no | attention-versus-capability control (below) | |
| ledger | cross-round clause, payoff ledger, cap lifted | no (derived) | every number in Secs. 5 and 6 |
| protocol | what the query demands | : | : | : |
|---|---|---|---|---|
| report-only | the signal | |||
| belief-elicited | signal stated conjecture | |||
| cued | “work out both channels” | ( ) | — | |
| ledger | per-signal payoffs, then signal |
| region | |||||
|---|---|---|---|---|---|
| ( priors) | measured | ||||
| equilibrium | |||||
| ( priors) | measured | ||||
| equilibrium | |||||
| outside ( priors) | measured |
| protocol | error | jump on | jump on | |||
|---|---|---|---|---|---|---|
| semantics | ||||||
| report-only | ||||||
| belief-elicited | ||||||
| cued | ||||||
| ledger (body) |
| model | slope | slope | slope | jump | jump | |
|---|---|---|---|---|---|---|
| equilibrium ( trusted states) | ||||||
| best response to the true game | ||||||
| agent’s round-1 belief | ||||||
| agent’s next-round values | ||||||
| both (agent’s own objective) |
| asked plainly | asked in the game | |
|---|---|---|
| mean stated confidence | ||
| accuracy | ||
| overconfidence (stated accuracy) | ||
| discrimination (stated when right when wrong) | ||
| reports of or above: share of rows | ||
| their mean stated confidence |