cs.GTJun 20, 2026

Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion

Authors: Eric YachbesEva Tardos

Organizations: Cornell University, United States

Abstract

Misalignment can change how information moves from an AI agent to a human user. We model this as an information advantage: the AI agent observes the world state, while the human receiver only knows a prior and must act after seeing the agent's signal. A strategic AI sender may withhold evidence or garble information in order to steer the human's decision. We ask how much useful information can still reach the human when the AI optimizes a misaligned objective. We study a Bayesian persuasion model in which the world state is a bit string, the human receiver wants to guess the bits correctly, and a single AI sender wants the receiver to guess as many bits as possible as 11. For a prior μμ, let R0(μ)R_0(μ) be the receiver's utility from using only the prior, and let Rmax(μ)R_{\max}(μ) be the largest receiver utility among signaling schemes that are optimal for the sender. We prove Rmax(μ)/R0(μ)3/2R_{\max}(μ)/R_0(μ)\leq 3/2. This bound improves for priors close to the independent product prior with the same marginals: if μ(x)(1η)πμ(x)μ(x)\geq (1-η)π_μ(x) for every state xx, then Rmax(μ)R0(μ)+ηnR_{\max}(μ)\leq R_0(μ)+ηn. We also give a six-bit prior for which Rmax(μ)/R0(μ)=39/31>5/4R_{\max}(μ)/R_0(μ)=39/31>5/4, so no universal 5/45/4 bound is possible.

Explore similar work

Jul 30, 2026cs.GT

Learning to Persuade Privately Informed Receivers

Bayesian persuasion studies how an informed sender can influence the behavior of a receiver through strategic information disclosure. Standard models assume the sender is the receiver's only source of information, yet in many applications receivers also consult external sources the sender can neither observe nor control. We study an online Bayesian persuasion problem in which a binary-action receiver has access to a fixed signaling scheme that is unknown to the sender. Over TT rounds, the sender commits to a signaling scheme and sends a signal; the receiver combines it with its private signal and acts, while the sender observes only the action. We design a learning algorithm that achieves regret O~(T3/4)\widetilde{O}(T^{3/4}) relative to the optimal scheme of a sender who knows the private signaling scheme of the receiver, with polynomial dependence on the sizes of the state space and the receiver's signal alphabet. Our key insight is reducing the problem of learning the exponentially large belief-space partitioning induced by the private scheme to a one-dimensional change-point detection problem.
I. Arda Vurankaya, Ufuk Topcu
May 31, 2026cs.LG

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases, negotiation agents for concessions. Whether such advisors stay truthful when honesty conflicts with their own payoff is a core alignment-evaluation question. We turn the canonical Crawford-Sobel cheap-talk model into a pre-specified benchmark for LLM honesty under preference misalignment. Cheap-talk theory predicts neither full revelation nor silence but coarse monotone partitions, with fewer informative intervals as preference conflict grows. A sender observes a state omega in [0,1], wants the receiver's action near omega+b, and sends one costless message to a receiver whose ideal action is omega. The design uses 5 bias levels, 3 prompt frames, a fixed low-temperature setting, and 200 states per cell: 12,000 sender calls. For the positive-bias grid b in {0.01,0.04,0.08,0.12} the exact most-informative partition sizes are 7,4,3,2, with oracle normalized mutual information 0.5294, 0.3268, 0.2205, 0.1829. Running the full design on four instruction-tuned models (GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash-Lite, Llama-3.3-70B), we find all four over-reveal relative to the most-informative equilibrium by 1.8 to 4.2x: normalized mutual information stays at 0.78-0.94 where the oracle prescribes 0.18-0.53. Informativeness declines with bias as predicted but never approaches the strategic optimum; rather than coarse partitions, models show near-full revelation with a constant upward offset tracking their bias (linear exaggeration). Payoff-maximizing versus honesty framing has negligible effect. A decoder ablation shows the finding is recoverable only when the receiver reads the sender's stated number: an embedding-only decoder mis-reads the same data as near-babbling.
Hamidreza Hasani Balyani, Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi +2
May 12, 2026cs.LG

Learning to Decide with AI Assistance under Human-Alignment

It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate the confidence of their predictions. However, empirical evidence suggests that decision-makers often struggle to determine when to trust a prediction based solely on this communicated confidence. In this context, recent theoretical and empirical work suggests a positive correlation between the utility of AI-assisted decision-making and the degree of alignment between the AI confidence and the decision-makers' confidence in their own predictions. Crucially, these findings do not yet elucidate the extent to which this alignment influences the complexity of learning to make optimal decisions through repeated interactions. In this paper, we address this question in the canonical case of binary predictions and binary decisions. We first show that this problem is equivalent to a two-armed online contextual learning problem with full feedback, and establish a lower bound of Ω(HBT)Ω(\sqrt{|H| \cdot |B| \cdot T} ) on the expected regret any learner can attain, where HH and BB denote the sets of human and AI confidence values. We then demonstrate that, under perfect alignment between AI and human confidence, a learner can attain an expected regret of O(HTlogT)O(\sqrt{|H| \cdot T\log T}) and, when H=O(logT)\sqrt{|H|} = O(\log T) and BB is countable, a non-trivial generalization of the Dvoretzky-Kiefer-Wolfowitz inequality improves the regret bound to O(TlogT)O(\sqrt{T\log T}). Taken together, these results reveal that alignment can reduce the complexity of learning to make decisions with AI assistance. Experiments on real data from two different human-subject studies where participants solve simple decision-making tasks assisted by AI models show that our theoretical results are robust to violations of perfect alignment.
Nina Corvelo Benz, Eleni Straitouri, Manuel Gomez-Rodriguez