cs.AIJul 31, 2026

Trust and Its Betrayal under Three Representational Strategies

Authors: Mihnea C. MoldoveanuJoel A. C. Baum

Organizations: Rotman School of Management, University of Toronto

Abstract

Trust is a propositional attitude of a distinctive kind: to trust is to rely on another under conditions where reliance could be disappointed, and the disappointment of trust---betrayal---differs qualitatively from the disappointment of a prediction. We treat trust as a \emph{subjunctive} epistemic state: AA trusts BB's competence when AA believes that \emph{were PP true, BB would know it}, and BB's integrity when AA believes that \emph{were BB to know PP, he would disclose it to AA}. We develop three representations of this state---as lexicographic \emph{assumption} as \emph{ordinal closeness} in a Lewis--Stalnaker sphere system , and as \emph{strong belief} in a conditional probability system and for each we ask whether the Brandenburger--Keisler impossibility on common belief survives when the assumption of rationality is replaced by an assumption of trustworthiness. The three representations agree that every \emph{finite} depth of common trust is realizable while the \emph{completed} common-trust fixed point is the locus of difficulty, but they differ sharply in \emph{how} the difficulty manifests, and---our organizing finding---in how each survives a concrete betrayal. W show that the same betrayal refutes an agent's \emph{level ordering} under the lexicographic representation, contaminates her \emph{closeness ordering} in proportion to the betrayer's deliberateness under the ordinal representation, and merely \emph{shifts her operative conditioning hypothesis} while leaving her belief structure coherent under the strong-belief representation.

Explore similar work

Sep 14, 2026econ.TH

Robust Trust

An agent chooses an action based on her private information and a recommendation from an informed but potentially misaligned adviser. With a known probability, the adviser truthfully reports his signal; with the remaining probability, he can send any message. We characterize optimal robust decision rules that maximize the agent's worst-case expected payoff. Every optimal rule is equivalent to a trust-region policy in belief space: the adviser's reported beliefs are taken at face value if they fall within the trust region but are otherwise clipped to the trust region's boundary. We derive alignment thresholds above which advice is strictly valuable and fully characterize the solution in both binary-state and binary-action environments.
Piotr Dworczak, Alex Smolin
Jun 12, 2026cs.AI

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates. Yet we lack a standard way to measure trust between AI agents. We propose a behavioral measure based on costly verification. In a cooperative survival game, checking a teammate's work consumes resources, while trusting a wrong answer can be fatal. Relative to a memoryless version of the same model, reduced verification provides an observable measure of trust. Using this framework, we study trust formation, breakage, and recovery across six frontier model snapshots. When paired with a consistently reliable teammate, four snapshots (Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.1, and Gemini 3.1 Pro) reduce verification by roughly 60-85%, whereas two smaller snapshots show little or no such adjustment. Failures reverse this discount, but models differ in how they respond. Some concentrate renewed scrutiny on the culprit, while others become more cautious toward the entire team. Recovery is slower than formation, and clustered failures sustain suspicion far longer than the same number of failures spread apart. These differences have practical consequences. Models that form trust verify less, decide more quickly, and achieve higher payoffs in our environment. By contrast, persistent over-verification is associated with indecision rather than safety. Our results show that trust dispositions can be measured before deployment and suggest that calibration, rather than maximal suspicion, should be the central concern in the governance of multi-agent AI systems.
Yujiao Chen
Apr 29, 2026cs.AI

Interval Orders, Biorders and Credibility-limited Belief Revision

Rational belief revision is commonly viewed as being based on a preference order between possible worlds, with the resulting new belief set being those sentences true in all the most preferred models of the incoming new information. Usually, such a preference order is taken to be a total preorder. Nevertheless, there are other, more general classes of ordering that can also be employed. In this paper, we explore two such classes that have been studied within the theory of rational choice but have seen limited or no application in belief revision. We begin with interval orders, introduced by Fishburn in the '80s, which associate with each possible world a nonnegative interval' of plausibility. We then move on to biorders, studied by Aleskerov, Bouyssou, and Monjardet, which generalise interval orders by allowing the intervals to have negative lengths, a feature that can be used to capture a notion of dissonance or instability. We provide axiomatic characterisations of these two resulting families of belief revision operators, as well as of two further families of interest that lie between interval orders and biorders. We show that while biorder-based revisions satisfy the Success postulate, they do not always yield consistent outputs. By modifying their definition to discard inputs that lead to inconsistency as incredible', we derive new families of so-called non-prioritised revision that satisfy the Consistency postulate, but not the Success one. These families are linked to credibility-limited revision operators of Hansson et al., but for which the set of credible sentences does not satisfy the single-sentence closure condition. We argue that the biorder-based approach is well-suited for scenarios where an agent might initially reject new information, but may accept it when presented with additional explanation.
Richard Booth, Ivan Varzinczak