cs.LGSep 27, 2026

Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces

Authors: Han Chen, Yingrui Li

Organizations: Independent Researcher

Abstract

Equivalent probability requests can lead to different decisions even when both reports are valid. We introduce probability contracts, a benchmark that connects exact finite-world posteriors, validated event alignment, interface coherence, and failure-aware decision evaluation. Across four model-interface configurations on 1,000 worlds, Kev has lower aggregate canonical posterior error than Jev but larger complement and coarsening residuals; accuracy ordering varies by stratum. Jev's Event and Choice interfaces change the binary action on 32.8% of valid pairs at defer cost 0.10. Post-hoc analyses show that disagreement certifies only 11-52% of mean binary pair error and does not consistently outperform confidence for selection. An action-region characterization and a standard scoring-rule identity explain averaging's expected Brier guarantee relative to random interface selection, but not a decision-loss guarantee at each cost. The loss contrast takes both signs on a 99-cost grid for every configuration; small penalties where both policies beat deferral have pointwise intervals containing zero. Secondary checks specified before collection include a separate 400-root cohort, where Event/Choice effects remain configuration-dependent. A joint surface-order and answer-ID intervention shifts posttrained probabilities. Both bounded reasoning arms yield no valid probability reports, leaving their probability accuracy undefined. Probability contracts make these distinctions measurable by evaluating event semantics, posterior error, coverage, and decision cost together.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do System One Decisions Add Up? A Study of Probabilistic Coherence

    Sep 27, 2026Saman Sarker JoyConfidence CalibrationDecisions

  2. Diagnosing and Improving Probabilistic Reasoning in Large Language Models

    Sep 29, 2026Huaman Sun, Dingcheng Wang, Jason Hartline +1LLM Reasoning StrategiesProbabilistic Model

  3. How reliable are LLMs when it comes to playing dice?

    Jun 5, 2026Luca Avena, Gianmarco Bet, Bernardo BusoniLLM Reasoning StrategiesProbabilistic Model