Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces
Organizations: Independent Researcher
Abstract
Equivalent probability requests can lead to different decisions even when both reports are valid. We introduce probability contracts, a benchmark that connects exact finite-world posteriors, validated event alignment, interface coherence, and failure-aware decision evaluation. Across four model-interface configurations on 1,000 worlds, Kev has lower aggregate canonical posterior error than Jev but larger complement and coarsening residuals; accuracy ordering varies by stratum. Jev's Event and Choice interfaces change the binary action on 32.8% of valid pairs at defer cost 0.10. Post-hoc analyses show that disagreement certifies only 11-52% of mean binary pair error and does not consistently outperform confidence for selection. An action-region characterization and a standard scoring-rule identity explain averaging's expected Brier guarantee relative to random interface selection, but not a decision-loss guarantee at each cost. The loss contrast takes both signs on a 99-cost grid for every configuration; small penalties where both policies beat deferral have pointwise intervals containing zero. Secondary checks specified before collection include a separate 400-root cohort, where Event/Choice effects remain configuration-dependent. A joint surface-order and answer-ID intervention shifts posttrained probabilities. Both bounded reasoning arms yield no valid probability reports, leaving their probability accuracy undefined. Probability contracts make these distinctions measurable by evaluating event semantics, posterior error, coverage, and decision cost together.
Figures & tables
| Relation | Aligned claims | Residual or condition | Eligibility |
|---|---|---|---|
| Complement | Exhaustive binary event | ||
| Coarsening | Same partition and information | ||
| Event/Choice | Absolute difference | Bijection to same event | |
| Intersection | Distance outside Fréchet bounds | Aligned | |
| Product | |||
| Permutation | Total variation after inverse map | Bijection; same information |
| Relation | Kev | Base | Posttrained | Jev |
|---|---|---|---|---|
| Event/Choice | 0.0629 | 0.2138 | 0.1126 | 0.1938 |
| Complement | 0.5055 | 0.1155 | 0.8000 | 0.0929 |
| Coarsening | 0.3554 | 0.2451 | 0.2836 | 0.1751 |
| Intersection bound | 0.1553 | 0.0814 | 0.0358 | 0.0200 |
| Product identity | 0.1328 | 0.2682 | 0.1406 | 0.0836 |
| Permutation TV | 0.1143 | 0.3224 | 0.3470 | 0.0534 |
| Canonical multiclass task | Binary task | |||
|---|---|---|---|---|
| Configuration | Expected loss | Oracle regret | Defer rate | Event/Choice changes |
| Kev | 0.1090 [0.1054, 0.1131] | 0.0193 [0.0158, 0.0232] | 0.9610 [0.9500, 0.9720] | 0.1490 [0.1290, 0.1710] |
| Base | 0.1000 [0.1000, 0.1000] | 0.0103 [0.0088, 0.0118] | 1.0000 [1.0000, 1.0000] | 0.0000 [0.0000, 0.0000] |
| Posttrained | 0.1000 [0.1000, 0.1000] | 0.0103 [0.0088, 0.0118] | 1.0000 [1.0000, 1.0000] | 0.3420 [0.3210, 0.3630] |
| Jev-1.13.0 | 0.1336 [0.1263, 0.1414] | 0.0438 [0.0366, 0.0515] | 0.8160 [0.7920, 0.8380] | 0.3280 [0.3020, 0.3540] |
| Configuration | Event (1) | Choice (1) | Mean (2) | Agreement (2) | Repeat mean (2) |
|---|---|---|---|---|---|
| Kev | 0.16695 | 0.15786 | 0.15281 | 0.14186 | 0.16695 |
| Base | 0.10000 | 0.10000 | 0.10000 | 0.10000 | 0.10000 |
| Posttrained | 0.25448 | 0.23257 | 0.19055 | 0.17239 | 0.25448 |
| Jev | 0.09654 | 0.17052 | 0.09890 | 0.09654 | 0.09645 |
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Defer cost | Action disagreement | 95% interval | Loss difference | 95% interval |
|---|---|---|---|---|
| 0.02 | 0.0535 | [0.0495, 0.0575] | [ , ] | |
| 0.05 | 0.1980 | [0.1865, 0.2095] | [ , ] | |
| 0.10 | 0.6085 | [0.5980, 0.6195] | [ , ] | |
| 0.20 | 0.6780 | [0.6660, 0.6905] | [ , ] | |
| 0.40 | 0.2705 | [0.2615, 0.2795] | [ , ] |
| Configuration | Valid | Accuracy | TV | Exp. Brier | Gap | Complement | Repeat |
|---|---|---|---|---|---|---|---|
| Kev (FP32) | 2,000/2,000 | 0.460 | 0.3426 | 0.6167 | 0.0772 | 0.5217 | 0.0000 |
| Base (FP32) | 2,000/2,000 | 0.365 | 0.4326 | 0.7185 | 0.2109 | 0.1186 | 0.0000 |
| Posttrained (FP32) | 2,000/2,000 | 0.420 | 0.4001 | 0.6786 | 0.1156 | 0.7929 | 0.0000 |
| Reasoning (BF16) | 0/2,000 | – | – | – | – | – | – |
| Variant | Intended | Eligible |
|---|---|---|
| Canonical | 1,000 | 1,000 |
| Event | 1,000 | 1,000 |
| Complement | 1,000 | 1,000 |
| Query of event | 1,000 | 1,000 |
| Intersection | 1,000 | 1,000 |
| Conditional | 1,000 | 913 |
| Configuration | Valid canonical roots | Posterior TV | Expected Brier |
|---|---|---|---|
| Kev | 1,000 | 0.2350 [0.2278, 0.2423] | 0.6333 [0.6229, 0.6436] |
| Base | 1,000 | 0.3740 [0.3655, 0.3824] | 0.7474 [0.7379, 0.7571] |
| Posttrained | 1,000 | 0.3202 [0.3124, 0.3280] | 0.6938 [0.6831, 0.7048] |
| Jev-1.13.0 | 990 | 0.2677 [0.2599, 0.2754] | 0.6560 [0.6415, 0.6705] |
| Configuration | Realized Brier |
|---|---|
| Kev | 0.6359 [0.6160, 0.6571] |
| Base | 0.7530 [0.7332, 0.7739] |
| Posttrained | 0.6918 [0.6696, 0.7137] |
| Jev-1.13.0 | 0.6469 [0.6094, 0.6851] |
| Configuration | Native validity | Exact- gap | Operational criterion |
|---|---|---|---|
| Kev | 1.0000 | 0.0618 | No |
| Base | 1.0000 | 0.1441 | No |
| Posttrained | 1.0000 | 0.0877 | No |
| Jev-1.13.0 | 0.9946 | 0.1952 | No |
| Relation | Mean [95% interval] | Valid roots | Eligible roots |
|---|---|---|---|
| Complement | 0.5055 [0.4971, 0.5138] | 1000 | 1000 |
| Event/Choice interface | 0.0629 [0.0610, 0.0648] | 1000 | 1000 |
| Coarsening | 0.3554 [0.3489, 0.3620] | 1000 | 1000 |
| Intersection bounds | 0.1553 [0.1507, 0.1598] | 1000 | 1000 |
| Product | 0.1328 [0.1300, 0.1356] | 913 | 913 |
| Exact repeat | 0.0000 [0.0000, 0.0000] | 1000 | 1000 |
| Relation | Mean [95% interval] | Valid roots | Eligible roots |
|---|---|---|---|
| Complement | 0.1155 [0.1133, 0.1177] | 1000 | 1000 |
| Event/Choice interface | 0.2138 [0.2111, 0.2165] | 1000 | 1000 |
| Coarsening | 0.2451 [0.2420, 0.2481] | 1000 | 1000 |
| Intersection bounds | 0.0814 [0.0807, 0.0821] | 1000 | 1000 |
| Product | 0.2682 [0.2673, 0.2691] | 913 | 913 |
| Exact repeat | 0.0000 [0.0000, 0.0000] | 1000 | 1000 |
| Relation | Mean [95% interval] | Valid roots | Eligible roots |
|---|---|---|---|
| Complement | 0.8000 [0.7982, 0.8017] | 1000 | 1000 |
| Event/Choice interface | 0.1126 [0.1085, 0.1168] | 1000 | 1000 |
| Coarsening | 0.2836 [0.2776, 0.2896] | 1000 | 1000 |
| Intersection bounds | 0.0358 [0.0342, 0.0375] | 1000 | 1000 |
| Product | 0.1406 [0.1389, 0.1424] | 913 | 913 |
| Exact repeat | 0.0000 [0.0000, 0.0000] | 1000 | 1000 |
| Relation | Mean [95% interval] | Valid roots | Eligible roots |
|---|---|---|---|
| Complement | 0.0929 [0.0890, 0.0967] | 1000 | 1000 |
| Event/Choice interface | 0.1938 [0.1873, 0.2002] | 1000 | 1000 |
| Coarsening | 0.1751 [0.1690, 0.1813] | 990 | 1000 |
| Intersection bounds | 0.0200 [0.0173, 0.0227] | 999 | 1000 |
| Product | 0.0836 [0.0804, 0.0868] | 912 | 913 |
| Exact repeat | 0.0188 [0.0176, 0.0200] | 1000 | 1000 |
| View | Question |
|---|---|
| Eligible | Was the semantic transformation defined? |
| Native valid | Did the returned object satisfy its interface contract? |
| Paired valid | Are all vectors needed for this contrast scoreable? |
| Intended pipeline | What loss results when failure triggers defer? |
| Stratum | Kev | Base | Posttrained | Jev |
|---|---|---|---|---|
| Finite population / joined tables | 0.2696 | 0.4061 | 0.2381 | 0.2429 |
| Finite population / nested records | 0.1866 | 0.3905 | 0.3408 | 0.1858 |
| Joint attributes / joined tables | 0.1764 | 0.3097 | 0.2468 | 0.4017 |
| Joint attributes / nested records | 0.1868 | 0.2937 | 0.3489 | 0.3831 |
| Missing rule / joined tables | 0.4545 | 0.4111 | 0.3430 | 0.1931 |
| Missing rule / nested records | 0.1598 | 0.3614 | 0.3516 | 0.1272 |
| Configuration | Mean | Minimum | Maximum | |
|---|---|---|---|---|
| Kev | 1.5055 | 1.0485 | 1.8153 | 1000 |
| Base | 1.1103 | 0.9344 | 1.3733 | 861 |
| Posttrained | 1.8000 | 1.5743 | 1.9150 | 1000 |
| Jev | 1.0641 | 0.6700 | 1.3900 | 786 |
| Configuration | Valid requests | Canonical TV ( ) | Complement gap ( ) | Loss | Defer |
|---|---|---|---|---|---|
| Mistral likelihood | 2385/2385 | 0.4799 (200) | 0.6469 (200) | 0.1734 | 0.8900 |
| Mistral JSON | 1454/2385 | 0.3937 (59) | 0.4848 (139) | 0.1119 | 0.9650 |
| OLMo likelihood | 2385/2385 | 0.4261 (200) | 0.2048 (200) | 0.1000 | 1.0000 |
| OLMo JSON | 1248/2385 | 0.4220 (60) | 0.2014 (103) | 0.1142 | 0.9700 |
| Reason | Mistral | OLMo |
|---|---|---|
| Invalid JSON | 528 | 4 |
| Wrong key support | 0 | 536 |
| Nonnumeric probability | 21 | 0 |
| Trailing content | 8 | 0 |
| Out-of-range probability | 62 | 1 |
| Simplex sum mismatch | 312 | 596 |
| Configuration | valid | TV | Expected Brier | Expected loss |
|---|---|---|---|---|
| Mistral likelihood | 200 | 0.4799 [0.4517, 0.5099] | 0.9696 [0.9275, 1.0132] | 0.1734 [0.1531, 0.1942] |
| Mistral JSON | 59 | 0.3937 [0.3506, 0.4358] | 0.8126 [0.7548, 0.8725] | 0.1119 [0.1001, 0.1271] |
| OLMo likelihood | 200 | 0.4261 [0.3955, 0.4583] | 0.8566 [0.8224, 0.8913] | 0.1000 [0.1000, 0.1000] |
| OLMo JSON | 60 | 0.4220 [0.3767, 0.4666] | 0.8224 [0.7598, 0.8884] | 0.1142 [0.1020, 0.1299] |
| Kev (historical) | 200 | 0.2349 [0.2179, 0.2522] | 0.6239 [0.6028, 0.6445] | 0.1079 [0.1016, 0.1149] |
| Base (historical) | 200 | 0.3785 [0.3601, 0.3967] | 0.7474 [0.7250, 0.7698] | 0.1000 [0.1000, 0.1000] |
| Relation | likelihood | Likelihood residual | JSON | JSON residual |
|---|---|---|---|---|
| Event/Choice | 200 | 0.0000 [0.0000, 0.0000] | 140 | 0.0000 [0.0000, 0.0000] |
| Complement | 200 | 0.6469 [0.6340, 0.6600] | 139 | 0.4848 [0.4423, 0.5269] |
| Coarsening | 200 | 0.2882 [0.2730, 0.3034] | 51 | 0.2055 [0.1664, 0.2448] |
| Intersection bound | 200 | 0.1091 [0.1052, 0.1131] | 124 | 0.2184 [0.1904, 0.2490] |
| Product | 185 | 0.1575 [0.1527, 0.1625] | 146 | 0.3279 [0.2957, 0.3626] |
| Permutation | 200 | 0.6167 [0.5891, 0.6432] | 39 | 0.4039 [0.3559, 0.4553] |
| Relation | likelihood | Likelihood residual | JSON | JSON residual |
|---|---|---|---|---|
| Event/Choice | 200 | 0.0000 [0.0000, 0.0000] | 196 | 0.0000 [0.0000, 0.0000] |
| Complement | 200 | 0.2048 [0.2012, 0.2084] | 103 | 0.2014 [0.1742, 0.2297] |
| Coarsening | 200 | 0.4655 [0.4599, 0.4711] | 58 | 0.2456 [0.1977, 0.2937] |
| Intersection bound | 200 | 0.0062 [0.0053, 0.0071] | 46 | 0.2206 [0.1730, 0.2713] |
| Product | 185 | 0.0984 [0.0958, 0.1010] | 45 | 0.3613 [0.3201, 0.4026] |
| Permutation | 200 | 0.4890 [0.4633, 0.5152] | 34 | 0.4036 [0.3343, 0.4757] |
| Configuration | |||||
|---|---|---|---|---|---|
| Mistral likelihood | 0.0200 | 0.0581 | 0.1734 | 0.3124 | 0.5635 |
| Mistral JSON | 0.0200 | 0.0523 | 0.1119 | 0.2140 | 0.4304 |
| OLMo likelihood | 0.0200 | 0.0500 | 0.1000 | 0.2000 | 0.4304 |
| OLMo JSON | 0.0299 | 0.0640 | 0.1142 | 0.2177 | 0.4195 |
| Kev (historical) | 0.0200 | 0.0500 | 0.1079 | 0.1999 | 0.3913 |
| Base (historical) | 0.0200 | 0.0500 | 0.1000 | 0.2000 | 0.3844 |
| Contrast | TV | TV difference | Expected-loss difference |
|---|---|---|---|
| Mistral likelihood – Mistral JSON | 59 | 0.1145 [0.0454, 0.1855] | 0.0615 [0.0342, 0.0871] |
| Mistral likelihood – OLMo likelihood | 200 | 0.0538 [0.0331, 0.0750] | 0.0734 [0.0531, 0.0942] |
| Mistral likelihood – OLMo JSON | 60 | 0.0790 [0.0015, 0.1582] | 0.0591 [0.0368, 0.0816] |
| Mistral JSON – OLMo likelihood | 59 | -0.0780 [-0.1505, -0.0077] | 0.0119 [0.0001, 0.0271] |
| Mistral JSON – OLMo JSON | 21 | -0.1047 [-0.1629, -0.0506] | -0.0024 [-0.0226, 0.0179] |
| OLMo likelihood – OLMo JSON | 60 | -0.0164 [-0.0902, 0.0558] | -0.0142 [-0.0299, -0.0020] |
| Configuration and estimand | F/T | F/R | J/T | J/R | M/T | M/R | N/T | N/R |
|---|---|---|---|---|---|---|---|---|
| Mistral likelihood, canonical | 25 | 25 | 25 | 25 | 25 | 25 | 25 | 25 |
| Mistral likelihood, Event/Choice | 25 | 25 | 25 | 25 | 25 | 25 | 25 | 25 |
| Mistral JSON, canonical | 8 | 1 | 8 | 2 | 15 | 16 | 8 | 1 |
| Mistral JSON, Event/Choice | 25 | 12 | 18 | 7 | 24 | 24 | 19 | 11 |
| OLMo likelihood, canonical | 25 | 25 | 25 | 25 | 25 | 25 | 25 | 25 |
| OLMo likelihood, Event/Choice | 25 | 25 | 25 | 25 | 25 | 25 | 25 | 25 |
| Configuration | Pair absolute error | Disagreement lower bound | Error beyond bound |
|---|---|---|---|
| Kev | 0.29397 [0.28537, 0.30287] | 0.03143 [0.03048, 0.03240] | 0.26254 [0.25384, 0.27140] |
| Base | 0.25692 [0.24884, 0.26514] | 0.10689 [0.10553, 0.10816] | 0.15004 [0.14149, 0.15864] |
| Posttrained | 0.37758 [0.36562, 0.38971] | 0.05632 [0.05423, 0.05836] | 0.32126 [0.30941, 0.33318] |
| Jev | 0.18567 [0.18023, 0.19105] | 0.09688 [0.09362, 0.10010] | 0.08879 [0.08376, 0.09393] |
| Configuration | Mean minus Random interface | Agreement minus mean |
|---|---|---|
| Kev | -0.00959 [-0.01347, -0.00581] | -0.01095 [-0.01503, -0.00733] |
| Base | 0.00000 [0.00000, 0.00000] | 0.00000 [0.00000, 0.00000] |
| Posttrained | -0.05297 [-0.05964, -0.04634] | -0.01816 [-0.02399, -0.01291] |
| Jev | -0.03463 [-0.03886, -0.03049] | -0.00236 [-0.00411, -0.00077] |
| Configuration | Event | Choice | Mean | Agreement | Repeat mean | Repeat agreement | Oracle | |
|---|---|---|---|---|---|---|---|---|
| Kev | 0.02 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0165 |
| Kev | 0.05 | 0.0512 | 0.0516 | 0.0506 | 0.0501 | 0.0512 | 0.0512 | 0.0411 |
| Kev | 0.10 | 0.1669 | 0.1579 | 0.1528 | 0.1419 | 0.1669 | 0.1669 | 0.0813 |
| Kev | 0.20 | 0.3128 | 0.3052 | 0.3122 | 0.2885 | 0.3128 | 0.3128 | 0.1571 |
| Kev | 0.40 | 0.3988 | 0.3986 | 0.3967 | 0.3950 | 0.3988 | 0.3988 | 0.2689 |
| Base | 0.02 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0200 | 0.0165 |
| Original actions | Midpoint action | |
|---|---|---|
| Same action | Same action | |
| No, defer | No | |
| No, defer | Defer | |
| Defer, yes | Defer | |
| Defer, yes | Yes | |
| No, yes | No |
| Configuration | Lower | Equal | Higher | Both below |
|---|---|---|---|---|
| Jev | 82 | 1 | 16 | 16 |
| Kev | 62 | 5 | 32 | 7 |
| Base | 47 | 36 | 16 | 0 |
| Posttrained | 47 | 1 | 51 | 2 |
| Configuration | Disagreement | Confidence | Random |
|---|---|---|---|
| Jev | 0.17311 | 0.17943 | 0.22314 |
| Kev | 0.25827 | 0.23729 | 0.25416 |
| Base | 0.29737 | 0.29122 | 0.29496 |
| Posttrained | 0.29110 | 0.27650 | 0.29496 |
| Runtime stage | Mode | Acquired | Valid | Capped | Interrupted |
|---|---|---|---|---|---|
| Original | Thinking | 6 | 3 | 2 | 1 |
| Original | Non-thinking | 6 | 6 | 0 | 0 |
| Optimized single replica | Thinking | 5 | 4 | 0 | 1 |
| Optimized single replica | Non-thinking | 4 | 4 | 0 | 0 |
| Four optimized replicas | Thinking | 69 | 64 | 5 | 0 |
| Four optimized replicas | Non-thinking | 69 | 69 | 0 | 0 |
| Request | TV | Brier | |||
|---|---|---|---|---|---|
| Canonical | 12 | -0.0789 | -0.0581 | +0.2902 | -0.1969 |
| Event | 15 | -0.1331 | -0.1517 | +0.1068 | -0.3008 |
| Choice | 12 | -0.2735 | -0.4217 | +0.0529 | -0.3893 |
| Complement | 13 | -0.0423 | -0.0287 | +0.1562 | -0.1480 |
| Repeat | 12 | -0.0968 | -0.1138 | +0.1313 | -0.2911 |