Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model
Abstract
Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: https://github.com/Syntheme/beyond-answer-confidence.
Figures & tables
| Ground-truth contrast | Intended to measure | Validated interpretation | Answer uncertainty | Question | AUROC | Wordings |
| Fixed but unknown fact vs. future chance event | outcome settled | outcome settled, also inferred from tense alone (1.00) | 0.96, only because Jev concentrates on the mode for chance events; ideal distributions identical | settled | 1.00 | 0.99–1.00 |
| Complete vs. incomplete evidence, same paragraph count | evidence sufficient | evidence sufficient, beyond length features (0.60 0.94) | 0.85 (confidence, higher = complete) | enough | 0.95 | 0.95–0.97 |
| Pseudo-word made-up vs. obscure real person (same relations) | model’s knowledge | name form: text alone 1.00 | 0.78 | known | 0.88 | 0.87–0.97 |
| Look-alike made-up vs. obscure real person | model’s knowledge | equivalent to answer uncertainty ( ) | 0.76 | known | 0.74 | — |
| News past vs. before the boundary | model’s knowledge | explicit dates: date alone 0.997 | 0.385 (0.615 reversed) | known | 0.91 | — |
| same, answer option | “not known” | 0.89 | — |
| Banking77 ( =77) | CLINC150 ( =150) | |||||
|---|---|---|---|---|---|---|
| Code carries | acc | acc | ||||
| nothing | .010 | .320 | .310 | .008 | .363 | .355 |
| 1 example | .681 | .828 | .170 | .911 | .904 | .024 |
| 8 examples | .899 | .927 | .037 | .979 | .974 | .016 |
| name | .821 | .893 | .084 | .926 | .933 | .025 |
| name + 8 | .909 | .939 | .039 | .981 | .978 | .015 |
| Setting | Ideal | |
|---|---|---|
| Opaque codes (Banking77/CLINC150) | .013/.007 | .32/.36 |
| same, open model | .013/.007 | .017/.010 |
| Fixed but unknown generated fact | .25 | .49 |
| Fabricated entities (PopQA templates) | .25 † | .52 |
| Winner of a made-up future event | .25 § | .76 |
| Fair die, all 720 option orders | .17 | .80 |
| Stated | 0 | .1 | .2 | .3 | .4 | .7 | MAE |
|---|---|---|---|---|---|---|---|
| Yes/no per outcome | .01 | .09 | .16 | .25 | .36 | .68 | .030 |
| Score per outcome | .05 | .13 | .22 | .34 | .45 | .74 | .028 |
| Instructed Choice | .00 | .00 | .00 | .01 | .99 | 1.0 | .230 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Set | mean of 3 | single, 2.5–97.5% | estimate | |
|---|---|---|---|---|
| Banking77, 8 examples | 770 | 0.037 | 0.031–0.040 | 100% |
| CLINC150, 1 example | 1,500 | 0.025 | 0.021–0.026 | 100% |
| PopQA | 14,267 | 0.021 | 0.021–0.025 | 100% |
| Daily Oracle, after change point | 1,680 | 0.305 | 0.304–0.310 | 0% |
| TriviaQA | 9,960 | 0.031 | 0.030–0.032 | 100% |
| SimpleQA Verified | 1,000 | 0.051 | 0.048–0.061 | 11% |
| Maps allowed | In-sample | Random months (descriptive), mean / 95th pct | Cross-fitted mean (range) |
|---|---|---|---|
| Value-wise (each value its own level) | 0.066 | 0.034 / 0.038 | 0.170 (0.147–0.191) |
| Contiguous partitions (exact; lower bound for monotone maps) | 0.047 | 0.005 / 0.013 | 0.057 (0.046–0.074) |
| Arbitrary groupings, 2–20 levels (search) | 0.043–0.046 | — | 0.093–0.099 |
| band | before | acc. before | after | acc. after | Difference [95% CI] |
|---|---|---|---|---|---|
| 0.5–0.6 | 988 | 0.526 | 206 | 0.466 | 0.060 [ , 0.139] |
| 0.6–0.7 | 943 | 0.621 | 204 | 0.456 | 0.166 [0.100, 0.228] |
| 0.7–0.8 | 777 | 0.699 | 224 | 0.469 | 0.230 [0.152, 0.306] |
| 0.8–0.9 | 821 | 0.764 | 370 | 0.484 | 0.280 [0.236, 0.327] |
| 0.9–1.0 | 1,111 | 0.880 | 676 | 0.571 | 0.309 [0.260, 0.354] |
| Contrast (items) | Text only | known | Second look | |
|---|---|---|---|---|
| Pseudo-word made-up vs. all real (1,600 / 14,267) | 1.000 | 0.911 | 0.746 | — |
| Pseudo-word made-up vs. real, second-look subset (1,600 / 2,000) | — | 0.916 | 0.747 | 0.80 |
| Pseudo-word made-up vs. least popular real | 1.000 | 0.833 | 0.685 | — |
| Pseudo-word made-up vs. obscure real persons (600 / 555) | — | 0.884 | 0.784 | — |
| Look-alike made-up vs. obscure real persons (600 / 555) | 0.492 | 0.741 | 0.761 | — |
| News after vs. before, full panel (1,680 / 4,640) | 0.997 | 0.913 | 0.389 | — |
| Condition | Period | Coverage | Accuracy | Conf. errors (all) | Conf. errors (answered) |
|---|---|---|---|---|---|
| Base | before | 100% | 0.697 | 2.4% | 2.4% |
| Base | after | 100% | 0.511 | 17.3% | 17.3% |
| Today’s date given | after | 100% | 0.511 | 12.7% | 12.7% |
| “Not known” allowed | before | 64.3% | 0.740 | 0.3% | 0.5% |
| “Not known” allowed | after | 6.6% | 0.631 | 0.1% | 0.9% |
| Date + “not known” | after | 28.5% | 0.491 | 0.1% | 0.2% |
| Set | Accuracy | Brier | Uniform | Log loss | Uniform | ||
|---|---|---|---|---|---|---|---|
| TriviaQA | 4 | 9,960 | 0.955 | 0.067 | 0.750 | 0.132 | 1.386 |
| PopQA (real subjects) | 4 | 14,267 | 0.708 | 0.378 | 0.750 | 0.719 | 1.386 |
| Daily Oracle yes/no | 2 | 6,320 | 0.651 | 0.474 | 0.500 | 0.710 | 0.693 |
| HotpotQA, non-supporting paragraphs | 2 | 951 | 0.751 | 0.342 | 0.500 | 0.542 | 0.693 |
| HotpotQA, both supporting paragraphs | 2 | 951 | 0.958 | 0.071 | 0.500 | 0.193 | 0.693 |
| MMLU-CF | 4 | 10,000 | 0.776 | 0.377 | 0.750 | 1.075 | 1.386 |
| Banking77 | CLINC150 | |||||||
|---|---|---|---|---|---|---|---|---|
| Code carries | acc | acc | ||||||
| nothing | 0.010 | 0.320 | 0.774 | 0.310 | 0.008 | 0.363 | 0.592 | 0.355 |
| 1 example | 0.681 | 0.828 | 0.109 | 0.170 | 0.911 | 0.904 | 0.057 | 0.024 |
| 2 examples | 0.796 | 0.879 | 0.074 | 0.099 | 0.958 | 0.953 | 0.028 | 0.019 |
| 4 examples | 0.861 | 0.914 | 0.056 | 0.066 | 0.969 | 0.964 | 0.021 | 0.015 |
| 8 examples | 0.899 | 0.927 | 0.046 | 0.037 | 0.979 | 0.974 | 0.016 | 0.016 |
| Setting | Accuracy | Confidence | Gap | SmoothECE | |
|---|---|---|---|---|---|
| Fabricated entities (uniform reference 0.25) | 1,600 | — | 0.518 | — | — |
| Daily Oracle yes/no, after the change point | 1,680 | 0.511 | 0.816 | 0.305 | |
| MMLU-CF | 10,000 | 0.776 | 0.914 | 0.145 | |
| ANLI rounds 1–3 | 3,200 | 0.739 | 0.839 | 0.105 | |
| HotpotQA, non-supporting paragraphs | 951 | 0.751 | 0.830 | 0.083 | |
| HotpotQA, one of two needed paragraphs | 951 | 0.775 | 0.871 | 0.104 |
| Experiment | Units | Replicates | Paid calls | Dates | Planning |
|---|---|---|---|---|---|
| Instrument checks | — | varied | 402 | 26 Sep | planned |
| Intent study, development pilot | 6,356 | 3 | 7,945 | 26 Sep | pilot before the freeze |
| Intent dose–response (test) | 27,240 | 3 | 76,953 | 26–28 Sep | plan frozen in a git tag |
| Out-of-scope detection | 3,570 | 3 formats | 10,710 | 28 Sep | planned |
| Knowledge chance scenarios | 6,000 | 3 | 18,000 | 28 Sep | planned ∗ |
| Knowledge boundary (PopQA, fabricated, news) | 26,927 | 3 | 80,781 | 28 Sep | planned ∗ |
| Dataset | Source @ revision | Licence |
|---|---|---|
| Banking77 [ 9 ] | PolyAI task-specific-datasets @ 57ec275d | CC BY 4.0 |
| CLINC150 [ 24 ] | clinc/clinc_oos @ 155b9c71 | CC BY 3.0 |
| PopQA [ 28 ] | akariasai/PopQA @ 098765c7 | none stated |
| Daily Oracle [ 10 ] | agentic-learning-ai-lab/daily-oracle @ 455b35b2 | CC BY 4.0 |
| HotpotQA [ 48 ] | hotpotqa/hotpot_qa @ 1908d6af | CC BY-SA 4.0 |
| Quizbowl [ 33 ] | community-datasets/qanta @ e3c56022 | unknown |