cs.AISep 27, 2026

Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation

Authors: Gowthamkumar Nandakishore

Abstract

The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is −0.214-0.214, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, 0.2140.214. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature (T=0.469T=0.469, sharpening) removes most of the miscalibration (held-out ECE 0.2040.204 to 0.0370.037) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy (0.7670.767 vs. 0.7660.766). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at q=0.05q=0.05 (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy 0.6170.617), and every score measures agreement with a synthetic teacher whose self-agreement ceiling (0.7350.735) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 27, 2026cs.CL

Do System One Decisions Add Up? A Study of Probabilistic Coherence

A decision model can give probabilities that sum to one for every question yet disagree with itself when the same decision is broken into smaller steps. We study this form of probabilistic coherence in Jev and the English Laya checkpoint, using 2,500 matched examples per system across TREC, CLINC150, and MASSIVE. Across 72,000 classification questions, we compare direct fine-label predictions with broad-category probabilities and predictions reconstructed through those categories. Both systems show substantial disagreement: mean category-level total variation ranges from 0.219 to 0.349 for Jev and from 0.424 to 0.689 for Laya, on a scale where zero means exact agreement. The consequences differ sharply. On CLINC150, reconstruction reduces Jev's accuracy by 22.9 percentage points (paired 95% bootstrap interval: [-24.9, -20.9]) and improves Laya's by 21.3 points ([18.0, 24.5]). The same directions hold across all three datasets, with all six unadjusted accuracy-change intervals excluding zero. Improved accuracy can also accompany less reliable confidence: on MASSIVE, Laya gains 9.2 accuracy points while its expected calibration error rises from 0.046 to 0.124. Error analysis identifies both broad-category mistakes and within-category confusions. These findings show why decision systems need joint evaluation of accuracy, confidence calibration, and probability coherence in the workflow used by an application.
Jan 12, 2026cs.CL

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the variability of language: confidence estimates should remain consistent under semantically equivalent prompts or answer variations, while changing when answer meaning differs, as this may indicate a change in correctness. Therefore, we introduce a novel evaluation framework based on three complementary properties: \textbf{robustness} to prompt perturbations, \textbf{stability} across semantically equivalent answers, and \textbf{sensitivity} to semantically different answers. We show that these metrics are largely independent from existing CE metrics, and that common CE methods often fail on them: while most methods achieve high robustness and stability, they struggle to distinguish semantically different answers, potentially because they do not effectively leverage generation-side information. Overall, our framework exposes overlooked limitations of current CE evaluations and provides guidance for selecting confidence estimators for real-world applications.
Jul 9, 2026cs.CL

Two Axes of LLM Abstention: Answer Correctness and Question Answerability

A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.