cs.AISep 27, 2026

Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation

Authors: Gowthamkumar Nandakishore

Abstract

The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is −0.214-0.214, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, 0.2140.214. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature (T=0.469T=0.469, sharpening) removes most of the miscalibration (held-out ECE 0.2040.204 to 0.0370.037) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy (0.7670.767 vs. 0.7660.766). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at q=0.05q=0.05 (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy 0.6170.617), and every score measures agreement with a synthetic teacher whose self-agreement ceiling (0.7350.735) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do System One Decisions Add Up? A Study of Probabilistic Coherence

    Sep 27, 2026Saman Sarker JoyConfidence CalibrationDecisions

  2. Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

    Jan 12, 2026Yuxi Xia, Dennis Ulmer, Terra Blevins +3Confidence Estimation

  3. Two Axes of LLM Abstention: Answer Correctness and Question Answerability

    Jul 9, 2026Benedikt J. WagnerAnswerInstruction-Tuned Models