cs.CLOct 6, 2026

Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers

Authors: Santhosh Kumar Kasa, Siva Rajesh Kasa, Sumit Negi

Organizations: Amazon

Abstract

Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains

    Jul 4, 2026Zhuoer Shen, Mingyi Wang, Shaofeng Zou +1Machine-Generated Text DetectionBert-Based Models

  2. Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

    May 30, 2026Yongxi Zhou, Lai Yun Choi, Jiaxi Wen +1Pass-Rate EvaluationHigh-Fidelity Conditional Generation