Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.
Figures & tables
Figure 1
Figure 2: Request-level shape changes within a fixed deployment: batch size (against the single-example reference) and padded width (static to dynamic), for full-precision GPU, bf16 GPU, and 8-bit integer CPU. Padded width is the dominant effect in every deployment; batch size is near the floor at full precision and grows only under reduced precision and 8-bit integer execution.
Figure 3: Two combined deployment changes a practitioner encounters: a whole-deployment migration (left) and serving load within a fixed deployment (right).
predicted label of checkpoint θ on x under context c
pθ(⋅∣x,c) ; z
class-probability vector; class logits
Flipθ(x;c0,c)
indicator that fθ(x;c)=fθ(x;c0)
TV(x,c,c′)
total variation between the probability vectors under c and c′
Appendix
Table 5
Configuration
Conditions
Result
CPU PyTorch fp32
BERT and GPT-2; 1 and 32 threads
bitwise (4 repeats)
GPU PyTorch
fp32, fp16, bf16; all four formulations
bitwise (4 repeats)
GPU PyTorch, held-out replication
fp32, bf16, fp16
2,304 / 2,304 bitwise
ONNX Runtime CPU
BERT SST-2, five repeats
64 / 64 bitwise
GPU fp32, no positional encoding
four identical contexts
768 / 768 bitwise
GPU, fixed-width length comparisons
IMDB and AG News; fp32 and bf16
768 / 768; 576 / 576 bitwise
Appendix
Table 3: Same-context repeatability floor. Identical inputs, identical context, repeated runs; entries are bitwise-identical comparisons over the total.
Figure 5: Intervening on the position assignment. Assigning positions by order among real tokens removes the observed padding-induced class changes; deliberately restoring column positions on models trained with order positions brings the changes back. Matching the assignment rule between training and inference is separately required to preserve accuracy (Section ).
Figure 6: Two further amplifiers of numerical drift: the decision margin suppresses numerical label changes but not padding (left), and a longer real-token input increases per-forward drift at fixed padded width (right).
Axis
Metric
Small
Medium
Large
Padding
Flip
0.2505
0.1840
0.2642
Padding
TV (no flip)
0.0716
0.0880
0.1441
Padding
∥Δz∥∞ (no flip)
5.09
10.06
20.84
Precision
Flip
0.0368
0.0103
0.00568
Deploy. profile
Flip
0.0330
0.00879
0.00666
Batch
Flip
7.5×10−6
1.09×10−5
3.56×10−4
Appendix
Table 4: Comparison across model sizes, restricted to models of comparable accuracy. Padding rows use the 16 GPT-2 model, dataset, and seed combinations for which all three sizes pass the accuracy threshold; the remaining rows use 40 matched combinations across all four families. Total variation (TV) is computed only over examples whose predicted class does not change. Paired uncertainty for the principal contrasts is reported in the text.
Learned absolute
Rotary (RoPE)
Left-padded
GPT-2: 35,591 / 176,778 (20.1%)
Qwen3: 0 / 29,463
Right-padded
BERT, MLM: 0 / 176,778
ModernBERT: 0 / 29,463
Appendix
Table 5: Padding-induced class changes by position mechanism and padding side. Changes occur only where padding moves real tokens that carry a learned embedding for each absolute position.
Dataset
Label changes
Max. logit drift
SST-2 (2 classes)
2 of 64 (3.1%)
about 0.45
SST-5 (5 classes)
4 to 5 of 64 (6.2 to 7.8%)
0.21 to 0.27
Appendix
Table 6: Label changes and logit drift for the fully generative GPT-2 classifier under the same bf16 batch-shape change, by number of classes (64 targets per cell).
fp32
dynamic int8
dynamic fake-quant
static int8
BERT max ∣Δ∣ logit
0 (bitwise)
0.09
0.015 to 0.066
0 (bitwise)
Gen. GPT-2 max ∣Δ∣ logit
0 (bitwise)
0.774
0.091 to 0.307
0 (bitwise)
Max TV
0
0.124
0.092
0
Label changes
0
up to 2 / 64
up to 2 / 64
0 / 64
Target slot (fixed shape)
bitwise
bitwise
—
bitwise
Appendix
Table 7: Effect of changing the other examples in the target’s batch at fixed batch shape (CPU, static width 128, 64 targets, SST-2 and AG News, threads 1 and 32). Only the batch-global dynamic activation scale lets those other examples reach the target.
Prompt (fp32 margin)
fp32 divergence step
bf16 divergence step
TV at divergence
mismatches at 128
Low (0.017)
none
14
0.122
91
Medium (0.579)
none
4
0.094
119
High (2.200)
none
9
0.123
116
Appendix
Table 8: Accumulation experiment: greedy decoding under two batch shapes (GPT-2, 128 steps, 12 of 12 same-context floor checks bitwise).
Figure 7: Latency of batch-invariant matrix-multiplication kernels, as the ratio of the batch-invariant to the standard whole-forward median, by batch size, for four classifier formulations. A ratio below one is faster. Batch invariance is roughly neutral at batch size 1 (a dispatch-bound effect under eager execution) and costs about 1.3 to 1.5 times at large batch under a precision-matched comparison (true fp32, TF32-matched, or bf16); the default fp32 comparison, in which the invariant kernel runs at reduced tensor-core precision against a full-precision baseline, is the only one that appears faster at large batch.
Figure 8: Batch size (against the single-example reference) and padded width (static to dynamic) by deployment; batch size is shown as bar shade. Rows are label-change rate, probability total variation, and maximum logit change; BERT solid, fully generative GPT-2 hatched.
Figure 9: Within-batch rearrangements at a fixed batch shape (target slot, order of the other examples, composition of the other examples) by deployment; batch size shown as bar shade. Reordering is bitwise at every precision; composition is bitwise at full precision and bf16 and moves only under dynamic 8-bit integer quantization; the target slot is bitwise on the integer CPU path but shows a small nonzero score movement at full precision and bf16 (a matrix-tiling effect, with no label change).
Figure 10: Batch size and padded width by deployment and model size (small, medium, large shown as bar shade), pooled over batch size.
Figure 11: Within-batch rearrangements by deployment and model size (small, medium, large shown as bar shade), pooled over batch size.
Batch size
Format
8
32
128
256
fp32 (GPU)
19/192 (9.90%)
12/192 (6.25%)
6/192 (3.13%)
3/192 (1.56%)
bf16 (GPU)
20/192 (10.42%)
16/192 (8.33%)
5/192 (2.60%)
3/192 (1.56%)
int8, dynamic (CPU)
16/192 (8.33%)
18/192 (9.38%)
7/192 (3.65%)
4/192 (2.08%)
int8, static (CPU)
15/192 (7.81%)
16/192 (8.33%)
9/192 (4.69%)
4/192 (2.08%)
Appendix
Table 9: GPT-2 label changes from dynamic versus static padding, by numerical format and batch size. BERT shows zero changes at every GPU cell and at most 1 of 192 at a single int8 cell.
BERT
GPT-2
Precision change
Changes
Max logit drift
Changes
Max logit drift
fp32 to fp16
0/192
0.0031
0/192
0.0638
fp32 to bf16
0/192
0.0283
2/192
0.564
bf16 to fp16
0/192
0.0273
2/192
0.570
Appendix
Table 10: Precision pairs on the matched six-checkpoint set, 192 comparisons per family (three model sizes, 64 targets each). Each cell gives label changes over the denominator and the maximum absolute logit difference.
BERT
GPT-2
Condition
Changes
Max drift
Changes
Max drift
Device changed, no quantization
0/128
2.86×10−6
0/128
1.80×10−3
Device held at GPU, dynamic int8
1/192
0.055
1/192
0.276
Device held at CPU, dynamic int8
0/128
0.054
1/128
0.585
Device held at CPU, static int8
0/128
0.030
0/128
0.301
Device changed and dynamic int8 applied
1.49% mean label change over 158 pairs
Appendix
Table 11: Decomposing the whole-migration change (GPU fp32 to CPU dynamic int8) into a device factor and a quantization factor. The GPU-held row uses a six-checkpoint set (192 comparisons per family); the CPU-held rows use a separate medium-size checkpoint per family on SST-2 and AG News (128 per family), so the rows are not pooled. The whole-migration rate is measured on the main grid with different checkpoints.
Device
Format
Batch size
Composition
Target slot
Order
Galaxy S25
fp32
bitwise
bitwise
bitwise
bitwise
Galaxy S25
dynamic int8
n.m.
0.0048
bitwise
bitwise
Galaxy S25
static int8
n.m.
bitwise
bitwise
bitwise
Pixel 8a
fp32
bitwise
bitwise
bitwise
bitwise
Pixel 8a
dynamic int8
n.m.
0.0048
bitwise
bitwise
Pixel 8a
static int8
n.m.
bitwise
bitwise
bitwise
Appendix
Table 12: Serving-context factors on physical phones. Entries give the maximum total variation of the target probability vector, or “bitwise” when all 64 of 64 targets were bitwise identical, for CPU execution. Cells marked n.m. were not measured because the int8 graphs were exported at a fixed batch size of 8.
Intervention
Pairs
Label
Threshold
Abstention
Pred. set
Top 1% rank
Padding policy
218,439
10.224
18.546
9.161
15.468
32.010
Precision (fp32 to bf16)
799,620
0.991
3.077
2.141
1.932
12.655
Batch size (32 to 8)
799,620
0.000
0.0004
0.000
0.0003
0.0042
Appendix
Table 13: Fraction of evaluation-example pairs (percent) with a changed decision under each intervention. Abstention and prediction sets use 90 percent target coverage. Ranking churn is the fraction of the top 1 percent set whose membership differs, ∣S△S′∣/2k for selected sets of size k per class, not a fraction of all examples.
Policy
Recompute rate
Disagreements detected
Label margin
22.45% (286 of 1,274)
44 of 44
Fixed threshold
4.00% (51 of 1,274)
18 of 21
Appendix
Table 14: Held-out replay results for GPT-2 medium on SST-2 at fast batch size 32 (1,274 held-out targets).
Comparison
Pairs
Mean flip
95% CI
Median
P90
Maximum
Padding
158
0.10297
[0.07862, 0.12987]
0
0.40835
0.671
Precision
790
0.01619
[0.00968, 0.02429]
0.00250
0.03032
0.404
Whole migration (GPU fp32 to CPU int8)
158
0.01492
[0.00940, 0.02229]
0.00439
0.03114
0.386
Batch size
3,792
0.00019
[0.00010, 0.00029]
0
0
0.052
Global order
948
≈0
[0, 0.00001]
0
0
0.001
Target slot
474
zero observed
Appendix
Table 15: Serving-context instability among the 158 from-scratch checkpoints that exceed 1.15 times chance validation accuracy. Means average designed pairs rather than production traffic. P90 is the 90th percentile of pair-level flip rates.