Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with five metric families, namely bitwise equality, logit differences, top-K consistency, token agreement, and task accuracy. We trace one source of the mismatch to accumulation order inside vendors' matrix instructions. Upcasting to FP32 reduces the dense model's logit error by 43% at three times the runtime, yet keeping only the MLPs in BF16 retains 94% of this gain at 1.3 times the runtime, so most of the cost of full upcasting buys little. In the MoE models, FP32 and FP16 both lower the probability error but raise the logit error and change expert selection, and FP16 fails in the dense model. An output-head low-rank adapter (LoRA) does not help either, since the final hidden state does not predict the mismatch. The mismatch also carries into training. With every seed fixed, a student distilled from a teacher running on AMD answers 431 MMLU questions differently from one distilled from the same teacher on NVIDIA. Under FP32 upcasting, bitwise equality barely changes while the output distributions move most of the way to the reference, so judging cross-vendor agreement by a single measure misreads both its cost and its gains. Code is available at https://github.com/crova-project/crova.
Figures & tables
Figure 1: Cross-vendor mismatch spans the execution hierarchy. Inter-framework studies address differences in models, frameworks, and tensor-operation choices. At lower levels, kernel and hardware arithmetic can introduce additional numerical differences. The hierarchy organizes the analysis of how these differences propagate through model outputs and behavior.
Figure 2: Cross-vendor differences in a Jais dense model ( Jais Team, 2026 , a–d) and OLMoE MoE model ( Muennighoff et al., 2025 , e–h) . Components (a/e) and first blocks (b/f) receive identical inputs. Last-block outputs (c/g) and layer curves (d/h) use independent full forwards, allowing differences to propagate. The curves show the mean absolute difference across all tokens and hidden features at each block output.
Path
Captured
Synthetic
Full projection
Vendor GEMM
14738/15552
14444/15552
Ordered-tile
15211/15552
14960/15552
Scalar
0/15552
0/15552
First output tile
Ordered-tile
1022/1728
398/1728
Table 1: Mismatch in the matmul case study.
Figure 3: Two constructed examples of numerical mismatch across the five metric families. Case A changes logit values and lower rankings while preserving the generated sequence and correct answer. Case B changes the leading token and leads to an incorrect answer.
Figure 4: BF16 can gives different value due to less finer on ULP steps. BF16 values near one lie one ULP apart, sums round to the nearest, with ties to the even step. With b=0.5 ULP and c=0.504 ULP, in A=(a+b)+c \scriptsize1⃝ 1+b lands halfway, \scriptsize2⃝ rounds down to even so b is lost, and \scriptsize3⃝–\scriptsize4⃝ c rounds it up to 1+1 ULP. In B=(a+c)+b the sum rounds up to 1+1 ULP, then b lands halfway and rounds to the even 1+2 ULP. FP16 and FP32 have finer steps, so both orders agree.
Method
Bit-equal (%) ↑
Logit MSE ↓
KL ( 10−3 ) ↓
Top-5 (%) ↑
Rollout (%) ↑
MMLU (%) ↑
AMMLU/ ARC (%) ↑
Jais (dense)
Native AMD (BF16)
9.26
0.2873
7.08
66.53
93.50
60.86
71.36
FP16
nonfinite outputs
FP32 upcasting
11.24
0.1644
4.45
71.58
93.00
61.00
71.40
Selective FP32
11.07
0.1717
4.62
71.38
94.00
60.98
71.36
Head LoRA (KL)
9.22
0.2888
7.67
65.89
92.50
60.88
71.35
Table 2: AMD against the NVIDIA reference. Top-5 is ordered top-5 agreement, and rollout the share of responses identical to the reference. Selective FP32 keeps only the MLPs in BF16. Bold is best and underline second best per model. FP16 and FP32 outputs are rounded to BF16 for bit-equal.
Time (s)
Memory (GB)
Jais
Native AMD (BF16)
6,111
18.7
FP32 upcasting
19,245 (3.15 × )
47.7 (2.55 × )
Selective FP32 †
(1.30 × )
(2.02 × )
OLMoE
Native AMD (BF16)
1,129
14.4
Table 3: Cost on one MI210, same workload within each model. ∗ Relative to a native run measured alongside (1,054 s). Native timings vary by about 7% between runs. † All except MLPs in FP32 ( Section 5.3 )
vs. NVIDIA-teacher student
Accuracy (%) ↑
Teacher
Teacher
KL ( 10−4 ) ↓
Top-1 (%) ↑
Same ans. (%) ↑
MMLU
ARC
KL ( 10−4 ) ↓
No distillation
–
–
20.77
24.79 − 1.14
26.37 − 0.51
–
NVIDIA, repeat
0
100.00
100.00
25.93
26.88
0
NVIDIA, other seed
73.4
99.15
82.59
25.67 − 0.26
28.50 + 1.62
0
Native AMD (BF16)
1.61
99.84
96.93
26.12 + 0.19
27.22 + 0.34
1.94
AMD FP32 upcasting
23.7
99.51
94.67
25.64 − 0.29
26.79 − 0.09
1.87
Table 4: Distilling OLMoE into OLMo-1B with the teacher run on each execution. Same ans. is the share of MMLU questions answered as the NVIDIA-teacher student does, and teacher KL is the teacher’s own distance from the NVIDIA teacher. Gray numbers give the accuracy change from the NVIDIA-teacher student, and bold marks the best AMD teacher (smallest change for accuracy). The other-seed student changes only the data order.
Figure 5: Selective FP32 in Jais. The dashed line joins the settings that no cheaper setting beats on KL ( Table 15 ).
Table 5: Individual loss formulas for one input, before normalization and weighting.
Jais head LoRA
OLMoE head LoRA
Distillation
Trained weights
head, r=32 , α=64
head, r=32 , α=64
all student weights
Optimizer
AdamW, 10−4 , (0.9,0.999)
AdamW, 10−4 , (0.9,0.999)
AdamW, 10−5 , (0.9,0.95)
Weight decay
0
0
0
Batch size
1
1
8
Training responses
1,200
512 ∗
512
Epochs
2
2
2
Appendix
Table 6: Training hyperparameters. The normalizers cj are the mean unadapted losses over each model’s training positions. ∗ First 128 tokens of each response.
Method
MMLU (%) ↑
ArabicMMLU (%) ↑
Native AMD (BF16)
61.05
71.56
Head LoRA (KL)
61.06
71.56
Head LoRA weighting profiles
Equal
61.05
71.55
MSE-heavy
61.07
71.56
KL-heavy
61.06
71.55
Appendix
Table 7: Jais task accuracy without the questions used for adapter training or development (13,182 MMLU and 13,602 ArabicMMLU questions). Parser failures count as incorrect.
Figure 6: Share of the 131,072 OLMoE development positions at which AMD selects the same eight experts as the NVIDIA reference, by layer. Native AMD (BF16) matches the reference almost exactly at the first layer, while FP32 upcasting and FP16 lose 6.5 and 6.9 percentage points there, and remain below Native AMD (BF16) at every layer.
Path
Source instruction
Machine instruction
Count
A100, K=16
mma.sync.aligned.m16n8k16
HMMA.16816.F32.BF16
8
A100, K=8+8
mma.sync.aligned.m16n8k8
HMMA.1688.F32.BF16
16
MI210, K=8+8
v_mfma_f32_32x32x8bf16_1k
v_mfma_f32_32x32x8bf16_1k
2
Appendix
Table 8: Matrix instructions in the controlled probes, for the complete logical 32×32×16 operation. Both PTX source names end in .row.col.f32.bf16.bf16.f32 .
Figure 7: (a) Change in raw-logit MSE and forward KL relative to Native AMD (BF16). Values left of zero are closer to the NVIDIA reference, and Jais FP16 is omitted because it produces nonfinite outputs. Only FP32 upcasting on Jais moves toward the reference on both measures, and every other intervention moves away on at least one. (b) Forward KL of each distilled student from the NVIDIA-teacher student.
Numerical values
Choices (%) ↑
Accuracy (%) ↑
Method
Bit-eq. (%) ↑
Logit MSE ↓
Prob. MSE ( 10−7 ) ↓
KL ( 10−3 ) ↓
Top-1
Ord. top-5
Top-5 set
Rollout
MMLU
AMMLU/ ARC
Jais
Native AMD (BF16)
9.26
0.2873
0.279
7.08
98.76
66.53
83.91
93.50
60.86
71.36
FP16
nonfinite outputs
FP32 upcasting
–
0.1644
0.186
4.45
98.68
71.58
86.55
93.00
61.00
71.40
Head LoRA (KL)
9.22
0.2888
0.298
7.67
98.52
65.89
83.23
92.50
60.88
71.35
Appendix
Table 9: Complete output comparison for Table 2 . Choices use 2,498 positions and 200 rollouts for Jais and 131,072 positions and 128 rollouts of 1,024 tokens for OLMoE. The greedy next token on the reference input equals Top-1 and is not repeated. Jais accuracy parses generated answers on 13,890 MMLU and 14,303 ArabicMMLU questions, including the adapter’s training questions ( Table 7 excludes them), and OLMoE accuracy scores answer-letter logits on the full MMLU and ARC-Challenge test sets. – marks non-BF16 outputs.
Profile
MSE
KL
Token
Top-5
Equal
0.25
0.25
0.25
0.25
MSE-heavy
0.70
0.10
0.10
0.10
KL-heavy
0.10
0.70
0.10
0.10
Token-heavy
0.10
0.10
0.70
0.10
Top-5-heavy
0.10
0.10
0.10
0.70
Numerical-heavy
0.40
0.40
0.10
0.10
Appendix
Table 10: Loss weights after normalization.
Figure 8: Effect of loss weighting on Jais at the fixed training endpoint, relative to Native AMD (BF16) (open square). The left panel plots the change in raw-logit mean squared error (MSE) against change in ordered top-5 agreement over 2,498 positions. The right panel plots the change in forward Kullback–Leibler (KL) divergence against change in rollout agreement over 200 responses, where one response is 0.5 percentage points. The shaded region would improve both measures. No weighting profile reaches it in the left panel, and only Decision-heavy gains a response on the right while still increasing KL.
Method
Bit-eq. (%) ↑
Logit MSE ↓
KL ( 10−3 ) ↓
Ord. top-5 (%) ↑
Rollout (%) ↑
MMLU (%) ↑
AMMLU/ ARC (%) ↑
Jais
Native AMD (BF16)
9.26
0.2873
7.08
66.53
93.50
60.86
71.36
Equal
8.65
0.2957
7.21
66.37
93.50
60.87
71.34
MSE-heavy
8.72
0.2988
7.33
65.85
93.00
60.89
71.36
KL-heavy
8.95
0.2918
7.38
66.01
93.00
60.89
71.34
Token-heavy
8.07
0.3048
7.33
65.93
92.50
60.89
71.36
Appendix
Table 11: Loss-weighting comparison at the fixed training endpoint. Positions, rollouts, and accuracy sets are as in Table 9 , and the last column is ArabicMMLU for Jais and ARC-Challenge for OLMoE. Jais accuracy includes the adapter’s training questions ( Table 7 excludes them).
Map
Removed
Rank 32
11.87%
Any rank (all directions)
26.09%
Outside any linear map
73.91%
Appendix
Table 12: Share of squared Jais logit mismatch a linear head map can remove, same inputs.
Bit-equal
Logit
KL
Top-5
Router
Rollout
MMLU
ARC
Method
(%) ↑
MSE ↓
( 10−3 ) ↓
(%) ↑
(%) ↑
(%) ↑
(%)
(%)
Llama-3.1-8B-Instruct (dense)
NVIDIA reference
–
–
–
–
–
–
66.93
83.45
Native AMD (BF16)
9.52
0.00533
0.218
80.27
–
16.41
67.04
83.62
FP32 upcasting
–
0.00449
0.130
81.00
–
8.59
67.06
83.45
FP16
–
0.00586
0.151
79.33
–
10.94
67.06
83.53
Appendix
Table 13: Two further models on AMD against the NVIDIA BF16 reference. Router is the share of positions and layers where the four routed experts match. Accuracy is on the full MMLU and ARC-Challenge test sets. – marks non-BF16 outputs, a dense model, or the reference itself.
Logit MSE
KL
Top-5
Router
Rollout
AMD
NVIDIA
( 10−3 ) ↓
( 10−4 ) ↓
(%) ↑
(%) ↑
(%) ↑
BF16
BF16
18.9
2.17
70.69
90.40
34.38
FP16
BF16
22.6
2.15
68.14
87.87
35.16
FP32 upcasting
BF16
21.4
2.06
69.03
88.38
33.59
FP16
FP16
14.9
0.882
78.40
93.25
60.16
FP32 upcasting
FP32 upcasting
0.086
0.0019
99.85
99.97
98.44
Appendix
Table 14: OLMoE with the same or a different format on each vendor. Router is the share of positions and layers where the eight selected experts match.
Figure 9: Selective FP32 settings on Jais, relative to Native AMD (BF16) (open square). (a) Reduction in raw-logit MSE and forward KL, where the shaded region improves both. (b) KL reduction against forward time, with cheaper settings to the right. The dashed line joins the settings that no cheaper setting beats on KL. Keeping everything except the MLPs in FP32 (star) lies on this frontier near full upcasting at 1.30 times the forward time.
FP32 modules
Bit-equal (%) ↑
Logit MSE ↓
KL ( 10−3 ) ↓
Top-5 (%) ↑
Rollout (%) ↑
Time ( × ) ↓
Native AMD (BF16)
9.26
0.2873
7.08
66.53
93.50
1.00
Single module type
Output head
9.26
0.2865
6.87
66.21
93.50
1.01
LayerNorm
9.20
0.2522
5.81
66.89
93.50
1.02
Attention
10.00
0.2077
6.31
68.73
91.50
1.28
MLP
9.31
0.2449
7.52
66.57
91.00
2.10
Appendix
Table 15: Jais with FP32 computation in selected modules
Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on different GPUs. The root cause is floating-point non-associativity combined with hardware-dependent kernel selection. Inference frameworks select different matrix-multiplication kernels on each architecture, with different parallel reduction orders and unspecified tensor-core arithmetic, and the resulting rounding differences can flip output tokens. Existing solutions have imperfect cross-architecture reproducibility and incur a significant performance penalty. We present a solution employing a set of fixed-configuration fused-upcast GEMM kernels that load 16-bit weights from memory, upcast them to FP32 in registers, and accumulate with IEEE-754 arithmetic in a reduction order that is a pure function of the problem shape and is therefore independent of the device, its SM count, or kernel scheduling. By fixing the floating-point reduction order as a function of problem shape alone, every GPU runs the same operation sequence, so cross-architecture reproducibility of the linear layers reduces to correct IEEE-754 arithmetic rather than to rounding differences staying below a tie-flip threshold. We confirm our solution's linear-layer outputs are bitwise identical across NVIDIA Ampere, Ada, and Hopper GPUs, while running 1.17 to 3.1× faster end-to-end than the state-of-the-art solution and cutting weight-memory traffic in half.
Liam Cooper, Shinnung Jeong, Hyeran Jeon +2
Georgia Institute of Technology · University of California, Merced
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.
Gaoyuan Du, Anam Nawaz Khan, Rex Zhou +4
University of Tennessee, Knoxville · University of Chicago · Amazon +1
Small language models are cheap to serve and feasible on local hardware, but strong public 135M-class systems are commonly trained with hundreds of billions to trillions of tokens on large clusters. We study a sharply resource-constrained regime: a complete 134.5M-parameter language-model pipeline executed on one NVIDIA L20 GPU. The released checkpoint, L20-Edu-135M, receives approximately 13B pretraining tokens: 10B FineWeb-Edu tokens followed by a 3B-token educational, mathematics, code, and reasoning mixture. We document the architecture, data gates, cross-source MinHash/LSH near-deduplication, segment deduplication, benchmark-overlap removal, throughput optimization, supervised fine-tuning (SFT) with weight interpolation, and reinforcement learning from verifiable rewards (RLVR) on GSM8K. In a self-run zero-shot six-task harness, L20-Edu-135M obtains a mean score of 0.4150. It trails SmolLM-135M (0.4767) and SmolLM2-135M (0.4917), but its mean is 87.1% of SmolLM-135M's while its nominal token count is 2.17% as large. This ratio is descriptive, not evidence of statistical equivalence or a controlled scaling law. The model exceeds several older 100M-160M public baselines under the same harness. Direct GRPO-style RLVR decreases GSM8K exact-match accuracy from 1.82% to 1.59% (192-token completions) and 1.21% (320-token completions). These single-run results identify a concrete failure mode rather than establishing a general lower bound on RLVR. The contribution is an auditable resource-constrained case study, not a state-of-the-art claim.