cs.LGOct 4, 2026

Measuring and Reducing Cross-Vendor Mismatch in Language Models

Authors: Erland Hilman Fuadi, Chong Tian, Xiaosong Ma, Qirong Ho

Organizations: Mohamed bin Zayed University of Artificial Intelligence

Abstract

Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with five metric families, namely bitwise equality, logit differences, top-K consistency, token agreement, and task accuracy. We trace one source of the mismatch to accumulation order inside vendors' matrix instructions. Upcasting to FP32 reduces the dense model's logit error by 43% at three times the runtime, yet keeping only the MLPs in BF16 retains 94% of this gain at 1.3 times the runtime, so most of the cost of full upcasting buys little. In the MoE models, FP32 and FP16 both lower the probability error but raise the logit error and change expert selection, and FP16 fails in the dense model. An output-head low-rank adapter (LoRA) does not help either, since the final hidden state does not predict the mismatch. The mismatch also carries into training. With every seed fixed, a student distilled from a teacher running on AMD answers 431 MMLU questions differently from one distilled from the same teacher on NVIDIA. Under FP32 upcasting, bitwise equality barely changes while the output distributions move most of the way to the reference, so judging cross-vendor agreement by a single measure misreads both its cost and its gains. Code is available at https://github.com/crova-project/crova.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

    Sep 22, 2026Liam Cooper, Shinnung Jeong, Hyeran Jeon +2LLM Inference OptimizationFloating-Point

  2. Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

    Sep 22, 2026Gaoyuan Du, Anam Nawaz Khan, Rex Zhou +4Constrained DecodingLLM Inference Optimization