Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms. This article focuses on two shared steps, the computation of the
X rotations,
R, and the
Y loadings,
Q, and accelerates both. For
R, term-by-term accumulation is replaced by a direct evaluation strategy that requires the same number of multiplications but parallelizes better on modern hardware. For
Q, I identify - to the best of my knowledge, for the first time - equivalences showing that each
Y loading is obtainable, up to explicitly derived constants, from quantities already computed earlier in the same iteration, and I exploit them in IKPLS to reduce the cost of each loading from
Θ(KM) to
Θ(M) operations whenever
M=1 or
2≤M<K, with
K predictor variables (number of columns in
X) and
M response variables (number of columns in
Y). Both improvements provably yield exactly the same
W,
P,
Q,
R, and
T as the original algorithms. Benchmarks with NumPy (CPU) and JAX (GPU) show speedups of up to two orders of magnitude for the isolated steps and of approximately
2× (CPU) and
6× (GPU) for entire fits. Both improvements are implemented in the free, open-source Python package \texttt{ikpls}.