Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power
r≥2, on product distributions over a scaled sign cube in dimension at least linear in
n, every proper learner with expected excess risk at most
ε on these distributions at the optimal sample size
n≍ε−2+2/r has worst-case ld-CMI of order
n bits, and
Θ(n/(1+(τ/ε)2)) bits under Gaussian noise of standard deviation
τ on the loss differences. The same holds without a regularizer, at
n≍ε−2. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is
O(n−1/2). We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.