cs.LGOct 8, 2026

Memorization and Malign Generalization in Conditional Diffusion Models with Random Features

Authors: Gwangho Kim, Sungyoon Lee

Organizations: Department of Computer Science Hanyang University, Seoul

Abstract

Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding diffusion models requires rethinking (again) generalization

    May 7, 2026Pierre Marion, Yu-Han WuMemorization in Generative ModelsDiffusion Models

  2. Smoothing the Score Function to Enhance Generalization in Diffusion Models

    Date pendingXinyu Zhou, Jiawei Zhang, Stephen J. WrightMemorization in Generative ModelsGenerative Modeling

  3. Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

    Apr 29, 2026Bao Pham, Mohammed J. Zaki, Luca Ambrogioni +2Memorization in Language ModelsNeural Network Memorization