cs.LGSep 23, 2026

CARE: Condition-Aware Representation Regularization for Diffusion Models

Authors: Fengjia Guo, Zhuoyi Yang, Jie Tang

Organizations: Department of Computer Science and Technology, Tsinghua University, Beijing, China.

Abstract

Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08% reduction in FID in 400k training steps, leading to a 3.5×\times speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

    May 26, 2026Nithesh Chandher Karthikeyan, Jonas Unger, Gabriel EilertsenText-Conditioned Diffusion ModelMulti-Reference Image Generation

  2. DICE: Distilling Classifier-Free Guidance into Text Embeddings

    Feb 6, 2025Zhenyu Zhou, Defang Chen, Can Wang +2Text-To-Image Diffusion ModelsMulti-Reference Image Generation

  3. Representation-Conditioned Diffusion Models for Guided Training Data Generation

    May 26, 2026Nithesh Chandher Karthikeyan, Jonas Unger, Gabriel EilertsenText-Conditioned Diffusion ModelData Augmentation