cs.IRSep 28, 2026

Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems

Authors: Cong Wang, Shoujin Wang, Yishuo Li, Qi Zhang, Liang Hu, Wenpeng Lu

Organizations: Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, Shandong, China · University of Technology Sydney, Sydney, NSW, Australia · Department of Computer Science and Technology, Tongji University, Shanghai, China · Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science; Shandong Academy of Artificial Intelligence, Jinan, Shandong, China

Abstract

Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.

Figures & tables

Explore similar work

CardsList
  1. FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models

    Sep 29, 2026Song-Li Wu, Xianquan Wang, Zhaocheng Du +2Real-World Content Recommendation ProblemDiffusion Models

  2. Time-Aware Diffusion based on Preference Disentanglement for Generative Recommendation

    Jun 1, 2026Bangguo Zhu, Peng Huo, Yuanbo Zhao +3Real-World Content Recommendation ProblemPreference Alignment Learning

  3. Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

    Jul 1, 2026Zhuoxuan Zhang, Kangqi Ni, Yuhang Chen +12Large Language Model RerankingCross-Encoder Reranking