Organizations: Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, Shandong, China · University of Technology Sydney, Sydney, NSW, Australia · Department of Computer Science and Technology, Tongji University, Shanghai, China · Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science; Shandong Academy of Artificial Intelligence, Jinan, Shandong, China
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.
Figures & tables
Figure 1 . (a) The number of papers published annually on diffusion-based recommender systems. (b) The distribution of diffusion-based recommender systems across different venues, showing only the top 10 venues.
Framework
Recommendation Scenario Coverage
Diffusion-aware Support
CF
Seq.
MM
POI
CDR
Process Config
Sampling Protocol
Diffusion Diagnostics
ReChorus ( Li et al., 2024a )
✓
✓
✗
✗
✗
✗
✗
✗
DaisyRec ( Sun et al., 2022 )
✓
✗
✗
✗
✗
✗
✗
✗
RecBole ( Zhao et al., 2021 )
✓
✓
✓
✗
✓
▲
▲
✗
Ludewig & Jannach ( Ludewig and Jannach, 2018 )
✓
✓
✗
✗
✗
✗
✗
✗
NewsRecLib ( Iana et al., 2023 )
✗
✓
✓
✗
✗
✗
✗
✗
Table 1 . Comparison of representative recommendation frameworks and evaluation studies. ✓, ▲ , and ✗denote native, partial or model-specific, and no explicitly documented support, respectively. CF, Seq., MM, POI, and CDR denote collaborative filtering, sequential, multimodal, point-of-interest, and cross-domain recommendation. Diffusion-aware support covers process configuration, sampling protocols, and diagnostic analyses.
Figure 2 . Overview of Eval4DiRec. The framework standardizes the end-to-end evaluation pipeline for diffusion-based recommender systems, including data preprocessing/splitting, model integration, training and inference, and metric reporting. It consists of four modules: Data prepares unified inputs and scenario-specific splits; DiRec integrates diffusion-based recommenders and manages sampling-related configurations; Execution orchestrates training, validation, and testing with structured outputs; and Utils provides utilities for configuration, randomness control, logging, and metric computation to support fair and reproducible comparisons.
Dataset
#User
#Item
#Interactions
Sparsity
Yelp
37,232
47,744
359,330
99.98%
Beauty
9,895
8,455
73,528
99.91%
ML-1M
5,652
2,468
223,456
98.40%
TikTok
9,319
6,710
59,541
99.90%
Baby
19,445
7,050
139,110
99.89%
Sports
35,598
18,357
256,308
99.96%
Table 2 . Statistics of processed datasets used in Eval4DiRec
Scenario
Model
Reason for Selection
Collaborative Filtering
CODIGEM ( Walker et al., 2022 )
An early diffusion-based model for reconstructing user–item interaction vectors.
DiffRec ( Wang et al., 2023b )
A representative interaction-vector diffusion model for collaborative filtering.
CF-Diff ( Hou et al., 2024 )
Represents the use of high-order collaborative signals in diffusion-based RSs.
DDRM ( Zhao et al., 2024 )
Represents diffusion-based denoising of user and item representations.
Sequential Recommendation
DiffuRec ( Li et al., 2023a )
An early diffusion-based model designed for sequential recommendation.
DiffuASR ( Liu et al., 2023 )
Represents the use of diffusion models for sequence augmentation.
Table 3 . Summary and selection rationale of the 14 diffusion-based recommendation models implemented in Eval4DiRec.
Method
Recall
NDCG
MRR
HR
Recall
NDCG
MRR
HR
Recall
NDCG
MRR
HR
CF
Yelp
Beauty
ML-1M
BPR-MF†
1.81
1.03
1.12
6.71
4.49
2.73
2.84
9.43
9.75
6.63
11.57
42.16
LightGCN†
3.31
1.79
2.14
9.09
6.32
3.55
3.61
11.86
12.31
8.74
14.06
48.28
CODIGEM
4.12
2.24
2.61
10.93
6.90
4.02
4.09
13.67
15.56
10.97
17.07
57.46
DiffRec
4.82*
2.66*
3.16*
12.69*
7.13
4.08
4.06
14.04
15.67*
11.05*
18.49*
58.02*
CF-Diff
3.36
1.91
2.34
9.28
7.18
3.94
3.89
14.21
15.52
10.76
17.62
57.72
Table 4. Overall Performance Comparison. Different recommendation tasks are evaluated on their respective datasets. All metrics are reported @20 and shown as percentages (%). Bold and underlined values indicate the best and second-best results on each dataset, respectively. Results marked with * are statistically significant based on a paired t -test ( p<0.05 ). Models marked with † are non-diffusion baselines.
Figure 3 . Efficiency Analysis. The left axis (Red bars) represents the Training Time per epoch, while the right axis (Blue hashed bars) denotes the Total Inference Time. Note that absolute inference time is dataset-dependent (e.g., number of users/items), thus we focus on within-dataset comparisons.
Figure 4 . Impact of noise schedule on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 5 . Impact of diffusion steps on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 6 . Impact of noise distribution on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 7 . Robustness to missing data: performance under different missing ratios across three datasets (Recall@20).
Figure 8 . Robustness to noisy data: performance under different noise ratios across three datasets (Recall@20).
Figure 9 . Long-tail item analysis across different popularity groups on three datasets (Recall@20).