Organizations: Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, Shandong, China · University of Technology Sydney, Sydney, NSW, Australia · Department of Computer Science and Technology, Tongji University, Shanghai, China · Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science; Shandong Academy of Artificial Intelligence, Jinan, Shandong, China
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.
Figures & tables
Figure 1 . (a) The number of papers published annually on diffusion-based recommender systems. (b) The distribution of diffusion-based recommender systems across different venues, showing only the top 10 venues.
Framework
Recommendation Scenario Coverage
Diffusion-aware Support
CF
Seq.
MM
POI
CDR
Process Config
Sampling Protocol
Diffusion Diagnostics
ReChorus ( Li et al., 2024a )
✓
✓
✗
✗
✗
✗
✗
✗
DaisyRec ( Sun et al., 2022 )
✓
✗
✗
✗
✗
✗
✗
✗
RecBole ( Zhao et al., 2021 )
✓
✓
✓
✗
✓
▲
▲
✗
Ludewig & Jannach ( Ludewig and Jannach, 2018 )
✓
✓
✗
✗
✗
✗
✗
✗
NewsRecLib ( Iana et al., 2023 )
✗
✓
✓
✗
✗
✗
✗
✗
Table 1 . Comparison of representative recommendation frameworks and evaluation studies. ✓, ▲ , and ✗denote native, partial or model-specific, and no explicitly documented support, respectively. CF, Seq., MM, POI, and CDR denote collaborative filtering, sequential, multimodal, point-of-interest, and cross-domain recommendation. Diffusion-aware support covers process configuration, sampling protocols, and diagnostic analyses.
Figure 2 . Overview of Eval4DiRec. The framework standardizes the end-to-end evaluation pipeline for diffusion-based recommender systems, including data preprocessing/splitting, model integration, training and inference, and metric reporting. It consists of four modules: Data prepares unified inputs and scenario-specific splits; DiRec integrates diffusion-based recommenders and manages sampling-related configurations; Execution orchestrates training, validation, and testing with structured outputs; and Utils provides utilities for configuration, randomness control, logging, and metric computation to support fair and reproducible comparisons.
Dataset
#User
#Item
#Interactions
Sparsity
Yelp
37,232
47,744
359,330
99.98%
Beauty
9,895
8,455
73,528
99.91%
ML-1M
5,652
2,468
223,456
98.40%
TikTok
9,319
6,710
59,541
99.90%
Baby
19,445
7,050
139,110
99.89%
Sports
35,598
18,357
256,308
99.96%
Table 2 . Statistics of processed datasets used in Eval4DiRec
Scenario
Model
Reason for Selection
Collaborative Filtering
CODIGEM ( Walker et al., 2022 )
An early diffusion-based model for reconstructing user–item interaction vectors.
DiffRec ( Wang et al., 2023b )
A representative interaction-vector diffusion model for collaborative filtering.
CF-Diff ( Hou et al., 2024 )
Represents the use of high-order collaborative signals in diffusion-based RSs.
DDRM ( Zhao et al., 2024 )
Represents diffusion-based denoising of user and item representations.
Sequential Recommendation
DiffuRec ( Li et al., 2023a )
An early diffusion-based model designed for sequential recommendation.
DiffuASR ( Liu et al., 2023 )
Represents the use of diffusion models for sequence augmentation.
Table 3 . Summary and selection rationale of the 14 diffusion-based recommendation models implemented in Eval4DiRec.
Method
Recall
NDCG
MRR
HR
Recall
NDCG
MRR
HR
Recall
NDCG
MRR
HR
CF
Yelp
Beauty
ML-1M
BPR-MF†
1.81
1.03
1.12
6.71
4.49
2.73
2.84
9.43
9.75
6.63
11.57
42.16
LightGCN†
3.31
1.79
2.14
9.09
6.32
3.55
3.61
11.86
12.31
8.74
14.06
48.28
CODIGEM
4.12
2.24
2.61
10.93
6.90
4.02
4.09
13.67
15.56
10.97
17.07
57.46
DiffRec
4.82*
2.66*
3.16*
12.69*
7.13
4.08
4.06
14.04
15.67*
11.05*
18.49*
58.02*
CF-Diff
3.36
1.91
2.34
9.28
7.18
3.94
3.89
14.21
15.52
10.76
17.62
57.72
Table 4. Overall Performance Comparison. Different recommendation tasks are evaluated on their respective datasets. All metrics are reported @20 and shown as percentages (%). Bold and underlined values indicate the best and second-best results on each dataset, respectively. Results marked with * are statistically significant based on a paired t -test ( p<0.05 ). Models marked with † are non-diffusion baselines.
Figure 3 . Efficiency Analysis. The left axis (Red bars) represents the Training Time per epoch, while the right axis (Blue hashed bars) denotes the Total Inference Time. Note that absolute inference time is dataset-dependent (e.g., number of users/items), thus we focus on within-dataset comparisons.
Figure 4 . Impact of noise schedule on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 5 . Impact of diffusion steps on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 6 . Impact of noise distribution on diffusion-based recommendation performance on three datasets (Recall@20).
Figure 7 . Robustness to missing data: performance under different missing ratios across three datasets (Recall@20).
Figure 8 . Robustness to noisy data: performance under different noise ratios across three datasets (Recall@20).
Figure 9 . Long-tail item analysis across different popularity groups on three datasets (Recall@20).
While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinforcing amplification of popularity bias. We identify that this phenomenon is driven by two compounding mechanisms. First, while optimization loss is universally dominated by high-frequency items across recommenders, DRMs suffer from a unique structural prior mismatch during generation. Because the forward terminal distribution of long-tailed data deviates significantly from the standard Gaussian prior, reverse sampling trajectories inherently collapse toward high-density popular items, fundamentally suppressing niche item generation. To dismantle this self-reinforcing loop, we propose FairDiff, a plug-and-play fairness-aware diffusion framework. To overcome the popularity-dominated loss, we introduce Popularity Condition Guidance (PCG). Rather than altering the training objective, PCG acts as an inference-time distributional reweighting mechanism, mathematically reshaping the score-based gradient field to penalize high-popularity regions and guide trajectories toward niche semantics. Furthermore, we design a Semantic Calibration (SC) Module to bridge the prior mismatch, aligning the forward and reverse distributions via one-step optimal transport. Comprehensive evaluations demonstrate that FairDiff achieves state-of-the-art performance while effectively mitigating the self-reinforcing Matthew Effect, highlighting its value as a general framework for DRMs.
Song-Li Wu, Xianquan Wang, Zhaocheng Du +2
Tsinghua University · University of Science and Technology of China · Huawei Noah’s Ark Lab
Recently, Generative Recommenders (GRs) have emerged as a transformative recommendation paradigm by replacing traditional item IDs with semantic indices (SIDs). Owing to the exceptional generative capabilities of diffusion models, a few pioneering works explore developing GRs with diffusion architectures as the backbone. However, a fatal limitation of existing diffusion-based GRs is that the diffusion process applies uniformly to all items within the historical interactions. In contrast, the user preference is shaped by multifaceted time-evolving factors and thus exhibits a non-stationary distribution in the temporal aspect. To bridge this gap, this study proposes a novel GR framework, named TDPM, by designing the time-aware diffusion on SID tokens. Specifically, TDPM explicitly integrates the impact of time-evolving user preferences into the diffusion process. In detail, the user preference is disentangled into (i) the period preference, which remains consistent over a long time-span, and (ii) the point preference, which is triggered by recent focal events. Extensive experiments on three public real-world datasets demonstrate the significant superiority of TDPM over the state-of-the-art baselines. TDPM achieves average improvements of up to 29.21% and 25.45% in terms of HR@20 and NDCG@20, respectively. The ablation study further underscores the necessity of time-aware token diffusion in diffusion-based GRs.
Bangguo Zhu, Peng Huo, Yuanbo Zhao +3
Central South University Changsha, China · National Super Computing Center Tianjin, China · Renmin University of China Beijing, China +1
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by 2.4--3.5× at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.