Enhancing High-order Interaction Awareness in LLM-based Recommender Model
Authors: Xinfeng Wang, Jin Cui, Fumiyo Fukumoto, Yoshimi Suzuki
Organizations: Graduate School of Engineering University of Yamanashi, Kofu, Japan · Interdisciplinary Graduate School University of Yamanashi, Kofu, Japan
Large language models (LLMs) have demonstrated prominent reasoning capabilities in recommendation tasks by transforming them into text-generation tasks. However, existing approaches either disregard or ineffectively model the user-item high-order interactions. To this end, this paper presents an enhanced LLM-based recommender (ELMRec). We enhance whole-word embeddings to substantially enhance LLMs' interpretation of graph-constructed interactions for recommendations, without requiring graph pre-training. This finding may inspire endeavors to incorporate rich knowledge graphs into LLM-based recommenders via whole-word embedding. We also found that LLMs often recommend items based on users' earlier interactions rather than recent ones, and present a reranking solution. Our ELMRec outperforms state-of-the-art (SOTA) methods in both direct and sequential recommendations.
Figures & tables
Figure 1: Illustration of our motivation. In (a), LLM-based recommenders bridge users (pink) and items (green) via text prompts (blue), failing to capture high-order interactive signals. Conversely, GNNs can capture these signals, e.g., 3-hop neighbors (red arrows) in (b).
Figure 2: Illustration of input and output words.
Figure 3: Illustration of integrating interaction graph awareness into LLMs. We first leverage random feature propagation based on LightGCN to obtain whole-word embeddings, which can reflect user–item positions in the interaction graph by their semantic similarity conveyed by red, blue, and yellow edges. We then merge whole-word and word embeddings to enhance LLMs with interaction graph awareness.
Figure 4: Influence of relative positions within the interaction graph on direct and sequential recommendations.
Dataset
#User
#Item
#Review
Avg.
Density (%)
Sports
35,598
18,357
296,337
8.3
0.0453
Beauty
22,363
12,101
198,502
8.9
0.0734
Toys
19,412
11,924
167,597
8.6
0.0724
Table 1: Statistics of dataset. “#User”, “#Item”, “#Review”, and “#Avg” denote the number of users, items, reviews, and the average user interactions, respectively.
Models
Sports
Beauty
Toys
H@5
N@5
H@10
N@10
H@5
N@5
H@10
N@10
H@5
N@5
H@10
N@10
Traditional Approach
SimpleX
0.2362
0.1505
0.3290
0.1800
0.2247
0.1441
0.3090
0.1711
0.1958
0.1244
0.2662
0.1469
Large Language Model-based Approach
P5
0.1955
0.1355
0.2802
0.1627
0.1564
0.1096
0.2300
0.1332
0.1322
0.0889
0.2023
0.1114
RSL
0.2092
0.1502
0.3001
0.1703
0.1564
0.1096
0.2300
0.1332
0.1423
0.0825
0.1926
0.1028
Table 2: Performance comparison on direct recommendation. Bold : Best, underline : Second best. “*” indicates that the improvement is statistically significant ( p -value < 0.05) in the 10-trial T-test.
Models
Sports
Beauty
Toys
H@5
N@5
H@10
N@10
H@5
N@5
H@10
N@10
H@5
N@5
H@10
N@10
Traditional Approach
Caser
0.0116
0.0072
0.0194
0.0097
0.0205
0.0131
0.0347
0.0176
0.0166
0.0107
0.0270
0.0141
GRU4Rec
0.0129
0.0086
0.0204
0.0110
0.0164
0.0099
0.0283
0.0137
0.0097
0.0059
0.0176
0.0084
HGN
0.0189
0.0120
0.0313
0.0159
0.0325
0.0206
0.0512
0.0266
0.0321
0.0221
0.0497
0.0277
SASRec
0.0233
0.0154
0.0350
0.0192
0.0387
0.0249
0.0605
0.0318
0.0463
0.0306
0.0675
0.0374
Table 3: Performance comparison between ELMRec and baselines in the sequential recommendation task.
Models
Sports
Beauty
Toys
H@10
N@10
H@10
N@10
H@10
N@10
Direct recommendation
w/o Text Prompt
0.1270
0.1006
0.0381
0.0296
0.0190
0.0180
w/o Graph-aware
0.2890
0.1783
0.2687
0.1650
0.2141
0.1243
ELMRec
0.6479
0.4852
0.6794
0.4973
0.6045
0.4141
Impv.
124.2%
172.1%
152.8%
201.4%
182.3%
233.1%
Table 4: Ablation study. “w/o Graph-aware” and “w/o Reranking” denote the ELMRec without the interaction graph awareness and reranking approach, respectively. “w/o Text Prompt” indicates that only whole-word embeddings are fed into the LLM for recommendations.
Models
Sports
Beauty
Toys
H@10
N@10
H@10
N@10
H@10
N@10
Direct recommendation
w/o Graph-aware
0.2890
0.1783
0.2687
0.1650
0.2141
0.1243
Prepending
0.3456
0.2174
0.2730
0.1675
0.2362
0.1309
Addition
0.6479
0.4852
0.6794
0.4973
0.6045
0.4141
Table 5: Results by various methods of incorporating graph-aware whole-word embeddings. “Prepending” refers to prepending directly whole word embedding before the input sequence. “Addition” indicates adding graph-aware embeddings to ID tokens.
Whole-word
Sports
Beauty
Toys
Embeddings
H@10
N@10
H@10
N@10
H@10
N@10
Direct recommendation
- Random
0.2758
0.1736
0.2850
0.1752
0.2127
0.1240
- Incremental
0.2989
0.1820
0.2881
0.1753
0.2207
0.1334
- Graph-aware
0.6479
0.4852
0.6794
0.4973
0.6045
0.4141
Sequential recommendation
Table 6: Effect of various whole-word embeddings. “Graph-aware” represents the interaction graph-aware whole-word embeddings. “Incremental” indicates that the graph-aware whole-word embeddings are replaced with incremental embeddings. The result of “Random” is obtained by disordering the indices for incremental whole-word embeddings.
Figure 5: Effect of α . The x-axis and the y-axis indicate the values of α and NDCG@10 (%), respectively.
Hyperprameter
Sports
Beauty
Toys
H@10
N@10
H@10
N@10
H@10
N@10
N = 0
0.0599
0.0435
0.0703
0.0471
0.0737
0.0572
N = 5
0.0614
0.0469
0.0744
0.0528
0.0763
0.0617
N = 10
0.0616
0.0471
0.0748
0.0528
0.0764
0.0618
N = 15
0.0616
0.0471
0.0750
0.0529
0.0764
0.0618
Table 7: Effect of N for reranking candidates.
Figure 6: Effect of σ and L in direct recommendations.
Figure 7: Visualization of whole-word embeddings of users and items after 1st and 3rd rounds of propagation. The dots in the same color denote users (or items) who have interacted with the same items (or users). The closer the dots, the greater their similarity. Further visualization results are provided in the Appendix. A.1.3 .
Datasets
Stages
Pre-training
DirRec
SeqRec
Sports
04h06m58s
17m07s
07m38s
Beauty
02h08m48s
10m32s
04m49s
Toys
03h56m09s
12m44s
04m04s
Table 8: Running time in various stages on three datasets. “DirRec” and “SeqRec” denote the cumulative time cost in direct and sequential recommendations, respectively. “h”, “m”, and “s” indicate “hours” and “minutes”, and “seconds” respectively.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
DirRec
SeqRec
α
σ
N
L
α
σ
N
L
Sports
5
5
10
4
1
5
10
4
Beauty
9
6
15
4
6
6
15
4
Toys
11
5
10
4
9
5
10
4
Appendix
Table 9: Best values of hyperparameter for the three datasets. “DirRec” and “SeqRec” denote direct and sequential recommendations, respectively.
Figure 8: Visualization of user and item whole-word embeddings at each round of random feature propagation. The dots in the same color denote users who have interacted with the same items or the items that are interacted with by the same user. Similar users and items are close to each other in the distribution.
Figure 9: Illustration of reranking approach. The gray nodes such as i4 and i5 indicate the user’ interacted items. The red arrows refer to the reranking processes.