Despite the adoption of large language models (LLMs) in recommendation systems, prevailing approaches mostly model single-type behaviors (e.g., views or purchases). Even when incorporating multiple behaviors, existing methods flatten heterogeneous actions into homogeneous token sequences, ignoring their distinct decision-making roles. This flattening fails to capture semantic hierarchies and contextual nuances in complex decision-making, such as trade-offs between price and quality. Consequently, performance degrades in critical ``difficult-choice'' scenarios involving highly similar items. To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence. MARI maintains a Decision Memory Bank (DMB) that archives users' past rationales as Structured Decision Memories (SDMs): concise records of goals, constraints, and trade-offs. These SDMs are generated offline via Post-Hoc Decision Distillation from heterogeneous behaviors and user-generated content. By retrieving relevant SDMs to augment LLM reasoning, MARI achieves interpretability and scalability without the prohibitive cost of processing long raw sequences. Extensive experiments show MARI significantly outperforms state-of-the-art baselines on standard next-item prediction and a newly introduced Difficult Choice Prediction task, incurring low latency overhead by decoupling memory construction from online inference. Qualitative analyses reveal actionable, human-readable insights into user decision-making, marking a concrete step toward reasoning-aware recommendation systems.
Figures & tables
Fig. 1: Workflow comparison. Top: Directly feeding long raw behavior sequences into an LLM introduces context limits, noise, and latent decision factors, hindering accurate reasoning. Bottom: MARI retrieves Structured Decision Memories (SDMs) from a Decision Memory Bank (DMB) based on recent behaviors. Augmenting the LLM with these explicit rationales (goals, constraints, trade-offs) enables interpretable, grounded reasoning and improves recommendation accuracy.
Fig. 2: Effect of input sequence length on Hit@1 for NIP and DCP tasks (using DeepSeek-R1).
Fig. 3: Algorithmic framework of MARI. The offline stage constructs DMB via PHDD; during online inference, MARI retrieves relevant SDMs and fuses them to prompt the LLM for interpretable recommendation.
Metric
Q4B
Q8B
Q14B
Q32B
R1
Q4B-PHDD(Ours)
Precision ( ↑ )
Factor Key
0.577
0.654
0.636
0.631
0.779
0.769
Decision Prof.
0.467
0.470
0.525
0.556
0.725
0.713
Decision Effic.
0.483
0.473
0.529
0.513
0.707
0.693
LLM-as-Judge ( ↑ )
Struct. Corr.
3.56
3.31
3.58
3.71
4.13
4.12
TABLE I: Quality of Structured Decision Memory (SDM) generation on the PHDD dataset. Models : Q4B (Qwen3-4B), Q8B (Qwen3-8B), Q14B, Q32B. Bold : Best, Underlined : Second-best.
Fig. 4: The role of structured decision memory: comparing recommendation rationales with and without decision memory augmentation.
Category
Model
Base
MARI † (Ours, w/o SFT)
NIP-in
NIP-out
DCP
NIP-in
NIP-out
DCP
ID-based (w/ train)
SASRec MB
0.1666
0.1697
0.4077
-
-
-
MULE
0.2363
0.1968
0.4206
-
-
-
COPF
0.2013
0.1795
0.4037
-
-
-
HGIB
0.2489
0.2251
0.4384
-
-
-
LLMs (w/o train)
Qwen3-4B
0.0751
0.0225
0.4617
0.1278
0.0157
0.4751
TABLE II: Full comparison of Hit@1. MARI † denotes an off-the-shelf open-source general-purpose model without any task-specific fine-tuning for recommendation. Base denotes vanilla LLMs and ID-based models. Bold : Best, Underlined : Second-best.
Variant
Q4B
Q14B
Q32B
GPT-OSS-20B
R1
NIP-in
Base
0.075
0.100
0.114
0.148
0.196
MARI-B (Raw Behav.)
0.105
0.125
0.127
0.132
0.247
MARI †⋆ (Undist. SDM)
0.110
0.134
0.157
0.170
0.249
MARI † (Dist. SDM)
0.128
0.154
0.174
0.196
0.253
SFT w/o SDM
0.458
-
-
-
-
TABLE III: Ablation study on Hit@1. MARI-B : Uses raw behavior sequences aligned with SDMs. MARI †⋆ : Uses undistilled SDMs generated by Qwen3-4B (w/o PHDD). MARI † : Uses distilled SDMs generated by Qwen3-4B (w/ PHDD). SFT variants : Compares standard SFT against SFT with SDMs. Bold indicates the best (MARI † ) and overall (MARI) results.
Fig. 5: MARI † ’s Hit@1 performance across different tasks as K increases.
Variant
Q4B
Q14B
Q32B
GPT-OSS
R1
NIP-in
Base
0.075
0.100
0.114
0.148
0.196
Err. Ret.
0.079
0.124
0.124
0.158
0.230
Err. Mem.
0.094
0.120
0.117
0.151
0.223
MARI †
0.128
0.154
0.174
0.196
0.253
NIP-out
TABLE IV: Memory Error Analysis. Err. Ret. : Lowest-similarity memories. Err. Mem. : Random same-category SDMs from other users.
Large Language Models (LLMs) have shown strong potential for recommendation by leveraging their semantic understanding and contextual modeling capabilities. Recent studies further introduce reasoning mechanisms to improve user preference modeling. However, explicit natural-language reasoning incurs substantial inference overhead, whereas existing latent reasoning methods mainly focus on generating or verifying intermediate states, leaving their layer-wise preference roles and contributions insufficiently characterized. We propose HiLaR, a Hierarchical Latent Reasoning framework with layer-aware reinforcement optimization for LLM-based recommendation. HiLaR constructs temporal-guided hierarchical user preference representations, aligns them with multiple LLM latent reasoning states, and organizes the reasoning process from broad preferences to fine-grained current intents. To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state. Experiments on four Amazon benchmark datasets show that HiLaR generally outperforms strong sequential, generative, and LLM-based recommendation baselines. Ablation and sensitivity analyses further verify the contribution of hierarchical representation learning, latent alignment, and process-level optimization. Our code is available in https://github.com/hupeiyu21/HiLaR.
Peiyu Hu, Siying Gu, Weihai Lu +8
1Xi’an Jiaotong-Liverpool University · 2Xiaohongshu · 3Peking University +1
Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and natural-language reasoning abilities. Despite recent progress, current LLM-based recommenders still face key challenges in constructing decision-relevant contexts from heterogeneous evidence. First, existing methods often rely on fixed context construction strategies: collaborative behavioral evidence and item-side metadata are typically incorporated through predefined prompts, static retrieval pipelines, or handcrafted injection mechanisms, making it difficult to determine what information is truly beneficial for each instance. Second, heterogeneous evidence introduces a severe context-efficiency bottleneck. Rich metadata and collaborative interaction records can quickly overwhelm the context window, while aggressive compression or heuristic filtering may discard fine-grained evidence critical for accurate recommendation. To address these challenges, we propose RRCM, a ranking-driven retrieval-and-reasoning framework over collaborative and metadata memories for LLM-based agentic recommendation. RRCM starts from a lightweight user-history context and learns whether to recommend directly, retrieve collaborative evidence, retrieve item metadata, or interleave both through reasoning. Both memories are represented in natural language and accessed through a unified retrieval interface, enabling flexible evidence acquisition without handcrafted CF injection or fixed retrieval rules. We optimize this memory-reading policy with an outcome-only ranking reward, instantiated using group relative policy optimization, so that retrieval decisions are directly driven by final top-k recommendation quality. Extensive experiments show that RRCM significantly outperforms traditional baselines and diverse LLM-based recommendation approaches.
Shijun Li, Pranav Belligundu, Tianxin Wei +3
The University of Texas at Austin · University of Illinois at Chicago · Capital One AI Foundations +1
Sequential recommender systems have achieved significant success in modeling temporal user behavior but remain limited in capturing rich user semantics beyond interaction patterns. Large Language Models (LLMs) present opportunities to enhance user understanding with their reasoning capabilities, yet existing integration approaches create prohibitive inference costs in real time. To address these limitations, we present a novel knowledge distillation method that utilizes textual user profile generated by pre-trained LLMs into sequential recommenders without requiring LLM inference at serving time. The resulting approach maintains the inference efficiency of traditional sequential models while requiring neither architectural modifications nor LLM fine-tuning.