Despite the adoption of large language models (LLMs) in recommendation systems, prevailing approaches mostly model single-type behaviors (e.g., views or purchases). Even when incorporating multiple behaviors, existing methods flatten heterogeneous actions into homogeneous token sequences, ignoring their distinct decision-making roles. This flattening fails to capture semantic hierarchies and contextual nuances in complex decision-making, such as trade-offs between price and quality. Consequently, performance degrades in critical ``difficult-choice'' scenarios involving highly similar items. To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence. MARI maintains a Decision Memory Bank (DMB) that archives users' past rationales as Structured Decision Memories (SDMs): concise records of goals, constraints, and trade-offs. These SDMs are generated offline via Post-Hoc Decision Distillation from heterogeneous behaviors and user-generated content. By retrieving relevant SDMs to augment LLM reasoning, MARI achieves interpretability and scalability without the prohibitive cost of processing long raw sequences. Extensive experiments show MARI significantly outperforms state-of-the-art baselines on standard next-item prediction and a newly introduced Difficult Choice Prediction task, incurring low latency overhead by decoupling memory construction from online inference. Qualitative analyses reveal actionable, human-readable insights into user decision-making, marking a concrete step toward reasoning-aware recommendation systems.
Figures & tables
Fig. 1: Workflow comparison. Top: Directly feeding long raw behavior sequences into an LLM introduces context limits, noise, and latent decision factors, hindering accurate reasoning. Bottom: MARI retrieves Structured Decision Memories (SDMs) from a Decision Memory Bank (DMB) based on recent behaviors. Augmenting the LLM with these explicit rationales (goals, constraints, trade-offs) enables interpretable, grounded reasoning and improves recommendation accuracy.
Fig. 2: Effect of input sequence length on Hit@1 for NIP and DCP tasks (using DeepSeek-R1).
Fig. 3: Algorithmic framework of MARI. The offline stage constructs DMB via PHDD; during online inference, MARI retrieves relevant SDMs and fuses them to prompt the LLM for interpretable recommendation.
Metric
Q4B
Q8B
Q14B
Q32B
R1
Q4B-PHDD(Ours)
Precision ( ↑ )
Factor Key
0.577
0.654
0.636
0.631
0.779
0.769
Decision Prof.
0.467
0.470
0.525
0.556
0.725
0.713
Decision Effic.
0.483
0.473
0.529
0.513
0.707
0.693
LLM-as-Judge ( ↑ )
Struct. Corr.
3.56
3.31
3.58
3.71
4.13
4.12
TABLE I: Quality of Structured Decision Memory (SDM) generation on the PHDD dataset. Models : Q4B (Qwen3-4B), Q8B (Qwen3-8B), Q14B, Q32B. Bold : Best, Underlined : Second-best.
Fig. 4: The role of structured decision memory: comparing recommendation rationales with and without decision memory augmentation.
Category
Model
Base
MARI † (Ours, w/o SFT)
NIP-in
NIP-out
DCP
NIP-in
NIP-out
DCP
ID-based (w/ train)
SASRec MB
0.1666
0.1697
0.4077
-
-
-
MULE
0.2363
0.1968
0.4206
-
-
-
COPF
0.2013
0.1795
0.4037
-
-
-
HGIB
0.2489
0.2251
0.4384
-
-
-
LLMs (w/o train)
Qwen3-4B
0.0751
0.0225
0.4617
0.1278
0.0157
0.4751
TABLE II: Full comparison of Hit@1. MARI † denotes an off-the-shelf open-source general-purpose model without any task-specific fine-tuning for recommendation. Base denotes vanilla LLMs and ID-based models. Bold : Best, Underlined : Second-best.
Variant
Q4B
Q14B
Q32B
GPT-OSS-20B
R1
NIP-in
Base
0.075
0.100
0.114
0.148
0.196
MARI-B (Raw Behav.)
0.105
0.125
0.127
0.132
0.247
MARI †⋆ (Undist. SDM)
0.110
0.134
0.157
0.170
0.249
MARI † (Dist. SDM)
0.128
0.154
0.174
0.196
0.253
SFT w/o SDM
0.458
-
-
-
-
TABLE III: Ablation study on Hit@1. MARI-B : Uses raw behavior sequences aligned with SDMs. MARI †⋆ : Uses undistilled SDMs generated by Qwen3-4B (w/o PHDD). MARI † : Uses distilled SDMs generated by Qwen3-4B (w/ PHDD). SFT variants : Compares standard SFT against SFT with SDMs. Bold indicates the best (MARI † ) and overall (MARI) results.
Fig. 5: MARI † ’s Hit@1 performance across different tasks as K increases.
Variant
Q4B
Q14B
Q32B
GPT-OSS
R1
NIP-in
Base
0.075
0.100
0.114
0.148
0.196
Err. Ret.
0.079
0.124
0.124
0.158
0.230
Err. Mem.
0.094
0.120
0.117
0.151
0.223
MARI †
0.128
0.154
0.174
0.196
0.253
NIP-out
TABLE IV: Memory Error Analysis. Err. Ret. : Lowest-similarity memories. Err. Mem. : Random same-category SDMs from other users.