Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16% across three recommendation agent tasks, namely searching, ranking, and judging.
Figures & tables
Figure 1. Motivation of ReMem. (a) Conventional RecAgents read the Web as text: they extract noisy HTML into long contexts and rely on LLMs to recover user intent, which is brittle across modalities and inefficient for long-context reasoning. (b) The proposed ReMem follows a more human-like principle: see item pages through OCR-based multimodal perception and memorize only evolving preference abstractions through dynamic memory, enabling robust modeling of time-evolving user preferences.
Acc. (%)
Qwen3.5
Qwen3VL
4B
9B
Avg.
4B
8B
Avg.
Baseline
49.13
54.27
51.70
49.32
50.94
50.13
+ 1)
41.51
45.27
43.39
52.51
55.77
54.14
+ 2)
-
-
-
46.38
48.56
47.47
+ 3)
57.02
60.25
58.64
56.62
58.60
57.61
Table 1. The effect of different perception strategies. Strategy 3) parses multimodal web pages into informative textual descriptions through OCR perception, achieving superior performance compared with the other strategies.
Figure 2. Accuracy scores of different long-context reasoning strategies on HotpotQA. Even models that use long-context continual pretraining or vectorized memory techniques fail to maintain consistent performance. In contrast, the agent equipped with dynamic memory demonstrates relative lossless performance extrapolation.
Figure 3. Overview of the proposed ReMem framework. Inspired by how humans browse item pages while maintaining time-evolving memory, ReMem consists of three key components. a) Instead of relying on raw HTML, ReMem adopts an OCR-based perception module that represents item pages through a unified multimodal view of natural text, images, and interactive elements, better matching users’ browsing behavior. b) User interaction history is processed as a sequential stream of chunks, from which the model maintains a compact, fixed-length memory that is continuously updated to capture evolving user preferences. c) To teach the model what to remember and how to update memory in multi-turn interactions, ReMem employs Multi-Mem GRPO, which propagates the final-answer advantage to all intermediate memories that contribute to the final response.
Dataset
#Users
#Items
#Int.
Avg. Seq.
#Tokens (Avg. | Max.)
Games
94,762
25,612
570,720
8.7764
26,217.33 | 722,968
Books
7,377
120,925
207,759
28.1631
62,281.48 | 5,972,639
MovieTV
5,649
28,987
79,737
14.1166
30,207.68 | 292,389
Table 2. Basic statistics of benchmark datasets.
‘ Tasks
Datasets
Metrics
DeepRec
LLMRec
RecAgent
LongAgent
ReMem (Ours)
Imp.*
SASRec
BERT4Rec
P5
TokenRec
ToolRec
iAgent
Qwen-L1
Mem0
Searching
Games
HR@1
0.2828
0.2955
0.2509
0.3011
0.1970
0.3824
0.3249
0.3995
0.4292 ±0.0243
7.43%
HR@3
0.3115
0.3225
0.2525
0.3711
0.3636
0.5096
0.4851
0.5184
0.5590 ±0.0367
7.83%
MovieTV
HR@1
0.1763
0.1838
0.1319
0.1799
0.1613
0.2308
0.2260
0.2480
0.2552 ±0.0164
2.92%
HR@3
0.2208
0.2320
0.1816
0.2501
0.1827
0.3536
0.3162
0.3342
0.3656 ±0.0415
3.37%
Books
HR@1
0.1761
0.1898
0.1411
0.2232
0.2172
0.2925
0.2843
0.3133
0.3588 ±0.0190
14.55%
Table 3. Performance comparison between representative baselines and ReMem across three commonly used datasets on three recommendation tasks. The best and second-best results are highlighted in bold and underlined fonts, respectively. For ReMem, we conduct independent inference five times and report the mean and standard deviation. The improvements over baselines are statistically significant ( p<0.01 ).
Figure 4. Analysis of reasoning capability across contexts of varying lengths on the Books dataset for three recommendation tasks.
Module
Games
MovieTV
Books
HR@1
HR@3
HR@1
HR@3
HR@1
HR@3
ReMem (Full)
0.4292
0.5590
0.2552
0.3656
0.3588
0.5069
- OCR
0.3959
0.5067
0.2466
0.3403
0.3145
0.4520
- GRPO
0.3767
0.5025
0.2213
0.3410
0.2750
0.4680
- TEM
0.3301
0.4819
0.2236
0.3159
0.2537
0.4021
Table 4. Results of Ablation Studies on the Searching task.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5. Examples of the item pages in the Games dataset and WebWalkerQA dataset. Here we also provide the perception results of these examples by DeepSeek-OCR-2 ( Wei et al., 2026 ) .
Context Length = 37,781 Tokens
Item Sequence ↓
Memory
0. Forza Horizon 3 \u2013 Xbox One 1. RIG 400HX Stereo Gaming Headset for Xbox One 2. Silicone Cover Set for Oculus Quest 2
m1= The user interacts with gaming-related items, including Xbox One games (Forza Horizon 3), Xbox accessories (RIG 400HX Stereo Gaming Headset), and VR accessories (Silicone Cover Set for Oculus Quest 2). They show interest in gaming hardware and peripherals, particularly for Xbox and VR platforms. Reviews indicate they prioritize functionality and comfort (e.g., headset fit, VR cover comfort), and are open to budget-friendly options. They may also have family-oriented gaming interests.
3. USB C Portable Charger 5000mAh 4. Oculus Quest 2 Battery Pack
m2= {Similar Content}. Additionally, they value portable and convenient accessories that enhance gaming experiences, such as battery packs for VR headsets (Oculus Quest 2) to extend playtime and avoid cable clutter, and compact power banks for mobile devices during gaming sessions.
Candidates
VINDIJA Head Strap for Oculus Quest 2 with Battery
Game racing wheel 270 degree
Appendix
Table 5. A case study demonstrating the effectiveness of the proposed memory mechanism in modeling evolving user preferences. To improve readability, only the item titles are shown in the interaction sequence.
Dataset
Games
MovieTV
Short
Medium
Long
Short
Medium
Long
#Samples
25,550
68,350
862
482
5,120
47
#Tokens (K)
0-14
14-112
112-722
0-14
14-112
112-292
Task: Searching, Metric: HR@1
Ours
0.4388
0.4111
0.3563
0.2609
0.2549
0.2201
QwenL1
0.3418
0.3113
0.3005
0.2363
0.2227
0.2072
Appendix
Table 6. Supplements on various length reasoning evaluation across three recommendation tasks on the Games and MovieTV datasets.
Figure 6. The effect of chunk window size W under various datasets and tasks.
Second per Sample
Books
MovieTV
Games
Avg. #Tokens
62,281.48
30,207.68
26,217.33
Qwen3.5-9B
9.8142
3.9504
3.6327
ReMem
224.6032
118.5262
98.2545
Appendix
Table 7. Inference time of our model on a single H20 (96 GB) GPU.
Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reasoning and self-awareness of uncertainty. We introduce ReasonRec, a reasoning-augmented multimodal agent structured around a three-stage explicit reasoning pipeline. Specifically, we propose a reasoning-aware visual instruction tuning strategy that systematically transforms diverse recommendation tasks into unified CoT prompts, enabling the VLM to explicitly articulate intermediate decision steps. Additionally, our evidence-horizon curriculum progressively enhances the reasoning complexity to better handle cold-start and long-tail user scenarios, significantly boosting model generalization. Furthermore, the uncertainty-guided delegation mechanism empowers the agent to assess its own confidence, strategically allocating computational resources to optimize both recommendation accuracy and inference efficiency. Comprehensive experiments on four standard recommendation tasks across five real-world datasets demonstrate that ReasonRec achieves over 30% relative improvement in key ranking metrics compared to state-of-the-art multimodal recommenders. Crucially, ReasonRec substantially reduces inference latency by dynamically delegating up to 35% of queries to efficient sub-models without compromising accuracy. Extensive ablation studies further confirm that each proposed reasoning and planning mechanism individually contributes substantially to ReasonRec's overall effectiveness. Collectively, our results illustrate a clear pathway towards interpretable, adaptive, and efficient multimodal recommendation through explicit reasoning and agentic design.
Yihua Zhang, Mingfu Liang, Jiyan Yang +11
Meta AI · Michigan State University · The University of North Carolina at Chapel Hill +1
Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vulnerable to feedback loops: repeated exposure is mistaken for preference, immediate clicks dominate delayed satisfaction, and fluent explanations need not reflect the ranking decision. We propose our method, a model-agnostic framework for long-horizon recommendation. Our method uses a frozen multimodal language model to convert item content and feedback into evidence-grounded semantic atoms, then maintains separate short-term, long-term, and exposure memories. Propensity-weighted updates reduce policy-induced exposure bias, while a conservative offline critic reranks candidates for delayed satisfaction under a behavior-support constraint. Explanations use only influential evidence atoms and are checked by counterfactual deletion. We provide an identification result and evaluate the framework in e-commerce-like, news-like, and short-video-like environments. Across ten seeds, our method improves discounted long-term value over the strongest alternative by 6.1%, 7.6%, and 6.7%, respectively. Twenty-seed paired ablations show significant value drops after removing propensity correction (0.739 +/- 0.191) or conservative support regularization (0.523 +/- 0.234). A frozen instruction language model also more than doubles semantic-atom NDCG over TF-IDF on a held-out paraphrase benchmark.
Memory-augmented LLM agents have advanced personalized recommendation, yet existing approaches universally adopt flat memory representations that conflate ephemeral signals with stable preferences, and none provides a complete lifecycle governing how memory should evolve. We propose MARS (Memory-Augmented Agentic Recommender System), a framework that treats recommendation as a partially observable problem and maintains a structured belief state that progressively abstracts noisy behavioral observations into a compact estimate of user preferences. MARS organizes this belief state into three tiers: event memory buffers raw signals, preference memory maintains fine-grained mutable chunks with explicit strength and evidence tracking, and profile memory distills all preferences into a coherent natural language narrative. A complete lifecycle of six operations -- extraction, reinforcement, weakening, consolidation, forgetting, and resynthesis -- is adaptively scheduled by an LLM-based planner rather than fixed-interval heuristics. Experiments on four InstructRec benchmark domains show that MARS achieves state-of-the-art performance with average improvements of 26.4% in HR@1 and 10.3% in NDCG@10 over the strongest baselines with further gains from agentic scheduling in evolving settings.