Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16% across three recommendation agent tasks, namely searching, ranking, and judging.
Figures & tables
Figure 1. Motivation of ReMem. (a) Conventional RecAgents read the Web as text: they extract noisy HTML into long contexts and rely on LLMs to recover user intent, which is brittle across modalities and inefficient for long-context reasoning. (b) The proposed ReMem follows a more human-like principle: see item pages through OCR-based multimodal perception and memorize only evolving preference abstractions through dynamic memory, enabling robust modeling of time-evolving user preferences.
Acc. (%)
Qwen3.5
Qwen3VL
4B
9B
Avg.
4B
8B
Avg.
Baseline
49.13
54.27
51.70
49.32
50.94
50.13
+ 1)
41.51
45.27
43.39
52.51
55.77
54.14
+ 2)
-
-
-
46.38
48.56
47.47
+ 3)
57.02
60.25
58.64
56.62
58.60
57.61
Table 1. The effect of different perception strategies. Strategy 3) parses multimodal web pages into informative textual descriptions through OCR perception, achieving superior performance compared with the other strategies.
Figure 2. Accuracy scores of different long-context reasoning strategies on HotpotQA. Even models that use long-context continual pretraining or vectorized memory techniques fail to maintain consistent performance. In contrast, the agent equipped with dynamic memory demonstrates relative lossless performance extrapolation.
Figure 3. Overview of the proposed ReMem framework. Inspired by how humans browse item pages while maintaining time-evolving memory, ReMem consists of three key components. a) Instead of relying on raw HTML, ReMem adopts an OCR-based perception module that represents item pages through a unified multimodal view of natural text, images, and interactive elements, better matching users’ browsing behavior. b) User interaction history is processed as a sequential stream of chunks, from which the model maintains a compact, fixed-length memory that is continuously updated to capture evolving user preferences. c) To teach the model what to remember and how to update memory in multi-turn interactions, ReMem employs Multi-Mem GRPO, which propagates the final-answer advantage to all intermediate memories that contribute to the final response.
Dataset
#Users
#Items
#Int.
Avg. Seq.
#Tokens (Avg. | Max.)
Games
94,762
25,612
570,720
8.7764
26,217.33 | 722,968
Books
7,377
120,925
207,759
28.1631
62,281.48 | 5,972,639
MovieTV
5,649
28,987
79,737
14.1166
30,207.68 | 292,389
Table 2. Basic statistics of benchmark datasets.
‘ Tasks
Datasets
Metrics
DeepRec
LLMRec
RecAgent
LongAgent
ReMem (Ours)
Imp.*
SASRec
BERT4Rec
P5
TokenRec
ToolRec
iAgent
Qwen-L1
Mem0
Searching
Games
HR@1
0.2828
0.2955
0.2509
0.3011
0.1970
0.3824
0.3249
0.3995
0.4292 ±0.0243
7.43%
HR@3
0.3115
0.3225
0.2525
0.3711
0.3636
0.5096
0.4851
0.5184
0.5590 ±0.0367
7.83%
MovieTV
HR@1
0.1763
0.1838
0.1319
0.1799
0.1613
0.2308
0.2260
0.2480
0.2552 ±0.0164
2.92%
HR@3
0.2208
0.2320
0.1816
0.2501
0.1827
0.3536
0.3162
0.3342
0.3656 ±0.0415
3.37%
Books
HR@1
0.1761
0.1898
0.1411
0.2232
0.2172
0.2925
0.2843
0.3133
0.3588 ±0.0190
14.55%
Table 3. Performance comparison between representative baselines and ReMem across three commonly used datasets on three recommendation tasks. The best and second-best results are highlighted in bold and underlined fonts, respectively. For ReMem, we conduct independent inference five times and report the mean and standard deviation. The improvements over baselines are statistically significant ( p<0.01 ).
Figure 4. Analysis of reasoning capability across contexts of varying lengths on the Books dataset for three recommendation tasks.
Module
Games
MovieTV
Books
HR@1
HR@3
HR@1
HR@3
HR@1
HR@3
ReMem (Full)
0.4292
0.5590
0.2552
0.3656
0.3588
0.5069
- OCR
0.3959
0.5067
0.2466
0.3403
0.3145
0.4520
- GRPO
0.3767
0.5025
0.2213
0.3410
0.2750
0.4680
- TEM
0.3301
0.4819
0.2236
0.3159
0.2537
0.4021
Table 4. Results of Ablation Studies on the Searching task.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5. Examples of the item pages in the Games dataset and WebWalkerQA dataset. Here we also provide the perception results of these examples by DeepSeek-OCR-2 ( Wei et al., 2026 ) .
Context Length = 37,781 Tokens
Item Sequence ↓
Memory
0. Forza Horizon 3 \u2013 Xbox One 1. RIG 400HX Stereo Gaming Headset for Xbox One 2. Silicone Cover Set for Oculus Quest 2
m1= The user interacts with gaming-related items, including Xbox One games (Forza Horizon 3), Xbox accessories (RIG 400HX Stereo Gaming Headset), and VR accessories (Silicone Cover Set for Oculus Quest 2). They show interest in gaming hardware and peripherals, particularly for Xbox and VR platforms. Reviews indicate they prioritize functionality and comfort (e.g., headset fit, VR cover comfort), and are open to budget-friendly options. They may also have family-oriented gaming interests.
3. USB C Portable Charger 5000mAh 4. Oculus Quest 2 Battery Pack
m2= {Similar Content}. Additionally, they value portable and convenient accessories that enhance gaming experiences, such as battery packs for VR headsets (Oculus Quest 2) to extend playtime and avoid cable clutter, and compact power banks for mobile devices during gaming sessions.
Candidates
VINDIJA Head Strap for Oculus Quest 2 with Battery
Game racing wheel 270 degree
Appendix
Table 5. A case study demonstrating the effectiveness of the proposed memory mechanism in modeling evolving user preferences. To improve readability, only the item titles are shown in the interaction sequence.
Dataset
Games
MovieTV
Short
Medium
Long
Short
Medium
Long
#Samples
25,550
68,350
862
482
5,120
47
#Tokens (K)
0-14
14-112
112-722
0-14
14-112
112-292
Task: Searching, Metric: HR@1
Ours
0.4388
0.4111
0.3563
0.2609
0.2549
0.2201
QwenL1
0.3418
0.3113
0.3005
0.2363
0.2227
0.2072
Appendix
Table 6. Supplements on various length reasoning evaluation across three recommendation tasks on the Games and MovieTV datasets.
Figure 6. The effect of chunk window size W under various datasets and tasks.
Second per Sample
Books
MovieTV
Games
Avg. #Tokens
62,281.48
30,207.68
26,217.33
Qwen3.5-9B
9.8142
3.9504
3.6327
ReMem
224.6032
118.5262
98.2545
Appendix
Table 7. Inference time of our model on a single H20 (96 GB) GPU.