Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions
Organizations: ByteDance · The University of Melbourne · University of Technology Sydney · AXON
Abstract
Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study (), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing
Figures & tables
| Statistic | Value |
|---|---|
| Raw Amazon categories | 34 |
| Top active users retained per category | 100 |
| Candidate user pool | 3,300 |
| Target users / dialogue instances | 327 |
| Dialogues per variant | 327 |
| Dialogue runs ( variants) | 1,635 |
| Variant | Prod. emb. | Pipeline | |||
|---|---|---|---|---|---|
| Baseline | ✗ | ✗ | ✗ | — | Multi-agent |
| TextRAG | ✓ | ✓ | ✓ | Multi-agent | |
| NoRev | ✓ | ✓ | ✗ | Multi-agent | |
| Single | ✓ | ✓ | ✓ | Single-agent | |
| Full | ✓ | ✓ | ✓ | Multi-agent |
| Metric | Baseline | TextRAG | NoRev 1 | Single 2 | Full 3 |
|---|---|---|---|---|---|
| GPC | 3.77 | 4.93 | 4.88 | 4.91 | 4.93 |
| CIS | 3.17 | 4.88 | 4.79 | 4.78 | 4.89 |
| IT | 3.02 | 4.47 | 4.39 | 4.19 | 4.45 |
| TC | 5.00 | 5.00 | 5.00 | 5.00 | 5.00 |
| Avg. | 3.74 | 4.82 | 4.76 | 4.72 | 4.82 |
| Metric | Baseline | TextRAG | NoRev 1 | Single 2 | Full 3 |
|---|---|---|---|---|---|
| Acc. | 10.31% | 42.73% | 41.90% | 48.20% | 43.73% |
| Dimension | Baseline | TextRAG | NoRev 1 | Single 2 | Full 3 |
|---|---|---|---|---|---|
| Overall | 2.20 | 3.40 | 3.80 | 3.80 | 4.60 |
| GPC | 1.20 | 3.60 | 4.00 | 4.00 | 4.60 |
| CIS | 1.60 | 3.60 | 4.20 | 4.00 | 4.80 |
| IT | 2.00 | 3.60 | 3.80 | 3.60 | 4.60 |
| TC | 4.20 | 4.60 | 4.60 | 4.80 | 5.00 |
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Items | Reviews | Avg. | Dlg. |
|---|---|---|---|---|
| Electronics | 27,292 | 9,907,789 | 363.03 | 10 |
| Home_and_Kitchen | 51,421 | 8,780,109 | 170.75 | 10 |
| Kindle_Store | 113,360 | 7,172,600 | 63.27 | 10 |
| Movies_and_TV | 67,650 | 6,112,901 | 90.36 | 8 |
| Health_and_Household | 34,382 | 5,492,657 | 159.75 | 10 |
| Clothing_Shoes_and_Jewelry | 63,329 | 5,360,811 | 84.65 | 10 |
| Method | Role calls / turn | Auto. Avg. | User Overall |
|---|---|---|---|
| Baseline | 1 | 3.74 | 2.20 |
| TextRAG | 4 | 4.82 | 3.40 |
| NoRev | 4 | 4.76 | 3.80 |
| Single | 1 | 4.72 | 3.80 |
| Full | 4 | 4.82 | 4.60 |
| Comparison / metric | Adjusted | |
|---|---|---|
| RAG vs. Baseline ∗ | ||
| Full–Single: IT | 0.433 | |
| Full–Single: CIS | 0.040 | 0.508 |
| Full–Single: Overall | 0.002 | 0.351 |
| Variant | SR@all | First-hit turn |
|---|---|---|
| TextRAG | 0.352 | 2.26 |
| NoRev | 0.309 | 2.29 |
| Single | 0.352 | 2.50 |
| Full | 0.361 | 2.49 |
| Metric | GPC | CIS | IT | Overall |
|---|---|---|---|---|
| Spearman | 0.508 | 0.533 | 0.471 | 0.538 |
| Memory | Overall | IT | CIS |
|---|---|---|---|
| Full | 4.796 | 4.387 | 4.882 |
| None | 4.610 | 3.796 | — |
| Summary only | 4.801 | — | 4.860 |
| Top-1 turn only | 4.634 | 3.935 | — |
| Work | Training / inference setup | Evidence and task resources |
|---|---|---|
| ChatCRS Li et al. (2025a) | LLM generation with knowledge retrieval and a LoRA-fine-tuned goal planner. | External knowledge base and annotated dialogue goals; multi-goal CRS datasets. |
| RevCore Lu et al. (2021) | Trained recommendation and response-generation components. | Item-linked reviews, entities, and conversational recommendation data. |
| MACRec Wang et al. (2024a) | Configurable LLM agents with task-dependent tools. | User/item information and task-dependent tools; multiple recommendation scenarios. |
| RA-Rec Kemper et al. (2024) | Prompt-driven state tracking and generation with a pretrained dense retriever. | Domain-specific state fields, item reviews and metadata; demonstrated on Yelp restaurants. |
| PEARL Kim et al. (2024) | Review-grounded dialogue dataset with downstream model experiments. | Review-derived personas and knowledge for synthetic movie recommendation dialogues. |
| MUSE Wang et al. (2025) | Multimodal conversational dataset with benchmark experiments. | Scenario-grounded profiles and clothing-domain multimodal conversations. |
| Criterion | Description and Scoring Rubric (1–5) |
|---|---|
| 1. Global Personalization Consistency (GPC) | Did the agent maintain the user’s persona constraints throughout the entire dialogue? |
| 1 (Inconsistent/Amnesic): Agent started well but forgot constraints (e.g., budget, size) in later turns, or hallucinated user details. | |
| 3 (Fragmented): Agent remembered explicit constraints (e.g., Size M) but forgot implicit ones (e.g., Style preferences) as the chat went on. | |
| 5 (Rock Solid): Every single response, from start to finish, aligned with the User Profile. The agent never “reset” or forgot context. | |
| 2. Cumulative Information Synthesis (CIS) | Looking at all agent responses combined, did it build a complete, factual picture? |
| 1 (Fragmented/Repetitive): Agent repeated the same facts multiple times or gave contradictory information across different turns. |
| Interactivity Macro-level Features | Definition |
|---|---|
| Topic Management | the strategies and techniques used to control and navigate the flow of topics. |
| Tone Choice Appropriateness | the suitability of the tone used in communication, ensuring it aligns with the context, audience, and purpose to convey the intended message. |
| Conversation Opening | the initial interaction or exchange that begins a dialogue, often setting the tone and context for the dialogue. |
| Conversation Closing | the process of ending a dialogue or interaction, which involves signaling the conclusion of the discussion, summarizing key points, and often expressing a farewell. |
| Recommender Agent Response |
|---|
| User: Can you help me find a good moisturizer for dry skin? |
| Recommender Agent: … To provide the best recommendation, a few details are needed: |
| (1) Skin Concerns: Are you dealing with flakiness, redness, irritation, or general dryness? |
| (2) Skin Type: Do you have sensitive skin, or are you okay with fragrance or other ingredients? |
| (3) Budget: Are you looking for something affordable, mid-range, or premium? |
| (4) Texture Preference: Do you prefer a lightweight gel, rich cream, or a lotion-like consistency? |
| Reasoner Agent Evaluation on Recommend Agent Output |
|---|
| Reasoner Agent: … 1. Personalized Inquiry First: The agent begins by asking questions about skin concerns, type, budget, texture preference, and allergies. This ensures the final recommendation is customized rather than generic, increasing the likelihood of satisfaction and effectiveness. |
| 2. Categorization by Skin Needs: The agent groups moisturizers by skin profiles (e.g., sensitive, very dry, lightweight, intense hydration), allowing users to quickly identify the most relevant category and reducing decision fatigue. |
| 3. Inclusion of Dermatologist-Backed Brands: Products such as CeraVe, La Roche-Posay, and Vanicream are widely recognized by dermatologists for safety, efficacy, and ingredient transparency, which builds user trust. |
| 4. Clear Ingredient Guidance: Highlighting key ingredients (ceramides, hyaluronic acid, glycerin, etc.) helps users understand why certain products work and supports informed decisions. |
| 5. Balance of Price and Performance: Offers both affordable and premium options (e.g., Neutrogena vs. Kiehl’s), ensuring accessibility without compromising quality. |
| 6. Texture and Use-Case Consideration: Distinguishes lightweight gels (for daytime) from rich creams (for nighttime or severe dryness), reflecting practical understanding of skincare routines. |