RealWorldShop: Benchmarking and Improving Conversational Shopping Agents in Real-World E-commerce
Organizations: College of Computer Science, Sichuan University · Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, China · Institute of Data Science, National University of Singapore
Abstract
Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing benchmarks are mostly outcome-oriented or execution-oriented, leaving this evolving decision process under-evaluated. We introduce REALWORLDSHOP, a benchmark built on 3.28M grounded products, structured shopping episodes, a profile-grounded and actioncontrolled user simulator, and role-play evaluation. Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially under ambiguous intent, bundle, and multi-intent scenarios. We further propose REALSHOP_AGENT, an executable session-control framework with explicit state management, shopping-flow control, catalog-grounded retrieval, and runtime guards. Experiments show that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.
Figures & tables
| Component | Illustrative Episode |
|---|---|
| A budget-conscious, low-patience user preferring compact products, with latent constraints on counter space, noise, and household size. | |
| Ambiguous bundle shopping for a small-apartment kitchen, starting from a vague request for “something useful for my kitchen.” | |
| The simulator reveals space constraints, rejects an oversized appliance, updates the budget, and requests a compatible bundle. | |
| Target | Clarification, state tracking, constraint updating, bundle coordination, and grounded convergence. |
| Turn-Level Quality | Session-Level Process Quality | Turn–Session | Holistic Outcome | User Signal | ||||||||||
| Model | Need | Rec. Acc. | Rationale | Turn Avg. | Clarif. | State | Constraint | Pace | Personal. | Sess. Avg. | Gap | Succ. | Conv. | Early Exit |
| GPT-5 singh2025openai | 1.62 | 1.41 | 1.33 | 1.45 | 1.68 | 1.23 | 1.09 | 0.94 | 1.08 | 1.20 | 0.25 | 0.71 | 2.49 | 0.10 |
| Gemini-2.5-Flash comanici2025gemini | 0.91 | 0.80 | 0.84 | 0.85 | 1.21 | 0.69 | 0.56 | 0.59 | 0.45 | 0.70 | 0.15 | 0.38 | 1.64 | 0.35 |
| GPT-4o hurst2024gpt | 1.30 | 0.71 | 0.75 | 0.92 | 0.95 | 0.62 | 0.49 | 0.37 | 0.35 | 0.56 | 0.36 | 0.38 | 1.48 | 0.39 |
| DeepSeek-V3.2 liu2025deepseek | 1.36 | 0.92 | 0.81 | 1.03 | 1.41 | 0.80 | 0.63 | 0.64 | 0.52 | 0.80 | 0.23 | 0.56 | 1.93 | 0.27 |
| Kimi-K2.6 team2025kimi | 1.45 | 1.22 | 1.21 | 1.29 | 1.54 | 0.90 | 0.88 | 0.70 | 0.95 | 0.99 | 0.30 | 0.67 | 2.23 | 0.16 |
| Turn-Level Quality | Session-Level Process Quality | Turn–Session | Holistic Outcome | User Signal | ||||||||||
| Variant | Need | Rec. Acc. | Rationale | Turn Avg. | Clarif. | State | Constraint | Pace | Personal. | Sess. Avg. | Gap | Succ. | Conv. | Early Exit |
| Prompt-only Backbones | ||||||||||||||
| Qwen3.5-27B Base yang2025qwen3 | 1.41 | 0.83 | 0.80 | 1.01 | 1.49 | 0.58 | 0.57 | 0.50 | 0.36 | 0.70 | 0.31 | 0.55 | 1.74 | 0.25 |
| Qwen3.5-27B + SFT/RL | 1.67 | 1.09 | 0.95 | 1.24 | 1.54 | 0.92 | 0.93 | 0.82 | 0.78 | 1.00 | 0.24 | 0.65 | 2.19 | 0.20 |
| Module Ablations on SFT/RL Backbone | ||||||||||||||
| w/o Session State Manager | 1.53 | 1.27 | 1.06 | 1.29 | 1.58 | 1.02 | 0.87 | 0.83 | 0.93 | 1.05 | 0.24 | 0.61 | 2.18 | 0.22 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Finding Source | Diagnostic Subset | Metric | Base | +SFT/RL | Full Agent | vs. Base |
| Figure 2: Early-exit sessions reveal weak session-level process control | ||||||
| Fig. 2 | Overall sessions | Early Exit Rate | 0.25 | 0.20 | 0.09 | |
| Fig. 2 | Early-exit-risk episodes | Sess. Avg. | 0.60 | 0.79 | 1.21 | |
| Fig. 2 | Early-exit-risk episodes | State Tracking | 0.51 | 0.75 | 1.33 | |
| Fig. 2 | Early-exit-risk episodes | Constraint Updating | 0.47 | 0.78 | 1.24 | |
| Figure 3: Interaction phenomena stress state revision and trajectory management | ||||||
| Factor | Values |
|---|---|
| Coping style | problem-focused, emotion-focused, avoidant, assertive, compliant |
| Interaction style | directive, collaborative, browsing, interrogative, efficiency-first, tradeoff-seeking |
| Verbosity | low, medium, high |
| Politeness | direct, neutral, warm |
| Skepticism | low, medium, high |
| Decisiveness | low, medium, high |
| Judgment Target | QWK / | MAE | Agreement |
|---|---|---|---|
| Need Understanding | 0.82 | 0.07 | 0.95 |
| Recommendation Accuracy | 0.81 | 0.09 | 0.91 |
| Rationale Quality | 0.81 | 0.13 | 0.90 |
| Clarification | 0.82 | 0.11 | 0.93 |
| State Tracking | 0.86 | 0.07 | 0.94 |
| Constraint Updating | 0.89 | 0.06 | 0.96 |
| Metric Group | ICC(2,1) | Mean Std. | Flip Rate |
|---|---|---|---|
| Turn-level Quality | 0.86 | 0.09 | 0.03 |
| Session-level Process Quality | 0.87 | 0.08 | 0.02 |
| Overall Success | 0.92 | 0.04 | 0.02 |
| Convergence Quality | 0.88 | 0.11 | 0.04 |
| Validation Aspect | Metric | Score / Rate | Description |
|---|---|---|---|
| Long-term preference consistency | Preference consistency rate | 0.96 | Generated turns preserve stable user preferences in the profile. |
| Hard-constraint consistency | Violation-free rate | 0.93 | The simulator avoids violating hard constraints such as budget, compatibility, and logistics. |
| Soft / negotiable constraint consistency | Soft-constraint alignment rate | 0.91 | The simulator reflects soft preferences without treating them as strict requirements. |
| Disclosure-plan compliance | Planned-disclosure compliance rate | 0.94 | Session-specific intents and constraints are revealed according to the predefined disclosure plan. |
| Hidden-intent control | Non-leakage rate | 0.88 | Ambiguous-intent profiles avoid prematurely leaking latent goals or hidden constraints. |
| Interaction-style fidelity | Style-consistency score | 0.93 | Behavioral factors such as skepticism, patience, decisiveness, and comparison orientation are reflected in user behavior. |
| Validation Aspect | Metric | Score / Rate | Description |
|---|---|---|---|
| Overall action compliance | Intended-action match rate | 0.94 | Generated user utterances semantically match the assigned target actions. |
| Transition validity | State-valid transition rate | 0.97 | The intended user action is appropriate given the current dialogue state and session progress. |
| Information provision | Provide-action compliance rate | 0.95 | When assigned to provide missing information, the simulator reveals the expected intent, preference, or constraint. |
| Constraint refinement and update | Refine / update compliance rate | 0.97 | The simulator correctly refines, updates, or overrides prior requirements according to the target action. |
| Recommendation feedback | Reject / compare compliance rate | 0.98 | The simulator follows target actions involving rejection, comparison requests, or feedback on recommended candidates. |
| Decision and termination control | Confirm / end compliance rate | 0.92 | The simulator confirms a decision or ends the session only when the target action specifies confirmation or termination. |
| Item | Value |
|---|---|
| SFT examples | 10K |
| RL prompt seeds | 2K |
| RL training prompts | 1.5K |
| Held-out development prompts | 0.5K |
| GRPO group size | 8 rollouts per prompt |
| Max assistant turns per rollout | 30 |
| Hyperparameter | Value |
|---|---|
| Learning rate | 1e-6 |
| Weight decay | 0.1 |
| GRPO clipping | 0.2 |
| High clipping threshold | 0.28 |
| Global batch size | 64 |
| Samples per prompt | 8 |
| Case | Scenario | Stress Factor | Diagnostic Target |
| Profile A | Compatibility-sensitive multi-intent shopping | Old-wall-box fit, narrow balcony space, urgent vs. deferrable items | State tracking, compatibility verification, split-order planning |
| Profile B | Ambiguous core-item lighting upgrade | Vague study-room lighting need; 75 mm opening and budget priority revealed later | Clarification, hidden-constraint elicitation, hard-constraint propagation, candidate verification |
| Profile C | Low-patience family purchase | Speaker compatibility, fast convergence, later add-on item | Verification before recommendation, low-patience decision pacing |
| Failure 1 | Missing clarification | Underspecified small-kitchen request | Avoid premature recommendation under hidden constraints |
| Failure 2 | Priority loss in multi-intent shopping | Urgent setup items vs. optional decorative items | Preserve subgoal priority and avoid stale planning |
| Failure 3 | Constraint update not propagated | 85 mm recommendation after user corrects to 75 mm | Invalidate stale candidates after hard-constraint update |