Organizations: College of Computer Science, Sichuan University · Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, China · Institute of Data Science, National University of Singapore
Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing benchmarks are mostly outcome-oriented or execution-oriented, leaving this evolving decision process under-evaluated. We introduce REALWORLDSHOP, a benchmark built on 3.28M grounded products, structured shopping episodes, a profile-grounded and actioncontrolled user simulator, and role-play evaluation. Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially under ambiguous intent, bundle, and multi-intent scenarios. We further propose REALSHOP_AGENT, an executable session-control framework with explicit state management, shopping-flow control, catalog-grounded retrieval, and runtime guards. Experiments show that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.
Figures & tables
Figure 1: Overview of RealWorldShop . The benchmark integrates a 3.28M-SKU product inventory, approximately 1,200 structured shopping scenarios, and over 2,000 synthesized user profiles to construct profile-grounded user simulators, together with a session-level role-play protocol for evaluating conversational shopping agents.
Component
Illustrative Episode
P
A budget-conscious, low-patience user preferring compact products, with latent constraints on counter space, noise, and household size.
S
Ambiguous bundle shopping for a small-apartment kitchen, starting from a vague request for “something useful for my kitchen.”
Π
The simulator reveals space constraints, rejects an oversized appliance, updates the budget, and requests a compatible bundle.
Target
Clarification, state tracking, constraint updating, bundle coordination, and grounded convergence.
Table 1: Shopping episode illustration.
Turn-Level Quality ↑
Session-Level Process Quality ↑
Turn–Session
Holistic Outcome
User Signal
Model
Need
Rec. Acc.
Rationale
Turn Avg.
Clarif.
State
Constraint
Pace
Personal.
Sess. Avg.
Gap ↓
Succ. ↑
Conv. ↑
Early Exit ↓
GPT-5 singh2025openai
1.62
1.41
1.33
1.45
1.68
1.23
1.09
0.94
1.08
1.20
0.25
0.71
2.49
0.10
Gemini-2.5-Flash comanici2025gemini
0.91
0.80
0.84
0.85
1.21
0.69
0.56
0.59
0.45
0.70
0.15
0.38
1.64
0.35
GPT-4o hurst2024gpt
1.30
0.71
0.75
0.92
0.95
0.62
0.49
0.37
0.35
0.56
0.36
0.38
1.48
0.39
DeepSeek-V3.2 liu2025deepseek
1.36
0.92
0.81
1.03
1.41
0.80
0.63
0.64
0.52
0.80
0.23
0.56
1.93
0.27
Kimi-K2.6 team2025kimi
1.45
1.22
1.21
1.29
1.54
0.90
0.88
0.70
0.95
0.99
0.30
0.67
2.23
0.16
Table 2: Overall benchmark results on RealWorldShop . Turn Avg. is the arithmetic mean of Need, Rec. Acc., and Rationale, while Sess. Avg. is the arithmetic mean of Clarif., State, Constraint, Pace, and Personal. Gap measures the degradation from turn-level quality to session-level process quality and is computed as the unrounded Turn Avg. minus the unrounded Sess. Avg. Succ. denotes success rate, Conv. denotes convergence quality, and Early Exit denotes the simulator-side early-exit rate. All aggregate metrics are computed before rounding. Best and second-best results across all evaluated systems are marked in bold and underline , respectively. Tied results receive the same formatting.
Figure 2: Early-exit sessions reveal weak session-level process control. Compared with non-exit sessions, early-exit sessions show larger degradation in state tracking, constraint updating, and decision pacing, indicating that user abandonment is closely tied to failures in maintaining and revising session state.
Figure 3: Interaction phenomena stress state revision and trajectory management. Performance declines under constraint updates, preference overrides, recommendation rejections, and bundle re-planning, showing that agents struggle to propagate revised user information through the full shopping trajectory.
Figure 4: Scenario complexity amplifies the turn-to-session quality gap. Agents perform best in explicit single-item sessions, while ambiguous, bundle, multi-intent, and long-horizon scenarios expose larger gaps between local response quality and session-level shopping success.
Figure 5: Demanding interaction styles increase session fragility. Efficiency-first, high-skepticism, and low-patience users lead to larger turn-to-session gaps and higher early-exit rates, highlighting the need to evaluate shopping agents under realistic convergence pressure.
Turn-Level Quality ↑
Session-Level Process Quality ↑
Turn–Session
Holistic Outcome
User Signal
Variant
Need
Rec. Acc.
Rationale
Turn Avg.
Clarif.
State
Constraint
Pace
Personal.
Sess. Avg.
Gap ↓
Succ. ↑
Conv. ↑
Early Exit ↓
Prompt-only Backbones
Qwen3.5-27B Base yang2025qwen3
1.41
0.83
0.80
1.01
1.49
0.58
0.57
0.50
0.36
0.70
0.31
0.55
1.74
0.25
Qwen3.5-27B + SFT/RL
1.67
1.09
0.95
1.24
1.54
0.92
0.93
0.82
0.78
1.00
0.24
0.65
2.19
0.20
Module Ablations on SFT/RL Backbone
w/o Session State Manager
1.53
1.27
1.06
1.29
1.58
1.02
0.87
0.83
0.93
1.05
0.24
0.61
2.18
0.22
Table 3: Ablation study of RealShop_Agent under the same three-level metric structure as Table 2 . Turn-level and session-level dimensions are scored on a {0,1,2} scale. Turn Avg. is the macro-average of Need, Recommendation Accuracy, and Rationale, while Sess. Avg. is the macro-average of Clarification, State Tracking, Constraint Updating, Decision Pacing, and Personalization. Gap is computed as Turn Avg. minus Sess. Avg. All aggregate metrics are computed before final rounding. Succ. and Conv. are holistic-judge outcome metrics, while Early Exit is a simulator-side user signal. Higher values are better for all metrics except Gap and Early Exit . Best results are shown in bold, and second-best results are underlined. Tied second-best results are underlined simultaneously.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Finding Source
Diagnostic Subset
Metric
Base
+SFT/RL
Full Agent
Δ vs. Base
Figure 2: Early-exit sessions reveal weak session-level process control
Fig. 2
Overall sessions
Early Exit Rate ↓
0.25
0.20
0.09
0.16↓
Fig. 2
Early-exit-risk episodes
Sess. Avg. ↑
0.60
0.79
1.21
0.61↑
Fig. 2
Early-exit-risk episodes
State Tracking ↑
0.51
0.75
1.33
0.82↑
Fig. 2
Early-exit-risk episodes
Constraint Updating ↑
0.47
0.78
1.24
0.77↑
Figure 3: Interaction phenomena stress state revision and trajectory management
Appendix
Table 4: Finding-oriented diagnostic results of REALSHOP_AGENT. Rows are aligned with the failure modes analyzed in Figures 2 – 5 . Base denotes the prompt-only Qwen3.5-27B backbone, +SFT/RL denotes the adapted backbone without the full execution harness, and Full Agent denotes REALSHOP_AGENT with the complete execution harness. For Figure 4 , Gap is computed as Turn Avg. minus Sess. Avg. Δ reports the absolute improvement of Full Agent over Base; for paired metrics in Figure 4 , the two gains are reported in the same order as the metric column.
Table 6: Interaction-profile factors used to control simulated user behavior. These factors affect how users communicate, compare options, tolerate clarification, react to recommendations, and decide whether to continue or exit a shopping session.
Judgment Target
QWK / κ
MAE
Agreement
Need Understanding
0.82
0.07
0.95
Recommendation Accuracy
0.81
0.09
0.91
Rationale Quality
0.81
0.13
0.90
Clarification
0.82
0.11
0.93
State Tracking
0.86
0.07
0.94
Constraint Updating
0.89
0.06
0.96
Appendix
Table 7: Human validation of the automatic judging protocol. For ordinal score dimensions, the first column reports quadratic weighted kappa (QWK). For outcome labels, it reports agreement-based κ when applicable. MAE denotes mean absolute error, and Agreement denotes exact agreement with aggregated human annotations.
Metric Group
ICC(2,1)
Mean Std.
Flip Rate
Turn-level Quality
0.86
0.09
0.03
Session-level Process Quality
0.87
0.08
0.02
Overall Success
0.92
0.04
0.02
Convergence Quality
0.88
0.11
0.04
Appendix
Table 8: Stability of the automatic judging protocol across repeated judging passes. We report ICC(2,1), mean standard deviation, and label flip rate to assess whether the main trends are robust to judge sampling variance. Flip Rate denotes the proportion of sessions whose final discrete label changes across repeated judging passes; for score-based dimensions, scores are first mapped to their nearest discrete rubric value.
Validation Aspect
Metric
Score / Rate
Description
Long-term preference consistency
Preference consistency rate
0.96
Generated turns preserve stable user preferences in the profile.
Hard-constraint consistency
Violation-free rate
0.93
The simulator avoids violating hard constraints such as budget, compatibility, and logistics.
Soft / negotiable constraint consistency
Soft-constraint alignment rate
0.91
The simulator reflects soft preferences without treating them as strict requirements.
Disclosure-plan compliance
Planned-disclosure compliance rate
0.94
Session-specific intents and constraints are revealed according to the predefined disclosure plan.
Hidden-intent control
Non-leakage rate
0.88
Ambiguous-intent profiles avoid prematurely leaking latent goals or hidden constraints.
Interaction-style fidelity
Style-consistency score
0.93
Behavioral factors such as skepticism, patience, decisiveness, and comparison orientation are reflected in user behavior.
Appendix
Table 9: Profile-level validation of the Kimi-K2.6 user simulator. All metrics are reported such that higher values indicate better simulator reliability. We evaluate whether simulated user turns remain faithful to structured profiles, follow the disclosure plan, avoid premature leakage of hidden intent, and reflect interaction-profile factors.
Validation Aspect
Metric
Score / Rate
Description
Overall action compliance
Intended-action match rate
0.94
Generated user utterances semantically match the assigned target actions.
Transition validity
State-valid transition rate
0.97
The intended user action is appropriate given the current dialogue state and session progress.
Information provision
Provide-action compliance rate
0.95
When assigned to provide missing information, the simulator reveals the expected intent, preference, or constraint.
Constraint refinement and update
Refine / update compliance rate
0.97
The simulator correctly refines, updates, or overrides prior requirements according to the target action.
Recommendation feedback
Reject / compare compliance rate
0.98
The simulator follows target actions involving rejection, comparison requests, or feedback on recommended candidates.
Decision and termination control
Confirm / end compliance rate
0.92
The simulator confirms a decision or ends the session only when the target action specifies confirmation or termination.
Appendix
Table 10: Action-level validation of the Kimi-K2.6 user simulator. All metrics are reported such that higher values indicate better simulator controllability. Percentage-based metrics are computed over simulated user turns from the 200 sampled validation profiles. We evaluate whether generated user utterances follow the intended target actions, satisfy dialogue-state transition constraints, and reflect scenario- and style-conditioned action patterns.
Item
Value
SFT examples
10K
RL prompt seeds
2K
RL training prompts
1.5K
Held-out development prompts
0.5K
GRPO group size
8 rollouts per prompt
Max assistant turns per rollout
30
Appendix
Table 11: Training scale for backbone adaptation.
Hyperparameter
Value
Learning rate
1e-6
Weight decay
0.1
GRPO clipping
0.2
High clipping threshold
0.28
Global batch size
64
Samples per prompt
8
Appendix
Table 12: Main hyperparameters used for GRPO optimization. Additional implementation details follow the SLIME slime_github training framework.
Figure 6: Profile synthesis prompt used to construct structured user profiles for RealWorldShop .
Figure 7: User simulator prompt used to generate controlled multi-turn user behavior in RealWorldShop .
Figure 8: Shopping-agent prompt used to control session-level decision making in RealWorldShop .
Figure 9: Evaluation judge prompt used to assess completed shopping-agent trajectories in RealWorldShop .
Case
Scenario
Stress Factor
Diagnostic Target
Profile A
Compatibility-sensitive multi-intent shopping
Old-wall-box fit, narrow balcony space, urgent vs. deferrable items
State tracking, compatibility verification, split-order planning
Profile B
Ambiguous core-item lighting upgrade
Vague study-room lighting need; 75 mm opening and budget priority revealed later
As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade only the final choice can neither pose this long-horizon challenge nor explain which requirement an agent missed. To address this gap, we introduce EComAgentBench, a benchmark of 662 tasks grounded in real Amazon products and reviews. Each task scatters these requirements across a visible query, a tool-gated profile, and scripted clarification; an agent must uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. Moreover, typed, source-tagged rubrics grade every task, attributing each failure to a requirement and its source. Construction is automated yet reliable, with every answer fixed in code before any text is generated and every sample validated. Our evaluation of seven models reveals that even the strongest attains only 57.1% overall accuracy, and rubric satisfaction degrades from visible to hidden sources. Overall, we believe EComAgentBench will serve as a reproducible foundation for moving shopping agents from single-query search toward dependable assistance over long horizons.
Conversational shopping assistants now serve hundreds of millions of customers, yet no existing benchmark jointly evaluates the open-ended multi-turn reasoning, domain expertise, and criterion-level quality that real shopping conversations demand. Shopping reasoning is unique among language model applications. Unlike factual question answering or verifiable code generation, it requires balancing subjective preferences, budget constraints, and cross-product trade-offs across multi-turn dialogue, capabilities absent from previous e-commerce and general-purpose benchmarks. We introduce the Shopping Reasoning Bench, an expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10863 importance-weighted binary rubrics authored by retail domain experts. These criteria are organized under a taxonomy of five reasoning categories and fifteen subcategories covering diverse demands such as preference refinement, trade-off analysis, and compatibility assessment. An evaluation of nine models across three families (GPT, Claude, Gemini) shows that pass rates reach only 57--77% overall. On multi-turn missions, all models score 13--29 points lower on optional above-and-beyond criteria than on required ones, and performance degrades 4--18 points as conversations progress. These gaps show that current models handle basic shopping assistance but fall short of expert-level advice, making Shopping Reasoning Bench a challenging testbed for future shopping assistant development.
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.