Language Models for Page-Level Layout Decisions in E-commerce Search
Organizations: Walmart Global Tech Hoboken, NJ, USA · University of Illinois at Urbana-Champaign Urbana, IL, USA
Abstract
E-commerce search pages are critical touchpoints for millions of online shoppers. While traditional search engines return a ranked list of results, modern E-commerce search pages increasingly incorporate recommender system modules -- for example, secondary stacks that surface alternative product groupings at specific positions. When introduced appropriately, secondary stacks can improve user engagement; however, suboptimal placement may disrupt browsing flow and degrade the primary results. Unlike traditional search ranking, where evaluation techniques such as interleaving are well established, evaluating page-level layout changes e.g., when and where to insert a secondary stack remains challenging without costly online A/B testing. To address this, we study offline methods for evaluating whether a given layout decision -- specifically, the inclusion of a secondary stack at a particular position -- is beneficial to users. We investigate language models as scalable evaluators by comparing direct prompt-based, prompt-derived feature, and representation-based methods. Our results show that representation-based approaches consistently outperform prompt-based judging in predicting user engagement, suggesting they provide a reliable foundation for offline layout evaluation in E-commerce search.
Figures & tables
| Approach | ROC-AUC | AP | Accuracy | F1 |
|---|---|---|---|---|
| DP | 0.5961 (0.5778–0.6144) | 0.5673 (0.5442–0.5904) | 0.5120 (0.4914–0.5327) | 0.6772 (0.6589–0.6951) |
| PDJ | 0.6730 (0.6530–0.6937) | 0.6564 (0.6317–0.6821) | 0.5943 (0.5718–0.6150) | 0.6933 (0.6723–0.7132) |
| PFC-Lin | 0.6927 (0.6702–0.7147) | 0.7022 (0.6728–0.7303) | 0.6351 (0.6145–0.6559) | 0.6479 (0.6224–0.6710) |
| PFC-Xgb | 0.7343 (0.7142–0.7548) | 0.7385 (0.7129–0.7631) | 0.6504 (0.6295–0.6705) | 0.6703 (0.6473–0.6913) |
| PFC-Snn | 0.7396 (0.7192–0.7615) | 0.7361 (0.7092–0.7623) | 0.6615 (0.6414–0.6823) | 0.6742 (0.6519–0.6965) |
| REC-Snn | 0.8530 (0.8381–0.8683) | 0.8728 (0.8574–0.8872) | 0.7680 (0.7505–0.7873) | 0.7645 (0.7444–0.7841) |
| Approach | ROC-AUC | AP | Accuracy | F1 |
|---|---|---|---|---|
| PFC-Lin-R | 0.6713 (0.6496–0.6946) | 0.6705 (0.6444–0.6969) | 0.6242 (0.6027–0.6450) | 0.6371 (0.6137–0.6591) |
| PFC-Xgb-R | 0.6948 (0.6727–0.7166) | 0.6850 (0.6595–0.7108) | 0.6353 (0.6145–0.6573) | 0.7007 (0.6802–0.7214) |
| PFC-Snn-R | 0.6909 (0.6692–0.7133) | 0.6819 (0.6548–0.7063) | 0.6265 (0.6068–0.6482) | 0.6647 (0.6428–0.6863) |
| REC-Snn-A | 0.8550 (0.8398–0.8697) | 0.8742 (0.8590–0.8888) | 0.7649 (0.7468–0.7823) | 0.7637 (0.7428–0.7821) |