cs.LGJul 22, 2026

SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

Authors: Chenyu Zhang

Organizations: Li Auto Inc.

Abstract

Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production. We identify three fundamental gaps responsible for this disconnect: offline-online metric mismatch, pointwise-listwise objective misalignment, and temporal distribution drift. To address these gaps, we propose SalesLoop, a reinforcement learning framework that establishes a closed feedback loop between model predictions and real-world business outcomes. Our approach introduces (1) a performance-aware reward that encodes conversion outcomes weighted by ranking position and conversion velocity, and (2) Discriminative GRPO, a listwise optimization objective that adapts Group Relative Policy Optimization to discriminative ranking models. SalesLoop improves NDCG@K by +7.9% and P@K by +15.8% over the strongest static baseline. A 160-day production A/B test at a New Energy Vehicle manufacturer, spanning 16.5M leads and 280 sales specialists across two provincial markets, validates statistically significant cumulative lift of +4.7% (p=0.047p=0.047) and +8.7% (p=0.002p=0.002). In production, the ranking backbone achieves Top-10% recall of 44.1% and surfaces high-intent leads at 2.3×2.3\times the conversion rate of specialist baselines.

Explore similar work

CardsList
  1. Active Learners as Efficient PRP Rerankers

    May 14, 2026Jeremías Figueiredo Paschmann, Juan Kaplan, Francisco Nattero +3RerankingRanking