cs.AISep 23, 2026

TRACER: Trajectory-Aligned Learning for Multi-Turn User Simulation

Authors: Geng Chen, Ruotong Pan, Zhirui Yang, Qiqi He, Jiawei Chen, Zhang Yunfei, Chongyuan Chen, Minxuan Lv, +4 more

Organizations: Peking University · Kuaishou Technology · Xi’an Jiaotong University · Institute of Software, Chinese Academy of Sciences · Institute of Information Engineering · Beijing University of Posts and Telecommunications

Abstract

Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. Yet current simulators often produce plausible individual responses without reproducing the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that models evolving user intent and aligns simulated trajectories with real ones. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly addressing reward sparsity and credit assignment challenges in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1 points, while outperforming all baselines on group-level conversion-rate error and semantic trajectory distance and generalizing to out-of-distribution scenarios. In human Turing tests, annotators identified TRACER conversations at near-chance accuracy. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion and response quality via simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning User Simulators with Turing Rewards

    Jun 17, 2026Yingshan Susan Wang, Cedegao E. Zhang, Linlu Qiu +5User SimulationSimulation-Based Reinforcement Learning

  2. Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

    Aug 10, 2026Bo Wang, Ruixing Zhang, Yunqi Liu +4User SimulationImitation

  3. SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

    May 8, 2026Yada Pruksachatkun, Elaine Wan, Lyanna Chen +2User SimulationMultimodal Large Language Models