Proactive Dialogue Policy Optimization via Cognitive-State Transition
Organizations: Northwestern Polytechnical University
Abstract
Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators do not explicitly model the evolution of user cognition, limiting the consistency and state dependence of feedback across turns. Moreover, representing each action only by a high-level strategy label overlooks the large utterance space and cannot distinguish alternative realizations of the same strategy. To this end, we jointly design a nitive User ulator and ognitive-tate ransition--Driven olicy ptimization . Cog-Sim maintains the user's cognitive and affective states and generates responses through constrained state transitions across turns, so feedback depends on both the realized utterance and the user's current state. CSTPO organizes each action as a hierarchical strategy--utterance representation: a high-level strategy label constrains utterance sampling, and utterances are optimized within each label. Sparse complete-branch sampling reuses shared dialogue prefixes and estimates separate strategy-level and utterance-level advantages, enabling fine-grained optimization at both levels. Across three tasks, Cog-Sim exhibits monotonic dose--response relationships and is preferred over prompt-based simulators for naturalness. CSTPO improves Qwen3-14B's performance to a level comparable to that of GPT-5.5-based planning methods.
Figures & tables
| ESConv | P4G | CB | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Comparator | Nat. | Cons. | App. | Nat. | Cons. | App. | Nat. | Cons. | App. |
| Persona Zhao et al. (2024) | +.66 | +.55 | +.53 | +.63 | +.26 | +.27 | +.33 | -.01 | -.01 |
| Resistance Zhang et al. (2024) | +.41 | +.21 | +.23 | +.39 | +.10 | +.10 | +.42 | -.09 | +.04 |
| BDI-only | +.60 | +.47 | +.47 | +.60 | +.23 | +.25 | +.47 | +.03 | +.05 |
| ESConv | P4G | CB | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Backbone | Comm. | SL | Deal | |||||
| Standard | GPT | 2.180 | 2.627 | .601 | 45% | .450 | .571 | 66% | .484 |
| Proactive Deng et al. (2023) | GPT | 2.487 | 2.250 | .592 | 43% | .430 | .511 | 65% | .457 |
| ProCoT ( Deng et al., 2023 ) | GPT | 2.490 | 2.430 | .615 | 44% | .440 | .509 | 61% | .435 |
| AnE ( Zhang et al., 2023 ) | GPT | 1.963 | 3.093 | .632 | 42% | .420 | .516 | 61% | .433 |
| MI-Prompt ( Chen et al., 2023 ) | GPT | 2.163 | 1.620 | .473 | 40% | .400 | .541 | 62% | .456 |
| Setting | ESConv | P4G | CB |
|---|---|---|---|
| Standard | .442 | .310 | .244 |
| SFT Init. | .494 | .350 | .321 |
| CSTPO | .555 | .410 | .435 |
| ESConv | P4G | CB | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Comparator | Ident. | Comf. | Sugg. | Pers. | Coh. | Nat. | Pers. | Coh. | Nat. |
| Standard | +.23 | +.29 | +.06 | +.13 | +.05 | -.09 | +.15 | +.07 | -.10 |
| ProCoT ( Deng et al., 2023 ) | +.11 | +.13 | -.11 | +.25 | +.10 | -.02 | +.31 | +.11 | -.03 |
| PPDPP ( Deng et al., 2024 ) | +.15 | +.17 | +.03 | +.21 | +.07 | -.03 | +.07 | +.05 | -.04 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Section | Topic | Contents |
|---|---|---|
| A | Cog-Sim details | State representation, cognitive routing, constrained transitions, affect updates, and recoverable checkpoints. |
| B | CSTPO details | Training algorithm, same-state continuations, Critic targets, field masks, and policy updates. |
| C | Theoretical properties | Training-state refinement, orthogonal hierarchical credit decomposition, and action aliasing. |
| D | Data and seeds | Dataset statistics, strategy spaces, seed schema, information boundaries, and seed quality control. |
| E | Reproducibility | SFT and RL configurations, model backends, compute, terminal utilities, stopping rules, and statistics. |
| F | Additional results | Complete Qwen3-14B results, descriptive differences from Standard, and the fine-grained strategy mapping underlying the main-paper analysis. |
| Route | Judgment | Permitted change | |||
|---|---|---|---|---|---|
| Central | Accept | Substantive evidence may change directly relevant beliefs; desires and intentions may move when supported. | 1.0 | .5 | .8 |
| Central | Noncommit | Relevant beliefs may loosen without reversal; desires remain stable and no strong commitment is created. | .4 | .2 | .3 |
| Central | Reject | The rejected claim cannot be reinforced; counter-beliefs may emerge and a target intention may weaken. | .4 | .2 | .5 |
| Peripheral | Accept | Cue-related trust or norm beliefs may strengthen; core issue beliefs and desires remain nearly fixed. | .3 | .1 | .6 |
| Peripheral | Noncommit | Only small cue-driven belief or intention movement is permitted. | .2 | 0 | .2 |
| Peripheral | Reject | Cue-related trust may weaken; core beliefs and desires remain stable and a target intention may decrease. | .2 | 0 | .5 |
| Task | Train/Dev/Test | Actor/User | Strategies | Task objective |
|---|---|---|---|---|
| ESConv | 1196/150/150 | supporter/seeker | 8 | Improve the user’s emotion or hope and support a feasible action plan. |
| P4G | 813/102/102 | persuader/persuadee | 13 | Obtain a clear, voluntary donation commitment that is not withdrawn. |
| CB | 5247/597/838 | buyer/seller | 6 | Reach a valid agreement while optimizing the buyer-side price utility. |
| ESConv | P4G | CB | ||||||
| Method | Comm. | SL | Deal | |||||
| Standard | 2.083 | 2.227 | .539 | 39% | .390 | .400 | 49% | .353 |
| Proactive ( Deng et al., 2023 ) | 2.127 | 1.860 | .498 | 31% | .310 | .495 | 54% | .412 |
| ProCoT ( Deng et al., 2023 ) | 2.177 | 2.287 | .558 | 36% | .360 | .386 | 45% | .327 |
| AnE ( Zhang et al., 2023 ) | 1.540 | 1.717 | .407 | 35% | .320 | .443 | 57% | .365 |
| MI-Prompt ( Chen et al., 2023 ) | 1.833 | 1.490 | .415 | 42% | .420 | .442 | 61% | .370 |
| Method | ESConv | P4G | CB |
|---|---|---|---|
| SFT Init. | |||
| CSTPO |
| Task | Strategy labels |
|---|---|
| Advance — Directly advance the task objective | |
| ESConv | Providing Suggestions |
| P4G | Donate Request ; Logical Appeal ; Credibility Appeal ; Foot in the Door |
| CB | Propose Price ; Stance |
| Soften — Build rapport and reduce interactional tension | |
| ESConv | Affirmation and Reassurance ; Reflection of Feelings ; Self-disclosure |
| Strategy | Definition |
|---|---|
| Question | Ask about the user’s situation, feelings, needs, or preferences to elicit relevant information. |
| Restatement or Paraphrasing | Restate the user’s expressed content in different words to confirm or clarify understanding. |
| Reflection of Feelings | Identify and reflect the emotion conveyed by the user. |
| Self-disclosure | Share a relevant personal experience or feeling to build rapport. |
| Affirmation and Reassurance | Validate the user’s feelings or strengths and provide encouragement or reassurance. |
| Providing Suggestions | Recommend concrete coping actions or feasible next steps. |
| Strategy | Definition |
|---|---|
| Donate Request | Propose, request, confirm, or negotiate a donation, including its amount or reasons for declining. |
| Logical Appeal | Use reasons or evidence to explain why donating is useful or impactful. |
| Credibility Appeal | Establish the charity’s trustworthiness, reputation, or accountability. |
| Foot in the Door | Seek a small initial commitment that can facilitate a subsequent donation. |
| Praise User | Compliment the user or affirm the user’s prosocial qualities. |
| Social | Perform conversational functions such as greeting, acknowledgment, thanks, or closing. |
| Strategy | Definition |
|---|---|
| Propose Price | Make an initial, counter, or deliberately underspecified price proposal. |
| Stance | Express agreement, disagreement, or insistence toward a negotiating position. |
| Social | Greet the counterpart or perform another rapport-oriented conversational act. |
| Inquire | Request information about the item, transaction, or counterpart’s preferences. |
| Inform | Provide relevant item, transaction, preference, or constraint information without proposing a price. |
| Other | Use a negotiation act that is not captured by the preceding categories. |
| Setting | ESConv | P4G | CB |
|---|---|---|---|
| Standard | .442 | .310 | .244 |
| SFT Init. | .494 | .350 | .321 |
| CSTPO | .555 | .410 | .435 |
| Task | Dimension | vs. Standard | vs. ProCoT | vs. PPDPP |
|---|---|---|---|---|
| ESConv | Identification | |||
| Comforting | ||||
| Suggestion | ||||
| P4G | Persuasive | |||
| Coherent | ||||
| Natural |