Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
Organizations: University of Michigan Ann Arbor, MI, USA
Abstract
High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate latent conversation structure using a Gaussian mixture density over key interaction features (e.g., confusion episode dynamics, scaffolding directive balance) enables sampling expert strategies in proportion to their prevalence in the training data. Across two structurally distinct domains, tutoring and emotional support dialogues, our GFlow based synthetic data generation approach offers a better balance of fidelity, mode coverage and authenticity than reinforcement-learning and end to end LLM baselines, without copying training data. Evaluated on three downstream outcome prediction tasks, classifiers trained on synthetic GFlowNet generated conversations provide a stronger training signal than competitive synthesis baselines.
Figures & tables
| MathMentorDB | ESConv | |
|---|---|---|
| Domain | 1:1 math tutoring | emotional-support dialogue |
| Roles | tutor, student | supporter, seeker |
| Move vocab. | 24 | 9 |
| Train convs. | 2,500 | 1,040 |
| Held-out convs. | 1,000 | 260 |
| Avg. turns/conv. | 23.8 | 29.5 |
| Math tutoring | ESConv | |||||
| Condition | pMSE (all features) | pMSE (excl. surface) | pMSE (all features) | pMSE (excl. surface) | ||
| Real conversations (null) | 0.063 | 2.0 | 1.7 | 0.003 | 4.6 | 3.2 |
| GPT-5.6-terra realization | ||||||
| LLM only ( B2 , best tier) | 0.367 | 97.3 | 94.3 | 0.067 | 98.4 | 95.6 |
| RL + LLM (REINFORCE) ‡ | 0.267 | 95.3 | 94.8 | 0.108 | 94.2 | 80.8 |
| GFlow + LLM (ours) | 0.145 | 81.1 | 65.1 | 0.015 | 97.4 | 58.1 |
| pMSE | Auth. | Outlier | |||
|---|---|---|---|---|---|
| Real held-out (reference) | 9.4 4.5 | 0.94 0.03 | 0.96 0.02 | 0.70 | 0.050 |
| LLM only (B2), GPT-5.6-terra | 95.0 3.7 | 0.44 0.04 | 0.09 0.03 | 0.96 | 0.233 |
| REINFORCE, GPT-5.6-terra | 94.1 4.1 | 0.81 0.03 | 0.17 0.04 | 0.92 | 0.006 |
| GFlowNet (TB), GPT-5.6-terra | 71.5 3.3 | 0.87 0.04 | 0.57 0.07 | 0.85 | 0.117 |
| LLM only (B2), Gemini 2.5 Flash | 96.0 2.6 | 0.21 0.03 | 0.00 0.01 | 0.99 | 0.738 |
| REINFORCE, Gemini 2.5 Flash | 93.4 2.5 | 0.85 0.02 | 0.16 0.04 | 0.90 | 0.007 |
| GPT-5.6-terra Realization | Gemini 2.5 Flash Realization | |||||||
|---|---|---|---|---|---|---|---|---|
| Label | Real oracle | LLM | RL | GFlow | LLM | RL | GFlow | Majority floor |
| resolved | 0.79 0.03 | 0.50 0.03 | 0.29 0.01 | 0.65 0.03 | 0.57 0.03 | 0.29 0.01 | 0.64 0.03 | 0.37 |
| has_breakthrough | 0.92 0.02 | 0.71 0.03 | 0.41 0.01 | 0.72 0.03 | 0.69 0.03 | 0.41 0.01 | 0.81 0.03 | 0.41 |
| has_off_topic | 0.89 0.04 | 0.48 0.00 | 0.48 0.00 | 0.50 0.03 | 0.48 0.00 | 0.48 0.00 | 0.48 0.00 | 0.48 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | MathMentorDB (tutoring), 47 feat. | ESConv (emotional support), 30 feat. |
|---|---|---|
| Surface (6 / 6) | avg_student_len, std_student_len, avg_tutor_len, std_tutor_len, n_messages, turn_ratio | avg_seeker_len, std_seeker_len, avg_supporter_len, std_supporter_len, n_messages, turn_ratio |
| Content (21 / 8) | direct_rate, scaffolding_rate, correction_rate, confirm_pos_rate, confirm_neg_rate, tutor_question_rate, knowledge_check_rate, attempt_rate, knowledge_gap_rate, understanding_rate, breakthrough_rate, explain_problem_rate, explain_reasoning_rate, student_question_rate, student_confirm_rate, frustration_rate, encouragement_rate, empathy_rapport_rate, confidence_express_rate, acknowledgment_rate, knowledge_recall_rate | question_rate, restatement_rate, reflection_rate, self_disclosure_rate, affirmation_rate, suggestion_rate, information_rate, others_rate |
| Structure (7 / 4) | pct_tutor_academic, pct_student_academic, pct_non_academic, pct_socio_emotional, greeting_rate, platform_rate, smalltalk_rate | pct_empathic, pct_directive, pct_exploratory, pct_personal |
| Dynamics (10 / 9) | n_confusion_episodes, avg_confusion_duration, max_confusion_duration, instant_resolution_rate, extended_confusion_rate, scaffolding_density_first_third, scaffolding_density_last_third, scaffolding_density_shift, first_attempt_position, first_understanding_position | empathy_density_first_third, empathy_density_last_third, directive_density_first_third, directive_density_last_third, empathy_to_directive_shift, strategy_entropy, n_strategy_types, suggestion_position, first_empathy_position |
| Outcome (3 / 3) | resolved, has_breakthrough, has_off_topic | resolved_esc, has_suggestion, has_reflection |
| Tier | Combined | Surface | Content | Structure | Dynamics | Outcome |
|---|---|---|---|---|---|---|
| Tutoring B1 (naive) | 98.7 | 75.0 | 95.9 | 79.0 | 71.0 | 10.9 |
| Tutoring B2 (scenario) | 98.4 | 72.2 | 90.5 | 61.3 | 47.9 | 5.8 |
| Tutoring B3 (scenario+mode) | 98.6 | 70.3 | 89.5 | 68.3 | 44.8 | 6.8 |
| ESConv B1 (naive) | 95.2 | 96.1 | 75.3 | 69.3 | 73.6 | 68.3 |
| ESConv B2 (scenario) † | 95.0 | 87.2 | 62.4 | 68.1 | 70.4 | 51.8 |
| ESConv B3 (scenario+mode) | 96.8 | 98.1 | 78.8 | 69.7 | 72.5 | 53.4 |
| Condition | pMSE | Auth. | Near-copy | Outlier | ||
|---|---|---|---|---|---|---|
| Real held-out (reference) | 7.0 1.6 | 0.96 0.01 | 0.96 0.01 | 0.67 | 0.072 | 0.050 |
| LLM only (B2), GPT-5.6-terra | 95.5 2.3 | 0.43 0.03 | 0.06 0.02 | 0.94 | 0.000 | 0.123 |
| REINFORCE, GPT-5.6-terra | 96.0 1.6 | 0.27 0.02 | 0.06 0.02 | 0.96 | 0.001 | 0.282 |
| GFlowNet (TB), GPT-5.6-terra | 93.8 2.7 | 0.29 0.02 | 0.22 0.03 | 0.93 | 0.001 | 0.410 |
| LLM only (B2), Gemini 2.5 Flash | 91.6 2.0 | 0.78 0.04 | 0.42 0.04 | 0.67 | 0.025 | 0.013 |
| REINFORCE, Gemini 2.5 Flash | 92.7 2.0 | 0.52 0.04 | 0.12 0.02 | 0.75 | 0.025 | 0.103 |