ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
Organizations: ModelsLive Inc.
Abstract
High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.
Figures & tables
| Tier | Definition | Example slots |
|---|---|---|
| 1 | Highest-priority, decision-critical information that should be completed before recommendation or specialist handoff. | Chief complaint, target anatomy, contraindications, budget |
| 2 | Preference and feasibility constraints that refine the advisory process after Tier-1 information is covered. | Desired effect, recovery window, risk preference, procedure interest |
| 3 | Supplementary and logistical information used for profile completion, appointment arrangement, and follow-up. | Skin type, medical history, appointment window, contact information |
| LLM Judge Score | Human by Dimension | ||||||||||
| Grp | System | SFR | BIV | Emp | Comp | Acc | Struct | Bound | Avg. | ||
| A | GPT-5.5 | 85.7 | 80.2 | 80.3 | 84.6 | 88.7 | 86.1 | 86.3 | 82.5 | 85.3 | 85.8 |
| Gemini-3.1 | 76.2 | 71.8 | 72.7 | 65.7 | 85.2 | 83.3 | 79.5 | 73.1 | 72.5 | 78.7 | |
| Grok-4.20 | 81.3 | 72.4 | 75.2 | 72.5 | 78.5 | 88.3 | 88.3 | 78.9 | 78.1 | 82.4 | |
| DeepSeek-V4 | 74.5 | 79.2 | 74.1 | 63.8 | 87.1 | 83.5 | 83.5 | 75.2 | 77.5 | 81.4 | |
| Claude-4.6 | 88.2 | 82.5 | 82.4 | 80.2 | 72.5 | 91.1 | 84.6 | 84.5 | 79.4 | 82.4 | |
| Grp | System | SFR 1 | SFR 2 | SFR 3 | Eff |
|---|---|---|---|---|---|
| A | GPT-5.5 | 78.2 | 73.8 | 88.9 | 8.5 |
| B | Qwen3-32B | 68.9 | 58.1 | 83.6 | 13.4 |
| C | AutoGen | 63.3 | 57.8 | 77.5 | 15.7 |
| D | ThuRunel | 90.3 | 82.2 | 92.4 | 5.2 |
| Scoring standard |
| 90–100: Excellent. The experience strongly satisfies the criterion. 80–89: Good. The experience satisfies the criterion with minor issues. 70–79: Adequate. The experience is acceptable but has weaknesses. 60–69: Weak. The experience only partially satisfies the criterion. 0–59: Poor. The experience fails the criterion or is clearly unsatisfactory. |
| Configuration | SFR | BIV | ||
|---|---|---|---|---|
| ThuRunel | 92.7 | 90.2 | 88.3 | 91.4 |
| w/o CoT | 84.3 | 81.8 | 80.1 | 82.1 |
| w/o Rou | 88.9 | 85.2 | 85.3 | 85.3 |
| w/o PSR | 82.2 | 80.5 | 75.3 | 81.2 |
| w/o BM | 79.1 | 77.2 | 74.1 | 80.9 |
| Grp | System | Nat | Und | Trust | Will | Overall |
|---|---|---|---|---|---|---|
| A | GPT-5.5 | 4.4 | 3.7 | 4.1 | 3.8 | 4.2 |
| B | Qwen3-32B | 3.8 | 3.3 | 3.8 | 3.5 | 3.5 |
| C | AutoGen | 3.2 | 3.1 | 3.9 | 3.1 | 3.2 |
| D | ThuRunel | 4.5 | 4.4 | 4.5 | 4.3 | 4.5 |
| Cond. | Sys | SFR 1 | Cls | Rep | |
|---|---|---|---|---|---|
| Direct | GPT-5.5 | 95.2 | 88.6 | - | 3.2 |
| ThuRunel | 97.1 | 96.7 | - | 1.3 | |
| Underspecified | GPT-5.5 | 90.4 | 81.7 | 64.8 | 18.5 |
| ThuRunel | 93.5 | 93.8 | 79.2 | 9.3 | |
| Deflective | GPT-5.5 | 82.3 | 70.4 | 55.7 | 24.2 |
| ThuRunel | 90.3 | 89.6 | 73.8 | 11.4 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.