HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
Organizations: Qualcomm AI Research · KAIST AI
Abstract
As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
Figures & tables
| Pause | Backchannel | Smooth Turn Taking | User Interruption | |||||||
| Model | Syn TOR | Candor TOR | TOR | Freq | JSD | TOR | Latency | TOR | Judge | Latency |
| Moshi ( Défossez et al., 2024 ) | 0.416 | 0.579 | 0.473 | 0.126 | 0.724 | 0.387 | 0.000 | 0.772 | 3.773 | 0.852 |
| + GRPO | 1.000 | 0.454 | 0.309 | 0.110 | 0.737 | 0.992 | 0.000 | 0.955 | 3.942 | 0.660 |
| + HiPLEX | 1.000 | 0.306 | 0.127 | 0.134 | 0.734 | 1.000 | 0.000 | 0.960 | 4.083 | 0.439 |
| PersonaPlex ( Roy et al., 2026 ) | 0.847 | 0.426 | 0.218 | 0.103 | 0.755 | 0.882 | 0.000 | 0.945 | 2.693 | 0.712 |
| + GRPO | 0.876 | 0.444 | 0.236 | 0.107 | 0.746 | 0.882 | 0.000 | 0.956 | 2.712 | 0.548 |
| Model | Turn taking | Backchannel rate | Backchannel length | Pause intrusion |
| Moshi | 4.698 | 1.798 | 0.292 | 0.068 |
| + GRPO | 1.781 | 1.512 | 0.266 | 0.012 |
| + HiPLEX | 0.902 | 0.327 | 0.238 | 0.080 |
| PersonaPlex | 3.551 | 1.738 | 0.080 | 0.088 |
| + GRPO | 3.511 | 1.652 | 0.054 | 0.109 |
| + HiPLEX | 2.627 | 1.187 | 0.168 | 0.091 |
| Pause | Backchannel | Smooth Turn Taking | User Interruption | |||||||
| Model | Syn TOR | Candor TOR | TOR | Freq | JSD | TOR | Latency | TOR | Judge | Latency |
| GRPO w/o | 1.000 | 0.384 | 0.509 | 0.102 | 0.744 | 1.000 | 0.000 | 0.955 | 3.859 | 0.611 |
| GRPO w/ | 1.000 | 0.454 | 0.309 | 0.110 | 0.737 | 0.992 | 0.000 | 0.955 | 3.942 | 0.660 |
| HiPLEX w/o | 1.000 | 0.338 | 0.364 | 0.108 | 0.754 | 1.000 | 0.000 | 0.952 | 2.332 | 0.479 |
| HiPLEX w/ | 1.000 | 0.306 | 0.127 | 0.134 | 0.734 | 1.000 | 0.000 | 0.960 | 4.083 | 0.439 |
| Credit rule | Pause | Int. lat. | Judge | Turn |
| flat GRPO | 0.454 | 0.660 | 3.942 | 0.992 |
| factorized, all frames | 0.264 | 0.874 | 4.330 | 0.958 |
| fixed window | 0.463 | 0.465 | 4.273 | 0.933 |
| random, matched sparsity | 0.028 | 3.277 | 3.579 | 0.067 |
| event-causal ( HiPLEX ) | 0.306 | 0.439 | 4.083 | 1.000 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Window axis | Reward | Implementation name | Moshi | PersonaPlex |
| Turn taking | turn_onset | 1 | 1 | |
| turn_delay | 1 | 2 | ||
| turn_early_start | 1 | 1 | ||
| Pause | interaction | 1 | 1 | |
| listen_false_alarm | 1 | 1 | ||
| Backchanneling | interaction | 1 | 1 |
| Credit rule | Pause | BC TOR | BC freq | Turn | Judge | Int. lat. |
| flat token GRPO | 0.454 | 0.309 | 0.110 | 0.992 | 3.942 | 0.660 |
| factorized, all frames | 0.264 | 0.455 | 0.081 | 0.958 | 4.330 | 0.874 |
| factorized, fixed window | 0.463 | 0.382 | 0.117 | 0.933 | 4.273 | 0.465 |
| factorized, matched random | 0.028 | 0.091 | 0.009 | 0.067 | 3.579 | 3.277 |
| factorized, event-causal | 0.306 | 0.127 | 0.134 | 1.000 | 4.083 | 0.439 |
| Split | Channel recordings | Event windows | Window hours |
| Train | 44,292 | 73,561 (18.3–18.4k/axis) | 271 |
| Validation | 1,665 | 2,048 (512/axis) | 7.5 |
| Moshi + HiPLEX | PersonaPlex + HiPLEX | |
| Model and data | ||
| Pretrained checkpoint | kyutai/moshika-pytorch-bf16 | nvidia/personaplex-7b-v1 |
| Conditioning | none | text prompt + voice prompt |
| Training windows | 73,561 (18.3–18.4k per axis), Seamless Interaction | |
| Frame rate | Hz ( s per frame) | |
| Context prepended | – s, probability , linear schedule | |
| Model | official | official + greet-excl. | no-trim, prepend 3s | trim, prepend 3s | official, no prime | official, 3s prime |
| Moshi (base) | 0.416 | 0.263 | 0.409 | 0.139 | 0.985 | — |
| +RL Seamless (released) | 1.000 | 0.292 | 0.993 | 0.496 | 0.993 | 0.978 |
| Model | Official TOR | Speaks before turn end | Post-boundary word activity / first-word offset |
| Moshi | 0.387 | 38.7% | 0 9.2% / 0.63 s |
| + GRPO | 0.992 | 99.2% | 69.7% / 1.27 s |
| +RL Seamless (released) | 0.992 | 99.2% | 57.1% / 1.65 s |
| + HiPLEX | 1.000 | 100.0% | 98.3% / 0.57 s |
| PersonaPlex | 0.882 | 89.1% | 34.5% / 0.85 s |
| + GRPO | 0.882 | 89.9% | 36.1% / 0.70 s |
| Metric | HiPLEX | GRPO | Difference, 95% CI | |
| CANDOR pause TOR | 216 | 0.306 | 0.454 | |
| Backchannel TOR | 55 | 0.127 | 0.309 |
| Credit rule | Pause TOR | BC TOR | BC freq. | Turn TOR | Interrupt TOR | Judge | Int. lat. |
| GRPO | |||||||
| fixed window | |||||||
| event-causal ( HiPLEX ) |
| Method | Pause | BC TOR | BC freq | Turn | Judge | Int. lat. | |
| GRPO | 0.454 | 0.309 | 0.110 | 0.992 | 3.942 | 0.660 | |
| GRPO | 0.574 | 0.400 | 0.166 | 0.924 | 3.894 | 1.067 | |
| GRPO | 0.528 | 0.073 | 0.174 | 0.992 | 3.872 | 0.491 | |
| GRPO | 0.431 | 0.000 | 0.171 | 0.992 | 3.890 | 0.464 | |
| HiPLEX | 0.306 | 0.127 | 0.134 | 1.000 | 4.083 | 0.439 |
| Reference KL | Decisions per update | ||||
| System | control | content contrib. | control | content | Repetition-3 |
| GRPO w/o | (single policy) | 200 (all frames) | 0.007 | ||
| GRPO | (single policy) | 200 (all frames) | 0.009 | ||
| HiPLEX w/o | 38.7 | 9.1 | 0.098 | ||
| HiPLEX | 40.6 | 8.6 | 0.002 | ||
| Quantity | Statistic | Seamless (improvised) | Seamless (naturalistic) | Fisher |
| Own pause (s) | mean sd | 1.07 1.02 | 0.85 0.81 | 1.40 1.22 |
| median (IQR) | 0.64 (0.39–1.35) | 0.55 (0.36–0.96) | 0.92 (0.46–1.99) | |
| p10 / p90 | 0.26 / 2.60 | 0.26 / 1.80 | 0.29 / 3.45 | |
| Backchannels per minute | mean sd | 4.54 2.24 | 2.82 2.11 | 3.86 1.29 |
| median (IQR) | 4.21 (2.77–6.35) | 2.38 (1.28–3.91) | 3.80 (3.00–4.68) | |
| p10 / p90 | 1.82 / 7.51 | 0.43 / 6.08 | 2.30 / 5.41 |
| Pause | Backchannel | Smooth Turn Taking | User Interruption | |||||||
| Judge rubric | Syn TOR | Candor TOR | TOR | Freq | JSD | TOR | Latency | TOR | Judge | Latency |
| HiPLEX with – | 1.000 | 0.306 | 0.127 | 0.134 | 0.734 | 1.000 | 0.000 | 0.960 | 4.083 | 0.439 |
| HiPLEX with – | 1.000 | 0.407 | 0.200 | 0.131 | 0.733 | 1.000 | 0.000 | 0.955 | 4.021 | 0.445 |
| Variant (checkpoint) | Pause Candor TOR | BC TOR | BC Freq | Turn TOR | Judge |
| semantic , loose content KL (e20) | 0.389 | 0.583 | 0.071 | 0.467 | 4.488 |
| semantic , loose content KL (e30) | 0.167 | 0.000 | 0.013 | 0.067 | 3.955 |
| semantic , loose content KL (e90) | 0.088 | 0.000 | 0.046 | 0.017 | 3.500 |
| semantic , content weight (e90) | 0.181 | 0.000 | 0.062 | 0.076 | — |
| semantic , content weight (e100) | 0.199 | 0.036 | 0.059 | 0.017 | 4.320 |
| continued from e100 of the main run (e130) | 0.074 | 0.000 | 0.011 | 0.033 | 3.458 |