As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
Figures & tables
Figure 1: Timing and semantic credit assignment in HiPLEX. (a) An LLM judge evaluates response relevance and coherence. HiPLEX normalizes these scores within each rollout group and applies the semantic advantage to the conditional content policy only where ct=cont . (b) Timing components for the active interaction type are normalized separately, and event-causal masks route their advantages to control decisions. Each cell represents an 80-ms frame.
Figure 2: Hierarchical policy factorization in HiPLEX.
Pause
Backchannel
Smooth Turn Taking
User Interruption
Model
Syn TOR ∗↓
Candor TOR ↓
TOR ↓
Freq ↑
JSD ↓
TOR ↑
Latency ↓
TOR ↑
Judge ↑
Latency ↓
Moshi ( Défossez et al., 2024 )
0.416
0.579
0.473
0.126
0.724
0.387
0.000
0.772
3.773
0.852
+ GRPO
1.000
0.454
0.309
0.110
0.737
0.992
0.000
0.955
3.942
0.660
+ HiPLEX
1.000
0.306
0.127
0.134
0.734
1.000
0.000
0.960
4.083
0.439
PersonaPlex ( Roy et al., 2026 )
0.847
0.426
0.218
0.103
0.755
0.882
0.000
0.945
2.693
0.712
+ GRPO
0.876
0.444
0.236
0.107
0.746
0.882
0.000
0.956
2.712
0.548
Table 1: Result of Full-Duplex-Bench v1. The trained Moshi rows are seed-42 terminal checkpoints. Three-seed results are reported in Table 12 . The lower block contains calibration references and is not used as a training baseline. Best per metric within a family in bold, second best underlined.
Figure 3: Moshi family on Full-Duplex-Bench v1. (a) shows post-boundary word activity, which may include speech initiated before the annotated turn end. (b) shows judged quality conditional on a qualifying interruption response. “+GRPO (Ohashi et al.)” is our evaluation of their released checkpoint.
Figure 4: Reward during training with otherwise identical settings within each model family. Moshi is stable at 4×10−6 . PersonaPlex (PP) collapses, remains suppression-dominated at half that rate, and recovers with the rebalanced configuration (RB). Dotted lines show endpoint benchmark scores.
Figure 5: Human timing in the training corpora. Each panel is one quantity the benchmark scores.
Model
Turn taking ↓
Backchannel rate ↓
Backchannel length ↓
Pause intrusion ↓
Moshi
4.698
1.798
0.292
0.068
+ GRPO
1.781
1.512
0.266
0.012
+ HiPLEX
0.902
0.327
0.238
0.080
PersonaPlex
3.551
1.738
0.080
0.088
+ GRPO
3.511
1.652
0.054
0.109
+ HiPLEX
2.627
1.187
0.168
0.091
Table 2: Wasserstein-1 distance to pooled human timing marginals (lower is closer).
Pause
Backchannel
Smooth Turn Taking
User Interruption
Model
Syn TOR ↓
Candor TOR ↓
TOR ↓
Freq ↑
JSD ↓
TOR ↑
Latency ↓
TOR ↑
Judge ↑
Latency ↓
GRPO w/o Rllm
1.000
0.384
0.509
0.102
0.744
1.000
0.000
0.955
3.859
0.611
GRPO w/ Rllm
1.000
0.454
0.309
0.110
0.737
0.992
0.000
0.955
3.942
0.660
HiPLEX w/o Rllm
1.000
0.338
0.364
0.108
0.754
1.000
0.000
0.952
2.332
0.479
HiPLEX w/ Rllm
1.000
0.306
0.127
0.134
0.734
1.000
0.000
0.960
4.083
0.439
Table 3: Matched 100-epoch ablation of Rllm with the 0 – 2 rubric. The reported HiPLEX checkpoint is reused. Best values within each method are bold.
Credit rule
Pause ↓
Int. lat. ↓
Judge ↑
Turn ↑
flat GRPO
0.454
0.660
3.942
0.992
factorized, all frames
0.264
0.874
4.330
0.958
fixed window
0.463
0.465
4.273
0.933
random, matched sparsity
0.028
3.277
3.579
0.067
event-causal ( HiPLEX )
0.306
0.439
4.083
1.000
Table 4: Moshi credit-routing ablation. Factorized rows differ only in credit eligibility. Int. lat. denotes post-interruption response latency. Random masking is excluded from ranking because its 0.067 turn rate indicates silence.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Raw timing rewards before group normalization. Both panels show valid responses starting at or after the event anchor. Left: the reported Moshi onset reward ( β=0.3 , σ=0.8 s, κ=4 s) favors prompt responses and retains a graded tail for late responses. The other curves vary one parameter at a time. Right: the delay reward decreases linearly with the fraction of the remaining window spent waiting. Each curve ends at its window boundary, where the score reaches −1 . Early speech and missing responses are scored separately as shown below the plots. The early-speech penalty shown here applies to turn taking. The two components are normalized independently rather than added as raw scores.
Window axis
Reward
Implementation name
Moshi
PersonaPlex
Turn taking
r(on)
turn_onset
1
1
r(delay)
turn_delay
1
2
r(early)
turn_early_start
1
1
Pause
r(intr)
interaction
1
1
r(frac)
listen_false_alarm
1
1
Backchanneling
r(F1)
interaction
1
1
Appendix
Table 5: Active timing components and their weights wk in the reported training configurations. Turn taking, pause, backchanneling, and interruption use 3, 2, 5, and 3 components, respectively. The component labels re-onset and re-delay denote re-entry onset and delay after an interruption. The implementation name interaction refers to different scores on different axes.
Credit rule
Pause ↓
BC TOR ↓
BC freq ↑
Turn ↑
Judge ↑
Int. lat. ↓
flat token GRPO
0.454
0.309
0.110
0.992
3.942
0.660
factorized, all frames
0.264
0.455
0.081
0.958
4.330
0.874
factorized, fixed window
0.463
0.382
0.117
0.933
4.273
0.465
factorized, matched random
0.028
0.091
0.009
0.067
3.579
3.277
factorized, event-causal
0.306
0.127
0.134
1.000
4.083
0.439
Appendix
Table 6: Full Moshi credit-routing ablation. Int. lat. denotes post-interruption response latency. The random row’s low pause TOR coincides with near-zero backchannel frequency and turn response, distinguishing silence from balanced restraint.
Split
Channel recordings
Event windows
Window hours
Train
44,292
73,561 (18.3–18.4k/axis)
271
Validation
1,665
2,048 (512/axis)
7.5
Appendix
Table 7: Training-data statistics after conversation-only filtering (§ F ) Window hours count extracted RL windows; both channel orientations of each interaction are used.
Moshi + HiPLEX
PersonaPlex + HiPLEX
Model and data
Pretrained checkpoint
kyutai/moshika-pytorch-bf16
nvidia/personaplex-7b-v1
Conditioning
none
text prompt + voice prompt
Training windows
73,561 (18.3–18.4k per axis), Seamless Interaction
Frame rate
12.5 Hz ( 0.08 s per frame)
Context prepended
0 – 30 s, probability 0.5 , linear schedule
Appendix
Table 8: Complete training configuration for the two reported HiPLEX models. Everything above the last block is shared; the last block contains the only differences, all of them consequences of the stability analysis in § 5.2 . Remaining reward-shaping constants (post-anchor onset width 0.8 s, graded-tail scale 4 s, missing-event and false-alarm penalties −0.5 , backchannel maximum duration 1.0 s, interruption grace 0.16 s) are shared by both runs.
Model
official
official + greet-excl.
no-trim, prepend 3s
trim, prepend 3s
official, no prime
official, 3s prime
Moshi (base)
0.416
0.263
0.409
0.139
0.985
—
+RL Seamless (released)
1.000
0.292
0.993
0.496
0.993
0.978
Appendix
Table 9: Pause Syn TOR for the released checkpoint and base Moshi under the tested protocol variants. “prime” = server-side silence warm-up stepped into the model before the conversation; “greet-excl.” drops a conversation-start greeting utterance before scoring. Published values are Moshi 0.445 and +RL 0.307 .
Model
Official TOR
Speaks before turn end
Post-boundary word activity ↑ / first-word offset ↓
Moshi
0.387
38.7%
0 9.2% / 0.63 s
+ GRPO
0.992
99.2%
69.7% / 1.27 s
+RL Seamless (released)
0.992
99.2%
57.1% / 1.65 s
+ HiPLEX
1.000
100.0%
98.3% / 0.57 s
PersonaPlex
0.882
89.1%
34.5% / 0.85 s
+ GRPO
0.882
89.9%
36.1% / 0.70 s
Appendix
Table 10: Post-boundary word activity on the same 119 generations. “Speaks before turn end” is the fraction of clips whose first ASR word precedes the user’s turn end. Post-boundary word activity applies the elapsed-time-or-word-count criterion to words starting after that boundary; it may include continuation of earlier speech. The reported offset is the median time to the first post-boundary word among qualifying clips. Bold and underline mark the best and second-best values within each family for the post-boundary measures.
Metric
n
HiPLEX
GRPO
Difference, 95% CI
CANDOR pause TOR ↓
216
0.306
0.454
−0.148[−0.222,−0.074]
Backchannel TOR ↓
55
0.127
0.309
−0.182[−0.327,−0.036]
Appendix
Table 11: Paired bootstrap over Full-Duplex-Bench v1 clips for Moshi. A negative difference favors HiPLEX for both takeover-rate metrics.
Credit rule
Pause TOR ↓
BC TOR ↓
BC freq. ↑
Turn TOR ↑
Interrupt TOR ↑
Judge ↑
Int. lat. ↓
GRPO
0.495±0.050
0.297±0.011
0.131±0.020
0.992±0.000
0.985±0.026
3.927±0.033
0.631±0.026
fixed window
0.500±0.060
0.261±0.107
0.132±0.014
0.958±0.030
0.990±0.010
4.300±0.024
0.480±0.071
event-causal ( HiPLEX )
0.329±0.083
0.109±0.018
0.145±0.020
1.000±0.000
0.987±0.023
3.914±0.260
0.466±0.061
Appendix
Table 12: Moshi replication over training seeds 42 , 43 , and 44 . Entries are mean ± sample standard deviation. Judge denotes the gpt-4o-2024-08-06 response-quality score. Int. lat. denotes post-interruption response latency. Bold marks the best mean in each column.
Method
η
Pause ↓
BC TOR ↓
BC freq ↑
Turn ↑
Judge ↑
Int. lat. ↓
GRPO
1×10−6
0.454
0.309
0.110
0.992
3.942
0.660
GRPO
5×10−7
0.574
0.400
0.166
0.924
3.894
1.067
GRPO
2×10−6
0.528
0.073
0.174
0.992
3.872
0.491
GRPO
4×10−6
0.431
0.000
0.171
0.992
3.890
0.464
HiPLEX
4×10−6
0.306
0.127
0.134
1.000
4.083
0.439
Appendix
Table 13: The GRPO baseline swept over learning rates, including the rate HiPLEX uses. Everything except the step size is held fixed. Int. lat. denotes post-interruption response latency. Best values are bold and second-best values are underlined; ties share the same marking.
Reference KL
Decisions per update
System
control
content contrib.
control
content
Repetition-3 ↓
GRPO w/o Rllm
2.6×10−2 (single policy)
∼ 200 (all frames)
0.007
GRPO
2.1×10−2 (single policy)
∼ 200 (all frames)
0.009
HiPLEX w/o Rllm
9.8×10−2
1.0×10−2
38.7
9.1
0.098
HiPLEX
8.6×10−2
1.4×10−2
40.6
8.6
0.002
Appendix
Table 14: What each objective actually moves. Reference KL reports control KL and the content KL contribution weighted by πθctrl(cont) . KL values and decision counts are averaged over the optimizer updates of the final 30% of training. Repetition is the fraction of trigrams that recur within a response on the Full-Duplex-Bench v1 interruption task. The flat-baseline rows report total token-policy KL. Separate control and content KL were not recorded for these rows. The best repetition value is bold and the second-best is underlined. These diagnostics describe the reported, balanced reward setting rather than invariance to reward scaling.
Quantity
Statistic
Seamless (improvised)
Seamless (naturalistic)
Fisher
Own pause (s)
mean ± sd
1.07 ± 1.02
0.85 ± 0.81
1.40 ± 1.22
median (IQR)
0.64 (0.39–1.35)
0.55 (0.36–0.96)
0.92 (0.46–1.99)
p10 / p90
0.26 / 2.60
0.26 / 1.80
0.29 / 3.45
Backchannels per minute
mean ± sd
4.54 ± 2.24
2.82 ± 2.11
3.86 ± 1.29
median (IQR)
4.21 (2.77–6.35)
2.38 (1.28–3.91)
3.80 (3.00–4.68)
p10 / p90
1.82 / 7.51
0.43 / 6.08
2.30 / 5.41
Appendix
Table 15: Distribution descriptors for the quantities in Figure 5 . Spread matters as much as location here: the standard deviations are comparable to the means, so no single value characterises “correct” timing in any of the corpora. The overlap row is the fraction of floor transfers where the next speaker starts before the current one stops.
Pause
Backchannel
Smooth Turn Taking
User Interruption
Judge rubric
Syn TOR ↓
Candor TOR ↓
TOR ↓
Freq ↑
JSD ↓
TOR ↑
Latency ↓
TOR ↑
Judge ↑
Latency ↓
HiPLEX with 0 – 2
1.000
0.306
0.127
0.134
0.734
1.000
0.000
0.960
4.083
0.439
HiPLEX with 0 – 5
1.000
0.407
0.200
0.131
0.733
1.000
0.000
0.955
4.021
0.445
Appendix
Table 16: Effect of the judge rubric range on HiPLEX . The 0 – 2 rubric performs better on most reported metrics, with ties on several metrics and a slightly higher backchannel JSD. This comparison does not establish why the rubrics differ. Changes in rating patterns may affect the training signal, but a smaller raw-score spread alone does not imply smaller standardized advantages. Best values are bold and second-best values are underlined; ties share the same marking.
Variant (checkpoint)
Pause Candor TOR ↓
BC TOR ↓
BC Freq ↑
Turn TOR ↑
Judge ↑
semantic ×1.5 , loose content KL (e20)
0.389
0.583
0.071
0.467
4.488
semantic ×1.5 , loose content KL (e30)
0.167
0.000
0.013
0.067
3.955
semantic ×1.5 , loose content KL (e90)
0.088
0.000
0.046
0.017
3.500
semantic ×2 , content weight 1.5 (e90)
0.181
0.000
0.062
0.076
—
semantic ×2 , content weight 1.5 (e100)
0.199
0.036
0.059
0.017
4.320
continued from e100 of the main run (e130)
0.074
0.000
0.011
0.033
3.458
Appendix
Table 17: PersonaPlex sensitivity to stronger semantic feedback and continued training. A tested checkpoint reaches judged quality 4.488 alongside a higher backchannel takeover rate, while several later checkpoints have low turn-response rates. These results complement the diagnostics in Table 14 by examining settings outside the reported configuration. The final blue row is the configuration we report in the main table. Best values are bold and second-best values are underlined; ties share the same marking. Rows are scored on a 25% sample of the benchmark, which is why one judge entry is missing.
Key Laboratory of Intelligent Information Processing Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS) · Key Laboratory of AI Safety, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, China