Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.
Figures & tables
Figure 1: VAP-BC with the optional components evaluated in Section 5.3 : listener/floor-holder FiLM conditioning, rank-8 adapters, listener-identity adversary, and lexical stream.
Figure 2Figure 3
Method (listener + floor-holder base)
Unseen AUROC
Δ
Baseline
0.7385
—
+ Adapter ( r=8 )
0.7307
−0.0078
+ Adversarial ( λ=0.1 )
0.7369
−0.0016
+ Oracle lexical
0.7409
+0.0025
+ Adapter + adversarial
0.7302
−0.0083
+ Adapter + adv. + lexical
0.7367
−0.0018
Table 1: Remedies for unseen-listener generalization (seed 13); every change is within the three-seed spread ( 0.0049 ).
Figure 5
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Meetings
Views
Listeners
Segments
BC events
Train
113
450
146
47,780
13,626
Validation
20
80
20
7,744
1,574
Test, seen
14
56
56
6,044
1,655
Test, unseen
24
96
24
10,020
1,842
Total
171
682
190
71,588
18,697
Appendix
Table 2: Masked-listener benchmark by split. Views are meeting-listener pairs; segments are 20-s windows. Listener counts do not sum to 190 because all 56 seen-test listeners are also training listeners. Backchannel events are attributed to their view; 19 of the 18,697 occur after the last complete 20-s segment and therefore appear in no segment.
Figure 8: Validation backchannel AUPRC per epoch for the fully fine-tuned trunk across three seeds and the rank-8 adapter for seed 13. Epoch 0 is the first completed pass; open markers indicate the checkpoints selected by early stopping.
λmax
Top-1 accuracy
× chance
0 (no adversary)
69.9%
102.1
0.1
68.6%
100.2
0.5
66.0%
96.4
1.0
55.1%
80.4
Appendix
Table 3: Top-1 listener-identity accuracy of a fresh probe on the frozen listener-stream representation for each adversarial strength (146 classes; chance 0.69% ).
Figure 9: Backchannel base rate as a function of the number of concurrently active other speakers in the gold test set.
Onset (horizon not compensated)
Window (compensated)
Configuration
Slice
P
R
F1
P
R
F1
None
seen
0.045
0.110
0.064±0.005
0.073
0.179
0.103±0.005
unseen
0.029
0.082
0.042±0.003
0.046
0.133
0.068±0.004
Listener
seen
0.057
0.103
0.070±0.008
0.091
0.163
0.112±0.012
unseen
0.029
0.083
0.041±0.005
0.047
0.131
0.066±0.002
L+FH (mixed)
seen
0.053
0.160
0.079±0.004
0.076
0.227
0.112±0.004
Appendix
Table 4: Event-level precision (P), recall (R), and F1 on the listener-active slice, using a ±300 ms collar and a validation-selected threshold. Values are means over three seeds, with s.d. reported for F1. L+FH denotes listener + floor-holder conditioning; the final configuration isolates the floor-holder’s audio in channel 2.
Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
Maike Züfle, Peter Polák, Sefik Emre Eskimez +3
Karlsruhe Institute of Technology, Germany · AppTek, Germany · Charles University, Czech Republic +2
Reliable turn-taking is essential for spoken dialogue systems. However, most existing methods are designed for two-speaker interaction and struggle with realistic multiparty audio containing overlap and rapid speaker changes. We study multiparty turn-taking on the VoxConverse dataset and propose an audio-only two-stage pipeline that separates when to trigger a turn boundary from whether the floor is actually transferring. A fast trigger scans the audio and proposes candidate end-of-turn times, while a lightweight verifier runs only at those times to decide \textsc{Hold} or \textsc{Shift} and support next-speaker prediction. We report results in the full multiparty setting and a controlled dyadic top-2 projection for comparability. We also investigate diffusion-based, label-preserving background-audio mixing as a data augmentation strategy. Results show improved shift detection over a baseline, with further improvements from diffusion augmentation.
Rutherford A. Patamia, Ming Liu, Wei Luo +2
Deakin University, Melbourne, Australia · Griffith University, Brisbane, Australia
We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, turn-change prediction, and next speaker prediction. We compare supervised models trained for these tasks, text-based LLMs, multimodal LLMs (MM-LLMs), and human subjects. Experiments on the AMI corpus showed that LLMs outperformed supervised models and humans in next speaker prediction, despite not being trained on the target domain and without access to audio or visual information. An MM-LLM performed better than text-based LLMs on addressee detection and turn-change prediction but remained below human performance, indicating difficulty leveraging raw audio-visual signals. Ablation analyses revealed that conversational context was critical, particularly for next speaker prediction. We observed that human and LLM prediction patterns were similar, and intervals with frequent turn changes were difficult for both.
Ryo Fukuda, Takatomo Kano, Siddhant Arora +7
NTT, Inc., Japan · Language Technologies Institute, Carnegie Mellon University, USA