Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.
Figures & tables
Figure 1: VAP-BC with the optional components evaluated in Section 5.3 : listener/floor-holder FiLM conditioning, rank-8 adapters, listener-identity adversary, and lexical stream.
Figure 2Figure 3
Method (listener + floor-holder base)
Unseen AUROC
Δ
Baseline
0.7385
—
+ Adapter ( r=8 )
0.7307
−0.0078
+ Adversarial ( λ=0.1 )
0.7369
−0.0016
+ Oracle lexical
0.7409
+0.0025
+ Adapter + adversarial
0.7302
−0.0083
+ Adapter + adv. + lexical
0.7367
−0.0018
Table 1: Remedies for unseen-listener generalization (seed 13); every change is within the three-seed spread ( 0.0049 ).
Figure 5
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Meetings
Views
Listeners
Segments
BC events
Train
113
450
146
47,780
13,626
Validation
20
80
20
7,744
1,574
Test, seen
14
56
56
6,044
1,655
Test, unseen
24
96
24
10,020
1,842
Total
171
682
190
71,588
18,697
Appendix
Table 2: Masked-listener benchmark by split. Views are meeting-listener pairs; segments are 20-s windows. Listener counts do not sum to 190 because all 56 seen-test listeners are also training listeners. Backchannel events are attributed to their view; 19 of the 18,697 occur after the last complete 20-s segment and therefore appear in no segment.
Figure 8: Validation backchannel AUPRC per epoch for the fully fine-tuned trunk across three seeds and the rank-8 adapter for seed 13. Epoch 0 is the first completed pass; open markers indicate the checkpoints selected by early stopping.
λmax
Top-1 accuracy
× chance
0 (no adversary)
69.9%
102.1
0.1
68.6%
100.2
0.5
66.0%
96.4
1.0
55.1%
80.4
Appendix
Table 3: Top-1 listener-identity accuracy of a fresh probe on the frozen listener-stream representation for each adversarial strength (146 classes; chance 0.69% ).
Figure 9: Backchannel base rate as a function of the number of concurrently active other speakers in the gold test set.
Onset (horizon not compensated)
Window (compensated)
Configuration
Slice
P
R
F1
P
R
F1
None
seen
0.045
0.110
0.064±0.005
0.073
0.179
0.103±0.005
unseen
0.029
0.082
0.042±0.003
0.046
0.133
0.068±0.004
Listener
seen
0.057
0.103
0.070±0.008
0.091
0.163
0.112±0.012
unseen
0.029
0.083
0.041±0.005
0.047
0.131
0.066±0.002
L+FH (mixed)
seen
0.053
0.160
0.079±0.004
0.076
0.227
0.112±0.004
Appendix
Table 4: Event-level precision (P), recall (R), and F1 on the listener-active slice, using a ±300 ms collar and a validation-selected threshold. Values are means over three seeds, with s.d. reported for F1. L+FH denotes listener + floor-holder conditioning; the final configuration isolates the floor-holder’s audio in channel 2.