This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining speaker diarization with an LLM pass, with every label then checked by a human annotator. The model pairs a Whisper encoder with task-specific classification heads. On a class-balanced test set drawn from a held-out podcast, it reaches 84.33% accuracy (95% CI 80.3 to 88.1) against 69.28% for the Smart-Turn v3 baseline, and lowers the false negative rate from 51.57% to 7.55% at the cost of a higher false positive rate. We report what encoder layer fine-tuning, multi-scale pooling and INT8 quantization each contribute, and latency stays within 165 to 191 ms end to end on CPU.
Figures & tables
Figure 1: Our proposed model architecture with Whisper-tiny encoder finetuned in Bangla, multi-scale pooling and a classifier.
Metric
Training
Validation
Test
Total samples
31,549
3,506
319
Endpoint
25,276
2,829
159
Non-endpoint
6,273
677
160
Table 1: BanglaTurn dataset statistics. Endpoint and non-endpoint partition each split.
Model
Acc
FPR
FNR
Prec
Rec
F1
Silence threshold †
49.22 [43.9, 54.9]
11.88
89.94
45.71
10.06
16.49 [9.6, 23.5]
Prosody + silence LR †
62.07 [56.7, 67.4]
50.00
25.79
59.60
74.21
66.11 [60.2, 71.4]
Smart-Turn v3
69.28 [63.9, 74.3]
10.00
51.57
82.80
48.43
61.11 [53.7, 67.7]
Ours
84.33 [80.3, 88.1]
23.75
7.55
79.46
92.45
85.47 [81.2, 89.3]
Table 2: Performance comparison on the BanglaTurn test set (all metrics in %, endpoint as the positive class). Brackets give 95% bootstrap confidence intervals. FPR is computed over the 160 non-endpoint clips and FNR over the 159 endpoint clips. † Fitted on the test podcast itself by 10-fold cross-validation, and therefore optimistic.
Encoder + Components
Unfrozen
Acc
FPR
FNR
Prec
Rec
F1
layers
(%)
(%)
(%)
(%)
(%)
(%)
Whisper-tiny
0
83.07
21.88
11.95
80.00
88.05
83.83
Whisper-tiny
2
81.19
29.38
8.18
75.65
91.82
82.95
Whisper-tiny
4
83.07
27.50
6.29
77.20
93.71
84.66
Whisper-tiny (Bangla)
0
84.01
20.62
11.32
81.03
88.68
84.68
Whisper-tiny (Bangla)
2
81.19
27.50
10.06
76.47
89.94
82.66
Table 3: Ablation study comparing different model variants on the BanglaTurn test set. MSP denotes Multi-Scale Pooling. FPR and FNR are computed over the non-endpoint and endpoint clips respectively. With 319 test clips, the 95% confidence interval on each accuracy spans about ± 4 points, wider than most gaps between rows.
Predicted
Model
Reference
Endpoint
Non-endpoint
Silence threshold †
Endpoint
16
143
Non-endpoint
19
141
Prosody + silence LR †
Endpoint
118
41
Non-endpoint
80
80
Smart-Turn v3
Endpoint
77
82
Table 4: Confusion matrices on the BanglaTurn test set (159 endpoint and 160 non-endpoint clips). Rows give the reference label, columns the prediction. † Fitted on the test podcast by 10-fold cross-validation.
Encoder + Components
Quant.
Acc
Size
Lat.
(%)
(MB)
(ms)
Whisper-tiny
FP32
83.07
148.2
182.92
Whisper-tiny
INT8
82.76
39.4
164.98
Whisper-tiny (Bangla)
FP32
84.01
148.2
180.04
Whisper-tiny (Bangla)
INT8
79.31
39.4
169.96
Whisper-tiny (BN) + MSP
FP32
84.33
149.1
191.00
Table 5: Impact of INT8 quantization on model size, inference latency, and accuracy.
In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text representations overlap substantially. These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.
Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com.
Endpoint detection (EPD) is essential for natural turn-taking in streaming speech systems. However, reliably determining the endpoint of an utterance is challenging because speakers often pause mid-utterance due to hesitations and disfluencies. Semantic EPD has emerged as a promising direction to address this issue but is hindered by ambiguous supervision and strict streaming constraints. We propose Next-Turn that uses the time-to-next-speech-onset as the training objective, where targets are derived directly from speech timestamps and require no additional annotation. Experiments show that the proposed method outperforms conventional acoustic and recent semantic EPD baselines, achieving a 25.9% absolute improvement in endpoint accuracy within 320 ms over the strongest baseline. In addition, joint training with the duration-aware objective complements standard binary EPD, with gains that increase monotonically with increasing pauses.
Tristan Tsoi, Jiajun Deng, Yingke Zhu +5
Central Media Technology Institute, Huawei · The Chinese University of Hong Kong · Nanyang Technological University