This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining speaker diarization with an LLM pass, with every label then checked by a human annotator. The model pairs a Whisper encoder with task-specific classification heads. On a class-balanced test set drawn from a held-out podcast, it reaches 84.33% accuracy (95% CI 80.3 to 88.1) against 69.28% for the Smart-Turn v3 baseline, and lowers the false negative rate from 51.57% to 7.55% at the cost of a higher false positive rate. We report what encoder layer fine-tuning, multi-scale pooling and INT8 quantization each contribute, and latency stays within 165 to 191 ms end to end on CPU.
Figures & tables
Figure 1: Our proposed model architecture with Whisper-tiny encoder finetuned in Bangla, multi-scale pooling and a classifier.
Metric
Training
Validation
Test
Total samples
31,549
3,506
319
Endpoint
25,276
2,829
159
Non-endpoint
6,273
677
160
Table 1: BanglaTurn dataset statistics. Endpoint and non-endpoint partition each split.
Model
Acc
FPR
FNR
Prec
Rec
F1
Silence threshold †
49.22 [43.9, 54.9]
11.88
89.94
45.71
10.06
16.49 [9.6, 23.5]
Prosody + silence LR †
62.07 [56.7, 67.4]
50.00
25.79
59.60
74.21
66.11 [60.2, 71.4]
Smart-Turn v3
69.28 [63.9, 74.3]
10.00
51.57
82.80
48.43
61.11 [53.7, 67.7]
Ours
84.33 [80.3, 88.1]
23.75
7.55
79.46
92.45
85.47 [81.2, 89.3]
Table 2: Performance comparison on the BanglaTurn test set (all metrics in %, endpoint as the positive class). Brackets give 95% bootstrap confidence intervals. FPR is computed over the 160 non-endpoint clips and FNR over the 159 endpoint clips. † Fitted on the test podcast itself by 10-fold cross-validation, and therefore optimistic.
Encoder + Components
Unfrozen
Acc
FPR
FNR
Prec
Rec
F1
layers
(%)
(%)
(%)
(%)
(%)
(%)
Whisper-tiny
0
83.07
21.88
11.95
80.00
88.05
83.83
Whisper-tiny
2
81.19
29.38
8.18
75.65
91.82
82.95
Whisper-tiny
4
83.07
27.50
6.29
77.20
93.71
84.66
Whisper-tiny (Bangla)
0
84.01
20.62
11.32
81.03
88.68
84.68
Whisper-tiny (Bangla)
2
81.19
27.50
10.06
76.47
89.94
82.66
Table 3: Ablation study comparing different model variants on the BanglaTurn test set. MSP denotes Multi-Scale Pooling. FPR and FNR are computed over the non-endpoint and endpoint clips respectively. With 319 test clips, the 95% confidence interval on each accuracy spans about ± 4 points, wider than most gaps between rows.
Predicted
Model
Reference
Endpoint
Non-endpoint
Silence threshold †
Endpoint
16
143
Non-endpoint
19
141
Prosody + silence LR †
Endpoint
118
41
Non-endpoint
80
80
Smart-Turn v3
Endpoint
77
82
Table 4: Confusion matrices on the BanglaTurn test set (159 endpoint and 160 non-endpoint clips). Rows give the reference label, columns the prediction. † Fitted on the test podcast by 10-fold cross-validation.
Encoder + Components
Quant.
Acc
Size
Lat.
(%)
(MB)
(ms)
Whisper-tiny
FP32
83.07
148.2
182.92
Whisper-tiny
INT8
82.76
39.4
164.98
Whisper-tiny (Bangla)
FP32
84.01
148.2
180.04
Whisper-tiny (Bangla)
INT8
79.31
39.4
169.96
Whisper-tiny (BN) + MSP
FP32
84.33
149.1
191.00
Table 5: Impact of INT8 quantization on model size, inference latency, and accuracy.