Knowledge distillation (KD) improves low-resource acoustic learning by enriching one-hot supervision with the softened predictive distribution of a fixed teacher network. However, a teacher trained with limited or imbalanced annotations may produce a biased distribution whose components are not uniformly reliable. Although this distribution can still encode useful knowledge, direct full-distribution matching may also transfer teacher-induced biases, thereby distorting the student's decision boundary and degrading its generalization performance. To address this limitation, we propose Boundary-Anchored Mass-Partitioned Distillation (BA-MPD), a logit-based distillation objective composed of Boundary-Anchored Correction (BAC) and Mass-Partitioned Distillation (MPD). BAC addresses missing ground-truth labels in the set of the teacher's top predictions by swapping the true label for the lowest-ranked entry of the set, thus keeping the mass and uncertainty of the set unchanged. MPD then distills this corrected distribution through separate losses that enforce relational consistency within the set, balance the mass between high- and low-confidence groups, and weight lower-confidence dependencies. Ultimately, BAC and MPD together suppress harmful ranking errors and noisy low-confidence details, while retaining all useful teacher information. Experiments on two acoustic benchmarks under multiple label budgets show that BA-MPD consistently improves over supervised-learning baselines and vanilla KD while remaining competitive with strong logit-based KD baselines. Cross-budget results further show that BA-MPD remains effective when the teacher and student models use mismatched label budgets, demonstrating its ability to exploit imperfect teachers across supervision gaps. Implementation available at https://github.com/ShuanglinLi/BA-MPD.
Figures & tables
Fig. 1: Overview of Boundary-Anchored Mass-Partitioned Distillation (BA-MPD). The figure illustrates how BAC converts a teacher top- k support with label-support mismatch into a label-consistent anchored support Si through a boundary-level label exchange, while MPD distills the anchored teacher distribution by separating support relation matching, support–complement mass matching, and complement relation matching. The resulting BA-MPD loss is combined with cross-entropy (CE) to update the student.
Manner
Method
TAU Urban Acoustic Scenes: PaSST teacher → CP-Mobile student
5%
10%
25%
Macro
Seen
Unseen
Macro
Seen
Unseen
Macro
Seen
Unseen
Reference
Teacher Accuracy
47.82
43.78
45.19
49.74
45.55
47.22
50.53
45.90
48.18
Supervised
CE
42.40
39.44
38.00
45.49
43.12
40.66
50.09
48.21
45.58
Label smoothing
41.12
37.91
36.54
44.81
42.31
40.06
49.84
47.93
45.71
Mixup
43.13
39.95
39.21
45.92
43.16
41.77
50.83
49.11
45.78
TABLE I: Main results on TAU Urban Acoustic Scenes under different low-resource budgets. Reference Teacher denotes the split-specific PaSST teacher evaluated on the official test split and kept fixed for all KD methods within each budget. Student results are averaged over three random seeds and reported as macro-average scene classification accuracy, seen-device accuracy, and unseen-device accuracy (%). Δ denotes the absolute improvement of BA-MPD over vanilla KD.
Manner
Method
BEANS-CBI: CNN14 teacher → CNN14 student
10%
25%
50%
Macro
F1
Micro
Macro
F1
Micro
Macro
F1
Micro
Reference
Teacher Accuracy
6.80
6.00
7.32
19.33
17.47
21.33
37.80
35.30
40.77
Supervised
CE
6.66
5.88
7.29
19.80
18.09
21.73
38.71
36.18
41.40
Label smoothing
6.44
5.65
7.15
18.72
16.75
20.52
39.28
36.87
42.04
Focal loss
6.20
5.51
6.89
19.11
17.37
20.64
36.32
34.39
39.49
TABLE II: Main results on BEANS-CBI under different low-resource budgets. Reference Teacher denotes the split-specific CNN14 teacher evaluated on the official test split and kept fixed for all KD methods within each budget. Student results are averaged over three random seeds and reported as macro accuracy, macro-F1, and micro accuracy (%). Δ denotes the absolute improvement of BA-MPD over vanilla KD.
Benchmark
Budget
Setting 1
Setting 2
Setting 3
TAU
5%
45.43
46.33
45.73
10%
48.81
50.76
49.78
25%
54.05
54.73
54.45
BEANS-CBI
10%
7.39
7.44
7.32
25%
22.57
22.43
22.37
50%
41.91
41.83
41.67
TABLE III: Top- k sensitivity of BA-MPD. Results are reported using the primary metric of each benchmark: macro-average scene classification accuracy for TAU Urban Acoustic Scenes and macro accuracy for BEANS-CBI (%). For TAU, the three settings correspond to k=1/2/3 , with default k=2 . For BEANS-CBI, they correspond to k=5/10/20 , with default k=5 . Best results are highlighted in bold, and default settings are underlined.
Variant
Loss Terms
TAU
BEANS-CBI
Lmass
Ltop
Lcomp
5%
10%
25%
10%
25%
50%
Full BA-MPD
✓
✓
✓
46.33
50.76
54.73
7.39
22.57
41.91
w/o Lmass
–
✓
✓
45.53
48.97
53.70
7.45
21.84
40.86
w/o Ltop
✓
–
✓
45.08
48.39
53.46
7.50
23.51
41.83
w/o Lcomp
✓
✓
–
44.89
48.96
54.18
6.62
20.28
38.78
TABLE IV: Loss-term ablation of BA-MPD under the default top- k setting. All variants use the same boundary-anchored support and training protocol, differing only in the distillation terms included. Results are reported using the primary metric of each benchmark: macro-average scene accuracy for TAU Urban Acoustic Scenes and macro accuracy for BEANS-CBI (%).
Teacher Budget rt
Student Budget rs
Teacher Acc.
CE Student
Vanilla KD
MPD
BA-MPD
ΔCE
ΔVKD
5%
5%
47.82
42.40
45.70
45.55
46.33
+3.93
+0.63
10%
45.49
48.68
48.96
49.61
+4.12
+0.93
25%
50.09
53.04
53.42
53.80
+3.71
+0.76
10%
10%
49.74
45.49
49.36
49.89
50.76
+5.27
+1.40
25%
50.09
54.44
54.58
54.72
+4.63
+0.28
25%
25%
50.53
50.09
53.48
54.58
54.73
+4.64
+1.25
TABLE V: Cross-budget transfer on TAU Urban Acoustic Scenes with rt≤rs . rt and rs denote the teacher and student label budgets, respectively, and only pairs satisfying rt≤rs are evaluated. Results are reported as macro-average scene classification accuracy (%). ΔCE denotes the improvement of BA-MPD over the CE student with the same student budget, and ΔVKD denotes the improvement over vanilla KD under the same teacher–student budget pair.
Fig. 2: Teacher top- k support miss rates across support sizes and label budgets. Each cell reports the percentage of samples whose ground-truth class is excluded from the teacher’s top- k support. Lower values indicate fewer label-support mismatches. Outlined columns mark the default k used by BA-MPD.
Fig. 3: Incremental effect of boundary anchoring across different teacher-support conditions. The x-axis shows the teacher top- k support miss rate at the default k , and the y-axis reports the primary-metric difference between BA-MPD and MPD. The dashed line marks no performance change, with points above and below the line indicating positive and negative effects of BAC, respectively. Bubble area is proportional to the mean teacher probability mass assigned outside the default top- k support.