Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by choosing a small training subset under a fixed audio-duration budget. We represent a corpus as a phonotactic graph that links each utterance to its most phonemically similar ones, and we first test whether this graph has structure. In Bangla and English corpora, its clustering is 199 and 56 times that of a size-matched random graph, and its modularity is more than twice that of a degree-preserving random graph. We then propose Community Representative, a selector that samples across graph communities and spreads its choices within each one, starting from utterances rich in rare phonemes. At every budget and in both languages, it covers more rare phoneme bigrams than random and entropy-based selection, and this lead holds on held-out utterances. TTS models trained on its 20% core-sets have a significantly lower character error rate (CER) than models trained on equal-duration random or entropy-based subsets in both languages. When all models train for the same number of epochs, the Bangla core-set model also outperforms full-corpus training (3.93% vs. 4.47% CER) with 4.5x less training time.
Figures & tables
Figure 1: Overview. (a) Utterances become TF-IDF vectors over IPA phonemes and within-word bigrams (§ 3.3 ). (b) The k -NN utterance graph is tested against Erdős–Rényi (ER) and configuration-model nulls before selection relies on it (§ 4 ); colours mark Louvain communities. (c) Community Representative splits the duration budget across communities and selects within each in farthest-point order from the most rare-phoneme-rich utterance (star; § 5 ). (d) Core-sets are scored by rare-bigram coverage and downstream TTS intelligibility (§ 6 – 7 ).
quantity
Bangla
English
utterances / hours
40,422 / 48.75
13,100 / 23.92
pool / held-out
36,389 / 4,033
11,769 / 1,331
IPA phonemes
37
41
within-word bigrams
775
1,073
rare bigrams (bottom qu.)
194
278
rare share of tokens (%)
0.099
0.351
Table 1: Corpus and graph descriptors. Every hyperparameter behind these numbers is shared across the two languages.
Bangla
English
Gu
ER
config.
Gu
ER
config.
N
36,388
matched
matched
11,769
matched
matched
L
462,066
matched
matched
150,416
matched
matched
⟨k⟩
25.4
25.4
25.4
25.6
25.6
25.5
C
0.1401
0.0007
0.0018
0.1211
0.0022
0.0056
Q
0.470
0.156
0.172
0.404
0.155
0.177
Table 2: Gu against matched nulls in both languages (ER = Erdős–Rényi, config. = degree-preserving configuration model), averaged over 5 realizations. Clustering C and modularity Q exceed both nulls in both languages.
Figure 2: Gu versus matched nulls (log scale), one panel per language. Clustering and modularity both exceed the degree-preserving configuration model in Bangla and in English.
Figure 3: Degree distribution and clustering spectrum of Gu against one realization (seed 13) of each null model. P(K≥k) is the complementary cumulative degree distribution; C(k) is the mean local clustering coefficient of utterances in logarithmic degree bins. The configuration model shares Gu ’s degree sequence, so it appears only in the clustering panels.
Figure 4: Communities of Gu and the utterances Community Representative selects at a 10% budget. Communities are the Louvain partition the selector uses on the full similarity-weighted graph. Because an induced subgraph of Gu on a sample has almost no edges, each panel lays out a k -NN graph ( k=8 ) rebuilt on a 2,500-utterance sample stratified by community; the three largest communities are coloured and the rest grey. Rings mark selected utterances.
Community Representative
Phoneme Balance
Random
Full data
(20%)
(20%)
(20%)
(100%)
Bangla utt. / hours
9,183 / 8.78
8,872 / 8.78
7,299 / 8.78
40,422 / 48.75
English utt. / hours
3,017 / 4.30
2,825 / 4.30
2,342 / 4.30
11,769 / 21.48
Bangla schedule
1500 epochs; batch 128 (Random: 98)
English schedule
30,000 gradient steps; batch 256
Table 3: Downstream training arms: utterances / hours per language, and training schedules. The three 20% subsets have identical duration.
Bangla
English
method
10%
20%
40%
10%
20%
40%
Community Representative (ours)
89.5 ± 0.6
96.0 ± 0.8
98.5 ± 0.5
73.8 ± 1.2
86.0 ± 1.1
95.7 ± 0.8
Phoneme Balance
73.2
82.5
93.8
55.0
71.9
86.0
Random
49.9 ± 4.0
69.8 ± 1.8
83.7 ± 2.1
46.3 ± 2.0
64.8 ± 3.7
80.6 ± 3.9
Table 4: Rare-bigram coverage (%) by method, budget and language, mean ± 95% CI over 5 seeds. Entries without an interval have zero seed variance (§ 6.1 ). Best per column in bold.
Figure 5: Rare-bigram coverage vs. duration budget in both languages; bands are 95% CIs over 5 seeds. The ranking of methods is the same in both panels.
Figure 6: Graph coverage against the duration budget: the share of pool utterances that are selected or among the 15 nearest neighbours of a selected utterance. Community Representative is re-run at every 2% budget; bands are 95% CIs over 5 seeds.
Bangla
English
method
10%
20%
10%
20%
Community Representative (ours)
99.1 ± 0.2
99.4 ± 0.1
98.0 ± 0.3
98.8 ± 0.2
Phoneme Balance
97.3
98.5
94.8
97.7
Random
94.4 ± 0.9
97.3 ± 0.5
94.1 ± 0.7
97.1 ± 0.3
Table 5: Held-out bigram coverage (%): the share of the bigram types occurring in held-out utterances, which no selector could see, that each selection covers.
method
10%
20%
Community Representative (ours)
44.7
66.0
Phoneme Balance
37.3
54.7
Random
22.7
43.3
Table 6: Rare-bigram coverage (%) on OpenSLR SLR37 (1,891 utterances, 6 speakers). Community Representative leads at both budgets, and the ordering of the three methods matches both single-speaker corpora.
Bangla
English
System
CER (%)
Median
# >0.3
CER (%)
Median
# >0.3
Reference (topline)
3.42 ± 0.44
2.63
0
1.85 ± 0.21
0.83
1
Random (20%)
4.90 ± 0.52
3.61
6
3.39 ± 0.38
1.57
1
Phoneme Balance (20%)
4.62 ± 0.63
3.28
6
3.30 ± 0.46
1.35
1
Community Representative (20%)
3.93 ± 0.51
3.06
0
2.96 ± 0.33
1.24
1
Full data (100%)
4.47 ± 0.40
3.19
7
2.84 ± 0.41
1.11
3
Table 7: Held-out intelligibility of the TTS models trained on each arm: character error rate (CER, %) as mean ± 95% CI, median CER, and the number of catastrophic failures (items with CER >0.3 ). Bangla arms are matched on epochs and scored on 496 of the 500 test items; English arms are matched on gradient steps (30,000 each) and scored on all 500 items. Bold marks the best trained system per language.