Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.
Figures & tables
Figure 1: Overview of the Scalable Multi-Label Classification Methodology.
Model Architecture
Parameters
Reviews (Macro-F1)
Conversations (Macro-F1)
P95 Latency (ms)
Discriminative (DeBERTa-V3)
186M
0.7047
0.4062
< 100
Generative (SLM-1B)
1B
0.6502
0.4501
163
Generative (SLM-4B)
4B
0.7405
0.4908
208
Generative (SLM-8B)
8B
0.8031
0.5404
241
Table 1: Master Performance and Latency Comparison Across Architectures
Training Support
DeBERTa-V3
ModernBERT
SLM-8B
SLM-4B
SLM-1B
(Macro)
(Macro)
(Macro)
(Macro)
(Macro)
≥1
0.350
0.304
0.473
0.457
0.436
≥10
0.350
0.304
0.475
0.461
0.440
≥50
0.357
0.309
0.500
0.482
0.461
≥100
0.376
0.330
0.527
0.513
0.489
≥200
0.407
0.354
0.572
0.556
0.528
Table 2: Macro-F1 performance metrics on Conversational dataset across different training support levels.
Training Support
DeBERTa-V3
ModernBERT
SLM-8B
SLM-4B
SLM-1B
(Macro)
(Macro)
(Macro)
(Macro)
(Macro)
≥1
0.670
0.687
0.796
0.653
0.646
≥200
0.670
0.687
0.796
0.653
0.646
≥250
0.703
0.690
0.798
0.652
0.646
≥500
0.742
0.718
0.811
0.642
0.643
≥1000
0.749
0.731
0.814
0.633
0.635
Table 3: Macro-F1 performance metrics on Product Review dataset across different training support levels.
Dataset
Expansion
DeBERTa-V3
SLM-8B
Conversations
10 → 112 Labels
-35.53%
-0.23%
Product Reviews
4 → 21 Labels
-15.01%
-4.46%
Table 4: F1 Macro performance change with label expansion.
Label Configuration
DeBERTa-V3
SLM-8B
Macro-F1
Macro-F1
10 Random Labels
0.5435
0.4739
10+15 Labels
0.5132
0.5553
25+25 Random Labels
0.4890
0.5824
All Random Labels (112)
0.3504
0.4728
Table 5: Macro-F1 performance with varying label configurations on Conversational dataset.
Category Configuration
DeBERTa-V3
SLM-8B
Macro-F1
Macro-F1
Tires Only
0.7879
0.8334
Tires and P&A
0.6987
0.7924
All 5 Categories
0.6696
0.7962
Table 6: Macro-F1 performance with varying category combinations on Product Review dataset.
Architecture
Model
Latency (ms)
Relative to DeBERTa-V3
Discriminative
DeBERTa-V3
58
1.0 ×
ModernBERT
79
1.4 ×
Generative
SLM-1B
163
2.8 ×
SLM-4B
208
3.6 ×
SLM-8B
241
4.2 ×
Table 7: Inference latency comparison between discriminative and generative model families
Model
# OOV
OOV label rate
OOV samples
Sample rate
Most frequent OOV labels
SLM-1B
22
19.6%
42 / 39,172
0.11%
shipping-issue , shipping-insurance
SLM-4B
17
15.2%
54 / 39,172
0.14%
shipping-complaint , shipping-issue
SLM-8B
29
25.9%
100 / 39,172
0.26%
shipping-issue , shipping-label
Table 8: Out-of-vocabulary (OOV) labels produced by generative models. The OOV label rate is computed against the 112 labels observed during training, while the sample rate measures the fraction of evaluation samples for which the model generated an OOV label.