Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.
Figures & tables
Figure 1: Overview of the Scalable Multi-Label Classification Methodology.
Model Architecture
Parameters
Reviews (Macro-F1)
Conversations (Macro-F1)
P95 Latency (ms)
Discriminative (DeBERTa-V3)
186M
0.7047
0.4062
< 100
Generative (SLM-1B)
1B
0.6502
0.4501
163
Generative (SLM-4B)
4B
0.7405
0.4908
208
Generative (SLM-8B)
8B
0.8031
0.5404
241
Table 1: Master Performance and Latency Comparison Across Architectures
Training Support
DeBERTa-V3
ModernBERT
SLM-8B
SLM-4B
SLM-1B
(Macro)
(Macro)
(Macro)
(Macro)
(Macro)
≥1
0.350
0.304
0.473
0.457
0.436
≥10
0.350
0.304
0.475
0.461
0.440
≥50
0.357
0.309
0.500
0.482
0.461
≥100
0.376
0.330
0.527
0.513
0.489
≥200
0.407
0.354
0.572
0.556
0.528
Table 2: Macro-F1 performance metrics on Conversational dataset across different training support levels.
Training Support
DeBERTa-V3
ModernBERT
SLM-8B
SLM-4B
SLM-1B
(Macro)
(Macro)
(Macro)
(Macro)
(Macro)
≥1
0.670
0.687
0.796
0.653
0.646
≥200
0.670
0.687
0.796
0.653
0.646
≥250
0.703
0.690
0.798
0.652
0.646
≥500
0.742
0.718
0.811
0.642
0.643
≥1000
0.749
0.731
0.814
0.633
0.635
Table 3: Macro-F1 performance metrics on Product Review dataset across different training support levels.
Dataset
Expansion
DeBERTa-V3
SLM-8B
Conversations
10 → 112 Labels
-35.53%
-0.23%
Product Reviews
4 → 21 Labels
-15.01%
-4.46%
Table 4: F1 Macro performance change with label expansion.
Label Configuration
DeBERTa-V3
SLM-8B
Macro-F1
Macro-F1
10 Random Labels
0.5435
0.4739
10+15 Labels
0.5132
0.5553
25+25 Random Labels
0.4890
0.5824
All Random Labels (112)
0.3504
0.4728
Table 5: Macro-F1 performance with varying label configurations on Conversational dataset.
Category Configuration
DeBERTa-V3
SLM-8B
Macro-F1
Macro-F1
Tires Only
0.7879
0.8334
Tires and P&A
0.6987
0.7924
All 5 Categories
0.6696
0.7962
Table 6: Macro-F1 performance with varying category combinations on Product Review dataset.
Architecture
Model
Latency (ms)
Relative to DeBERTa-V3
Discriminative
DeBERTa-V3
58
1.0 ×
ModernBERT
79
1.4 ×
Generative
SLM-1B
163
2.8 ×
SLM-4B
208
3.6 ×
SLM-8B
241
4.2 ×
Table 7: Inference latency comparison between discriminative and generative model families
Model
# OOV
OOV label rate
OOV samples
Sample rate
Most frequent OOV labels
SLM-1B
22
19.6%
42 / 39,172
0.11%
shipping-issue , shipping-insurance
SLM-4B
17
15.2%
54 / 39,172
0.14%
shipping-complaint , shipping-issue
SLM-8B
29
25.9%
100 / 39,172
0.26%
shipping-issue , shipping-label
Table 8: Out-of-vocabulary (OOV) labels produced by generative models. The OOV label rate is computed against the 112 labels observed during training, while the sample rate measures the fraction of evaluation samples for which the model generated an OOV label.
Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks incurs prohibitive latency, cost, and data-privacy overhead. We present a hybrid framework that fine-tunes a small language model (LLaMA 3.1 8B, 2.05% trainable parameters via LoRA) on only 219 curated examples and couples it with a deterministic rule-based postprocessing layer. Applied to multi-label compliance evaluation of conversational transcripts (18 heterogeneous output fields), our system achieves 100% JSON structural validity, 83.0% human-validated overall accuracy, and 100% accuracy on the most critical classification field in blind evaluation on 53 unseen production transcripts. On a single NVIDIA A100 GPU, inference completes in ∼2 seconds -- 2--5x faster than frontier APIs -- at USD 0.013 per evaluation versus USD 0.025--0.055 for proprietary alternatives, yielding 46--76% cost savings. We introduce targeted hard-negative augmentation for critical decision boundaries and formalize the hybrid neural-symbolic decomposition, demonstrating that domain-adapted small language models with postprocessing can match frontier model accuracy while dramatically reducing operational cost, latency, and privacy risk.
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-K candidate labels by embedding similarity and prompt the LLM to choose among them. However, top-K retrieval reduces the number of candidates but does not help the model tell similar ones apart. When two similar labels both appear as candidates, the model lacks the signal to choose correctly between them. We propose a framework that (1) identifies which label pairs the model struggles to distinguish, (2) expands the candidate set to include confusable labels, and (3) generates targeted rules to differentiate between similar candidates. The framework requires no fine-tuning, and the generated rules transfer to smaller, cheaper models. On three benchmarks (WOS, Flipkart, LEDGAR), our approach improves Macro F1 by up to 10.0pp over retrieval baselines, with smaller models (2B--20B) gaining up to 11.5pp via cross-model transfer.
A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.
Xiangwei Wang, Peng Wang, Saman Halgamuge
The University of Melbourne, Australia · Shanghai Jiao Tong University, China