Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental question remains unanswered: which pooling strategy maximizes intent classification accuracy under realistic conversational noise in production language models? This work addresses this gap through 360 controlled experiments spanning four pooling configurations (mean, max, last-token, attention, and FFT-augmented variants) using Llama-3.2-1B-Instruct on BANKING77 and CLINC150 datasets under clean/noisy conditions with ten random seeds. Key findings reveal that attention pooling consistently outperforms alternative strategies under noisy conditions (~+2.6-2.8 F1 over the default), while mean pooling degrades performance by up to ~5 F1 points. Frequency-domain filtering does not produce consistent accuracy improvements and functions primarily as a structural variation rather than an accuracy-enhancing component. These results provide concrete, evidence-based guidance for building noise-robust conversational classifiers: attention pooling is recommended for noisy interfaces, mean pooling should be avoided, and last-token pooling is appropriate for clean-query scenarios.
Figures & tables
Figure 2: Overview of the nine experimental configurations. The shared backbone (input, tokenizer, Llama, H ) feeds into three groups: C0 (baseline with default head), C1 – C4 (custom pooling head applied directly to H ), and C5 – C8 (FFT low-pass filtering applied before pooling, H~=IFFT(FFT(H)⊙M40%) ).
Configuration
Head
Pooling
FFT
C0
Baseline
Default
Last
✗
C1
Mean
Custom
Mean
✗
C2
Max
Custom
Max
✗
C3
Last-token
Custom
Last
✗
C4
Attention
Custom
Attention
✗
C5
FFT + Mean
Custom
Mean
✓
Table 1: Complete experimental design matrix. C0 is the baseline using the default Llama classification head; C1 – C4 use a custom pooling head without filtering; C5 – C8 add FFT low-pass filtering before pooling.
Clean BANKING77
Noisy BANKING77
Clean CLINC150
Noisy CLINC150
F1 (%)
Train FLOPs
F1 (%)
Train FLOPs
F1 (%)
Train FLOPs
F1 (%)
Train FLOPs
default
88.80 ± 0.98
2.95e17 ± 2.57e16
87.67 ± 0.64
2.91e17 ± 2.58e16
86.66 ± 1.15
3.60e17 ± 2.09e16
86.42 ± 1.56
3.73e17 ± 1.76e16
custom mean
88.65 ± 0.54
3.97e17 ± 3.26e16
86.43 ± 0.53
4.09e17 ± 1.91e16
83.11 ± 1.25
4.97e17 ± 2.14e16
83.66 ± 1.59
4.94e17 ± 4.84e16
custom max
89.63 ± 0.39
3.96e17 ± 2.10e16
88.89 ± 0.37
3.83e17 ± 1.71e16
88.26 ± 1.16
4.72e17 ± 2.37e16
88.88 ± 0.98
4.94e17 ± 3.93e16
custom last
90.79 ± 0.37
2.73e17 ± 1.97e16
89.97 ± 0.50
2.80e17 ± 1.96e16
89.56 ± 0.58
3.26e17 ± 1.79e16
88.78 ± 0.71
3.45e17 ± 1.02e16
custom attention
90.74 ± 0.43
2.85e17 ± 2.28e16
90.31 ± 0.36
2.98e17 ± 2.02e16
88.81 ± 0.98
3.18e17 ± 1.81e16
89.25 ± 1.12
3.22e17 ± 2.11e16
Table 2: Performance metrics (F1 scores and training FLOPs) across experiment configurations for each dataset.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
LoRA configuration
Rank ( r )
4
Scaling factor ( α )
16
Target modules
q_proj , v_proj
LoRA dropout
0.1
Bias
None
Appendix
Table B.1: Training hyperparameters used in all experiments. The setup emphasizes parameter efficiency, early convergence, and reduced computational overhead.
Figure D.1: Mean ± standard deviation F1 scores across all experiments on the BANKING77 dataset.
Figure D.2: Mean ± standard deviation F1 scores across all experiments on the CLINC150 dataset.
Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefined classes. For instance, given a user prompt, the system must determine whether it primarily concerns mathematics, coding, or general text processing. Such classification enables routing prompts to specialized models optimized for specific domains, improving both accuracy and computational efficiency. In this work, we conduct a systematic study comparing training-free vs training-based approaches for intent classification. For this purpose, we consider two lightweight, training-free methods based on statistics of internal representations and compare them against MLP classifiers and linear probes. Our comprehensive empirical evaluation reveals that 1) Both training-free and training-based methods saturate easy benchmarks (mathematics vs. coding vs. natural language), 2) Training-based classifiers have an advantage on harder classification tasks (e.g. Java vs Python), and 3) Training-free methods are generally more robust to mixed-intent and adversarial prompts.
Nan Chen, Zhouhao Yang, Soufiane Hayou
Department of Applied Mathematics and Statistics Johns Hopkins University
We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset of 1,724 difficult safety prompts, each paired with an intent description and harm label. We use AIMS to evaluate intent-aware training across supervised fine-tuning, preference learning, reasoning distillation, and reinforcement learning. Despite its size, AIMS enables competitive safety classifiers across training regimes: DPO from model-generated intent errors improves over SFT, and intent-conditioned distillation outperforms reasoning-only distillation in most teacher-student pairs. Most notably, directly rewarding intent faithfulness with GRPO yields the strongest average performance across five external safety benchmarks, while our intent-aware models form the inference latency-F1 Pareto frontier. These results show that faithful intent modeling is a compact, high-quality supervision signal for more robust safety classifiers.
Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints. We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M--9B parameter range across eight English single-label intent-classification datasets. A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result. The evaluation includes standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy, we analyze confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Our results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models. Instruction tuning's effect on confidence calibration is inconsistent rather than uniformly harmful. These findings provide practical guidance for selecting and evaluating open-weight language models for intent classification.
Parishruthi Ganesh, Gerry Dozier, Cheryl Seals
Department of Computer Science and Software Engineering Auburn University Auburn, Alabama, USA