Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.
Figures & tables
Figure 1: Overview of the main tasks of our study
scenario
π\textsubscript21
π\textsubscript12
odds ratio
sample size
Conservative ( 5% increase)
0.10
0.15
1.5
796
Moderate ( 10% increase)
0.10
0.20
2.0
240
Optimistic ( 20% increase)
0.05
0.25
1.5
57
Table 1: A priori analysis for McNemar and Cochran’s Q tests with 3 scenarios. π\textsubscript21 is the proportion that worsened. π\textsubscript12 is the proportion that improved
scenario
effect size
dof
sample size
Small dof
Conservative (small effect)
0.10
1
785
Moderate (medium effect)
0.30
1
88
Optimistic (large effect)
0.50
1
32
Large dof
Conservative (small effect)
0.10
46
2919
Table 2: A priori analysis for goodness-of-fit tests with 3 scenarios, and two degree of freedom bounds for each scenario
Rank
Model
# Params (B)
Score
HumanEval
19
Qwen2.5-Coder 7B Instruct
7
0.884
35
Qwen2.5 7B Instruct
7.61
0.848
40
IBM Granite 4.0 Tiny Preview
7
0.824
44
Qwen2 7B Instruct
7.62
0.799
45
Qwen2.5-Omni-7B
7
0.787
Table 3: Performance of Small Language Models on the HumanEval, BigCodeBench and EvalPlus benchmarks. The selected SLMs for comparison appear in grey. Each SLM can appear in multiple benchmarks.
Metric
κ
Overall Agreement
Fleiss’ κ (3 experts)
0.89
Pairwise Agreement
Cohen’s κ ( E 1 , E 2 )
0.84
Cohen’s κ ( E 1 , E 3 )
1.00
Cohen’s κ ( E 2 , E 3 )
0.84
Table 4: Inter-annotator agreement measures across the 3 experts.
Figure 2: F 1 score vs perplexity for the selected SLMs. The dashed lines represent the median for the F 1 score (horizontal) and for the perplexity (vertical).
SLM name
accuracy
F 1 score
mcc
perplexity
error rate (%)
Qwen2.5-Coder-7B-Instruct
0.586
0.754
0.494
1.172
0.0095
IBM Granite 4.0
0.326
0.624
0.306
1.223
0.152
Gemma 3 4B
0.330
0.641
0.365
1.035
0
Phi-3-mini-128k-instruct
0.114
0.329
0.159
1.196
62.674
CodeQwen1.5-7B-Chat
0.507
0.571
0.225
1.120
0
Magicoder-S-DS-6.7B
0.516
0.602
0.251
1.184
0.114
Table 6: Classification performance of the selected SLMs, on DASWOW test dataset
Figure 3: Example of taxonomy mutation, replacing Model Deployment with Prediction
unified
synonyms
Data Collection
Data Acquisition, Data Gathering, Load Data
Data Preparation
Data Preprocessing, Data Wrangling, Data Cleaning
Data Modeling
Modeling, Model Building, Model Construction
Model Deployment
Prediction, Model Prediction, Inference
Model Evaluation
Evaluation, Model Testing, Performance Assessment
Save Results
Persist Results, Store Results, Record Results
Table 7: Collected synonyms
Figure 4: SLM classification performance for the different mutations. The mutation −1 corresponds to the original taxonomy.
Figure 5: Normalized mono-label classification matrix of SLM best classification against DS-Pipelines
Precision
Recall
F 1 score
Support
MCC
Data Collection
0.98
0.77
0.86
470
0.86
Data Modeling
0.95
0.83
0.89
898
0.87
Data Preparation
0.92
0.96
0.94
3144
0.83
Model Evaluation
0.58
0.76
0.66
156
0.65
Model Prediction
0.88
0.94
0.91
307
0.90
Save Results
0.00
0.00
0.00
0
0
Table 8: Mono-label classification report for SLM best on DS-Pipelines
Figure 6: Perplexity per class for SLM best on DS-Pipelines
Figure 7: MCC per class gaps between SLM best and the DASWOW classifier
Figure 8: Normalized multi-label classification matrix of SLM best classification against DASWOW
Precision
Recall
F 1 score
Support
MCC
Data Collection
0.72
0.93
0.81
187
0.79
Data Modeling
0.86
0.64
0.74
314
0.70
Data Preparation
0.72
0.94
0.81
1030
0.39
Model Evaluation
0.41
0.80
0.54
239
0.47
Model Prediction
0.43
0.82
0.57
87
0.56
Save Results
0.30
0.53
0.39
30
0.39
Table 9: Multi-label classification report for SLM best on DASWOW
Figure 9: Perplexity per class for SLM best on DASWOW
Figure 10: Proportion of stages in ML pipelines. The sum of all proportions per classifier can be lower than 100% (some cells may be unlabelled) or higher than 100% (some cells contain multiple labels).
Figure 11: Number of occurrences for each classifier compared to the ground truth. The ground truth was scaled to have the same population than the different classifiers. This scaling allows conducting the goodness of fit test.
Figure 12: Residual matrices representing the differences between the probability matrices (Markov chains) obtained from the different classifiers and the ground truth. The ground truth was scaled to have the same population than the different classifiers. This scaling allows conducting the goodness of fit test.
stage
definition
Data Acquisition
In the beginning of DS pipeline, data are collected from appropriate sources. Data can be acquired manually or automatically. Data acquisition also involves understanding the nature of the data, collecting relevant data, and integrating available datasets.
Data Preparation
Data are generally acquired in a raw format that needs certain preprocessing steps. This involves exploration and filtering, which helps identify the correct data for further processing. Well prepared data reduces the time required for data analysis and contributes to the success of the DS pipeline.
Storage
It is important to find an appropriate hardware-software combination to preserve data so that it can be processed efficiently. For example, Miao et al. used graph database system Neo4j [52] to build a collaborative analysis pipeline [49], since Neo4J supports querying graph data properties.
Feature Engineering
The entire dataset might not contribute equally to decision making. In this stage, appropriate features that are useful to build the model are identified or constructed. Features that are not readily available in the dataset, require engineering to create them from raw data.
Modeling
When data are preprocessed and features are extracted, a model is built to analyze the data. Model building includes model planning, model selection, mining and deriving important properties of data. Appropriate data processing strategies and algorithms are selected to create a good model.
Training
For a specific model, we need to train the model with available labeled data. By each training iteration, we optimize the model and try to make it better. The quality of the training dataset contributes to the training accuracy of the model.
Table 10: Definition of each stage of the DS-Pipelines taxonomy T dspipelines .
stage
definition
helper_functions
Code that is not directly related to the data science activity at hand, but provides useful scripting functions (e. g. importing or configuring libraries).
load_data
The process of loading a dataset of any type (e.g., .csv, .pkl) into a Jupyter notebook environment.
data_preprocessing
The process of preparing the dataset(s) for the subsequent analysis. It includes tasks such as cleaning, instance selection, normalisation, data transformation, and feature selection.
data_exploration
The process of inspecting the content and shape of a dataset to understand the nature and characteristics of the data. Note that it may involve the usage of visualisation techniques but differs in its purpose.
modelling
The process of applying statistical models and learning-based algorithms to learn from sample data.
evaluation
The process of assessing a model using one/various evaluation metric(s) such as goodness of fit and accuracy.
Table 11: Definition of each stage of the DASWOW taxonomy T daswow .
stage
definition
Data Collection
The process of gathering and importing data from various sources into a data analysis environment for further processing and analysis.
Data Preparation
The initial phase of data science where raw data is cleaned, filtered, transformed, and normalized to ensure quality and relevance for analysis.
Data Modeling
The process of creating mathematical or computational representations of real-world phenomena using statistical and machine learning techniques to extract insights and make predictions.
Model Prediction
The application of a previously trained machine learning model to new, unseen data for the purpose of making predictions or decisions.
Model Evaluation
The process of assessing a trained model’s performance using specific metrics and comparing it against new datasets or real-world scenarios to determine its effectiveness and reliability.
Save Results
The process of serializing and persisting data for future use or analysis.
Table 12: Definition of each stage of the best performing taxonomy mutation. Highlighted stage corresponds to the mutated headword.
Figure 13: Normalized multi-label classification matrix of DASWOW classification
stage
adjusted residual
p-value raw
p-value corrected
significant
DASWOW
Data Collection
-1.26E+00
2.06E-01
4.13E-01
False
Data Modeling
-2.16E+00
3.08E-02
9.23E-02
False
Data Preparation
9.97E+00
2.16E-23
1.30E-22
True
Model Evaluation
-7.91E+00
2.60E-15
1.30E-14
True
Model Prediction
-4.76E+00
1.98E-06
7.90E-06
True
Table 15: Adjusted Pearson residuals for the distribution of pipeline stages. Statistically significant residuals appear in grey.
Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability. We introduce SLMJury, a framework for evaluating small language models (SLMs) as judges across two paradigms: closed-ended binary correctness and open-ended quality scoring. We benchmark 16 SLM judges (0.6B-14B parameters) from four model families across ten benchmarks: eight closed-ended tasks spanning mathematical, scientific, and general reasoning (N=64,824 judgments per configuration), plus SummEval and MT-Bench for summarization and conversational scoring. We formalize judging as a budget-conditioned function and study five dimensions. Four findings emerge. (1) The overthinking effect is domain-dependent: for most judges quick 10-token verdicts match or beat extended reasoning on mathematical judging (by 2-7% where they help), while reasoning wins on general tasks by up to 23%. (2) Domain generalization separates model families, with math-to-general accuracy gaps ranging from under 10% to nearly 40%. (3) Closed-ended and open-ended judging draw on different capabilities: the best binary judge (Phi-4) drops to rank 9 on MT-Bench, while reasoning-trained models invert this ordering. (4) Under the Reflect-Critique-Refine (RCR) debate protocol, multi-agent debate degrades accuracy across all tested configurations, whereas the top judges resist six adversarial personas with <=0.55% variance. Reliable automated evaluation does not require large proprietary models, yet no single SLM dominates. The leaderboard is available at https://anishh15.github.io/SLMJury/, and our framework code and pip package are publicly available at https://github.com/anishh15/SLMJury and https://pypi.org/project/slmjury/.
Anish Laddha, Nitesh Pradhan, Gaurav Srivastava
Department of Computer Science and Engineering, LNMIIT, Jaipur, India · Department of Computer Science, Virginia Tech, Blacksburg, VA, USA
Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code. Some recurring source-level quality problems in these pipelines, known as ML code smells, may not cause immediate failures but can harm reproducibility, robustness, efficiency, or maintainability. Detecting ML code smell occurrences is challenging because the decisive evidence is often non-local, spanning helper functions, wrappers, imports, control-flow, and data-flow relations. We present SpecDetect4ML, a static analyser that operationalises 22 ML code smells using CPG views with project-level resolution. We evaluate it on 890 Python ML-based systems comprising more than 20M LOC and a system-level recall benchmark over the complete ML-relevant source subset of 10 selected systems. Under identical ML code smell specifications, CPG-based reasoning raises recall from 68.62% to 88.14% compared with AST-only analysis, while keeping CPG precision comparable at 90.32%. These results show that project-level static reasoning expands the detectable portion of non-local ML code smell occurrences, while configuration-dependent and runtime-only occurrences remain outside our source-only static claims.
Recently, language models have made rapid progress across various domains and applications. However, their capability for self-improvement, i.e., whether they are adept at recognising and correcting flaws in their own reasoning, remains dubious. In this study, we address this question by constructing a sufficiency test to rigorously examine the self-correction capabilities of small language models (SLMs). We propose a minimal three-step self-correction pipeline that collects initial SLM answers, prompts the same model to generate hints for its incorrect responses given the ground truth, and feeds the model the same question with its own feedback to refine the initial answer. We evaluate a variety of instruction-tuned and reasoning SLMs in this experimental setup on arithmetic and logical reasoning benchmarks. Our findings show that SLMs with injected hint sentences yield only a 4.4 percent gain over initial question-answering accuracy. Even though the correct answer was provided alongside the model's incorrect reasoning, the evaluated SLMs fail to understand what was missing in their reasoning and show minimal semantic difference between hints that lead to corrections and ones that do not. Furthermore, our experiments show that longer hints are positively correlated with incorrect final answers, suggesting that longer deliberation on problems can hinder the reasoning process, meaning that SLMs do not necessarily scale in performance with a larger compute budget.