Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
Organizations: Université Côte d’Azur, Inria, CNRS, I3S, Sophia Antipolis, France · McSCert, McMaster University, Hamilton, ON, Canada
Abstract
Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.
Figures & tables
| scenario | odds ratio | sample size | ||
| Conservative ( increase) | 0.10 | 0.15 | 1.5 | 796 |
| Moderate ( increase) | 0.10 | 0.20 | 2.0 | 240 |
| Optimistic ( increase) | 0.05 | 0.25 | 1.5 | 57 |
| scenario | effect size | dof | sample size |
| Small dof | |||
| Conservative (small effect) | 0.10 | 1 | 785 |
| Moderate (medium effect) | 0.30 | 1 | 88 |
| Optimistic (large effect) | 0.50 | 1 | 32 |
| Large dof | |||
| Conservative (small effect) | 0.10 | 46 | 2919 |
| Rank | Model | # Params (B) | Score |
| HumanEval | |||
| 19 | Qwen2.5-Coder 7B Instruct | 7 | 0.884 |
| 35 | Qwen2.5 7B Instruct | 7.61 | 0.848 |
| 40 | IBM Granite 4.0 Tiny Preview | 7 | 0.824 |
| 44 | Qwen2 7B Instruct | 7.62 | 0.799 |
| 45 | Qwen2.5-Omni-7B | 7 | 0.787 |
| Metric | |
| Overall Agreement | |
| Fleiss’ (3 experts) | 0.89 |
| Pairwise Agreement | |
| Cohen’s ( E 1 , E 2 ) | 0.84 |
| Cohen’s ( E 1 , E 3 ) | 1.00 |
| Cohen’s ( E 2 , E 3 ) | 0.84 |
| SLM name | accuracy | F 1 score | mcc | perplexity | error rate (%) |
| Qwen2.5-Coder-7B-Instruct | 0.586 | 0.754 | 0.494 | 1.172 | 0.0095 |
| IBM Granite 4.0 | 0.326 | 0.624 | 0.306 | 1.223 | 0.152 |
| Gemma 3 4B | 0.330 | 0.641 | 0.365 | 1.035 | 0 |
| Phi-3-mini-128k-instruct | 0.114 | 0.329 | 0.159 | 1.196 | 62.674 |
| CodeQwen1.5-7B-Chat | 0.507 | 0.571 | 0.225 | 1.120 | 0 |
| Magicoder-S-DS-6.7B | 0.516 | 0.602 | 0.251 | 1.184 | 0.114 |
| unified | synonyms |
| Data Collection | Data Acquisition, Data Gathering, Load Data |
| Data Preparation | Data Preprocessing, Data Wrangling, Data Cleaning |
| Data Modeling | Modeling, Model Building, Model Construction |
| Model Deployment | Prediction, Model Prediction, Inference |
| Model Evaluation | Evaluation, Model Testing, Performance Assessment |
| Save Results | Persist Results, Store Results, Record Results |
| Precision | Recall | F 1 score | Support | MCC | |
| Data Collection | 0.98 | 0.77 | 0.86 | 470 | 0.86 |
| Data Modeling | 0.95 | 0.83 | 0.89 | 898 | 0.87 |
| Data Preparation | 0.92 | 0.96 | 0.94 | 3144 | 0.83 |
| Model Evaluation | 0.58 | 0.76 | 0.66 | 156 | 0.65 |
| Model Prediction | 0.88 | 0.94 | 0.91 | 307 | 0.90 |
| Save Results | 0.00 | 0.00 | 0.00 | 0 | 0 |
| Precision | Recall | F 1 score | Support | MCC | |
| Data Collection | 0.72 | 0.93 | 0.81 | 187 | 0.79 |
| Data Modeling | 0.86 | 0.64 | 0.74 | 314 | 0.70 |
| Data Preparation | 0.72 | 0.94 | 0.81 | 1030 | 0.39 |
| Model Evaluation | 0.41 | 0.80 | 0.54 | 239 | 0.47 |
| Model Prediction | 0.43 | 0.82 | 0.57 | 87 | 0.56 |
| Save Results | 0.30 | 0.53 | 0.39 | 30 | 0.39 |
| stage | definition |
| Data Acquisition | In the beginning of DS pipeline, data are collected from appropriate sources. Data can be acquired manually or automatically. Data acquisition also involves understanding the nature of the data, collecting relevant data, and integrating available datasets. |
| Data Preparation | Data are generally acquired in a raw format that needs certain preprocessing steps. This involves exploration and filtering, which helps identify the correct data for further processing. Well prepared data reduces the time required for data analysis and contributes to the success of the DS pipeline. |
| Storage | It is important to find an appropriate hardware-software combination to preserve data so that it can be processed efficiently. For example, Miao et al. used graph database system Neo4j [52] to build a collaborative analysis pipeline [49], since Neo4J supports querying graph data properties. |
| Feature Engineering | The entire dataset might not contribute equally to decision making. In this stage, appropriate features that are useful to build the model are identified or constructed. Features that are not readily available in the dataset, require engineering to create them from raw data. |
| Modeling | When data are preprocessed and features are extracted, a model is built to analyze the data. Model building includes model planning, model selection, mining and deriving important properties of data. Appropriate data processing strategies and algorithms are selected to create a good model. |
| Training | For a specific model, we need to train the model with available labeled data. By each training iteration, we optimize the model and try to make it better. The quality of the training dataset contributes to the training accuracy of the model. |
| stage | definition |
| helper_functions | Code that is not directly related to the data science activity at hand, but provides useful scripting functions (e. g. importing or configuring libraries). |
| load_data | The process of loading a dataset of any type (e.g., .csv, .pkl) into a Jupyter notebook environment. |
| data_preprocessing | The process of preparing the dataset(s) for the subsequent analysis. It includes tasks such as cleaning, instance selection, normalisation, data transformation, and feature selection. |
| data_exploration | The process of inspecting the content and shape of a dataset to understand the nature and characteristics of the data. Note that it may involve the usage of visualisation techniques but differs in its purpose. |
| modelling | The process of applying statistical models and learning-based algorithms to learn from sample data. |
| evaluation | The process of assessing a model using one/various evaluation metric(s) such as goodness of fit and accuracy. |
| stage | definition |
| Data Collection | The process of gathering and importing data from various sources into a data analysis environment for further processing and analysis. |
| Data Preparation | The initial phase of data science where raw data is cleaned, filtered, transformed, and normalized to ensure quality and relevance for analysis. |
| Data Modeling | The process of creating mathematical or computational representations of real-world phenomena using statistical and machine learning techniques to extract insights and make predictions. |
| Model Prediction | The application of a previously trained machine learning model to new, unseen data for the purpose of making predictions or decisions. |
| Model Evaluation | The process of assessing a trained model’s performance using specific metrics and comparing it against new datasets or real-world scenarios to determine its effectiveness and reliability. |
| Save Results | The process of serializing and persisting data for future use or analysis. |
| stage | adjusted residual | p-value raw | p-value corrected | significant |
| DASWOW | ||||
| Data Collection | -1.26E+00 | 2.06E-01 | 4.13E-01 | False |
| Data Modeling | -2.16E+00 | 3.08E-02 | 9.23E-02 | False |
| Data Preparation | 9.97E+00 | 2.16E-23 | 1.30E-22 | True |
| Model Evaluation | -7.91E+00 | 2.60E-15 | 1.30E-14 | True |
| Model Prediction | -4.76E+00 | 1.98E-06 | 7.90E-06 | True |