cs.SEOct 7, 2026

Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

Authors: Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino, Sebastien Mosser

Organizations: Université Côte d’Azur, Inria, CNRS, I3S, Sophia Antipolis, France · McSCert, McMaster University, Hamilton, ON, Canada

Abstract

Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.

Figures & tables

Explore similar work

CardsList
  1. SLMJury: Can Small Language Models Judge as Well as Large Ones?

    Jun 5, 2026Anish Laddha, Nitesh Pradhan, Gaurav SrivastavaLarge Language Model JudgesRaw Judge Outputs

  2. SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs

    Sep 24, 2025Brahim Mahmoudi, Naouel Moha, Quentin Stiévenart +1Program AnalysisCode Quality

  3. More Yap Less Meaning: Uncovering Self-Improvement Behavior in SLMs

    Jun 7, 2026Marina Igitkhanian, Erik ArakelyanSmall Large Language ModelsLLM Reasoning Strategies