cs.LGSep 28, 2026

Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation

Authors: Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann

Organizations: University of Applied Sciences Aalen · Honda Research Institute Europe

Abstract

Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Why Large Language Models Fail at Tabular Prediction

    Aug 3, 2026Marta Garnelo, Wojciech M. CzarneckiLarge Language Models FailTabular Data

  2. EdgeLM: Edge Demonstrations for Language Models' Table Understanding

    Aug 5, 2026Soroush Omidvartehrani, Mohammadamin Habibollah, Mohammadreza Daviran +1Tabular DataRetrievers

  3. LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots

    May 23, 2026Daria Grushina, Kseniia Kuvshinova, Alina Kostromina +3Tabular Learning