Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.
Figures & tables
Figure 1: Two-level taxonomy of the benchmark dataset collection.
Figure 2: Workflow of the proposed system. Given a dataset, target feature, and optional textual description, the system extracts target-specific statistics, generates a compact serialized representation, and combines the information into a prompt for predicting the data domain and downstream ML task via the LLM.
Target feature name; target feature name + dataset description; target feature name + dataset description + target-specific statistics
Table 1: Hyperparameter space evaluated in the LLM experiments from January to March 2026, including model choice, reasoning mode, prompting strategy, and dataset information configuration.
Figure 3: F1-macro distributions across validation runs for Experiments 1–3. Each violin aggregates all runs for a given hyperparameter across the remaining ones. BP denotes the base prompt, DD dataset descriptions, ST target-specific statistics, and DD+ST their combination.
Few-shot
Zero-shot
Model
Reasoning
Info.
F1 Macro
Balanced Acc.
F1 Macro
Balanced Acc.
AutoGluon
–
–
–
–
0.93 ± 0.01
0.93 ± 0.01
NaiveAutoML
–
–
–
–
0.80 ± 0.01
0.81 ± 0.01
H2O
–
–
–
–
0.59 ± 0.01
0.58 ± 0.01
Qwen3-14B
✗
DD+ST
0.98 ± 0.01
0.98 ± 0.01
0.98 ± 0.01
0.98 ± 0.00
Qwen3-14B
✓
DD+ST
0.98 ± 0.00
0.98 ± 0.00
0.98 ± 0.01
0.98 ± 0.01
Table 2: Test-set results of the best-performing models in Experiment 1 for tabular ML task identification compared to the AutoGluon baseline. Performance metrics for the LLM-based methods are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Table 3: Test-set results of the best-performing models in Experiment 2 for cross-domain ML task identification across tabular and time series datasets. Performance metrics are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Figure 5: Confusion matrices of the best-performing models in the cross-domain setting. 5(a) : best state-of-the-art model in Experiment 2. 5(b) : best locally deployable model in Experiment 3. TS denotes time series, and TB denotes tabular.
Few-shot
Zero-shot
Model
Reasoning
Info.
F1 Macro
Balanced Acc.
F1 Macro
Balanced Acc.
Qwen3-4B-Instruct
✗
DD+ST
0.75 ± 0.02
0.74 ± 0.02
0.73 ± 0.02
0.73 ± 0.02
Qwen3-4B-Thinking
✓
DD+ST
0.62 ± 0.03
0.64 ± 0.03
0.45 ± 0.04
0.45 ± 0.04
Qwen3-4B-Instruct
✗
ST
0.50 ± 0.01
0.58 ± 0.02
0.48 ± 0.03
0.56 ± 0.02
Qwen3-4B-Thinking
✓
ST
0.46 ± 0.02
0.55 ± 0.02
0.45 ± 0.04
0.48 ± 0.04
Table 4: Test-set results of the best-performing locally deployable models in Experiment 3 for cross-domain task identification under resource-constrained settings. Performance metrics are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Table 5: Repository links for the benchmark dataset, implementation, and LM Contamination Index.
Fee type
Model
In-context learning
Avg. latency per prompt
Cost per prompt / total cost
Paid API
GPT-5.3
Zero-shot
–
0.003/1.13
Few-shot
–
0.009/3.38
Free
AutoGluon
–
0.0004s
-
Free
NaiveAutoML
–
0.01s
-
Free
H2O
–
0.009s
-
Free
Qwen3-14B
Zero-shot
6.03s
-
Appendix
Table 6: Latency and cost by model, baseline, and in-context learning strategy. Latency denotes the average per-prompt inference time for locally deployed models and baselines. GPT-5.3 was evaluated using API batch processing to reduce costs, preventing comparable per-prompt latency measurement. Reported costs are the API charges incurred for each experimental setting between December 2025 and March 2026. GPT-5.3 is a commercial API model, whereas Qwen and AutoGluon incurred no API costs.
Method
Total Tokens
System Prompt Tokens
User Prompt Tokens
Min.
Max.
Avg.
Min.
Max.
Avg.
Min.
Max.
Avg.
Zero-shot
537
15585
2502
238
238
238
287
15335
2252
Few-shot
5435
20483
7400
5136
5136
5136
287
15335
2252
Appendix
Table 7: Token usage statistics per prompt across all datasets.
Figure 6: High-level overview of a representative dataset in the collection.
Figure 7: Scaling behavior of model performance in Dataset Phase 2.