Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.
Figures & tables
Figure 1: Two-level taxonomy of the benchmark dataset collection.
Figure 2: Workflow of the proposed system. Given a dataset, target feature, and optional textual description, the system extracts target-specific statistics, generates a compact serialized representation, and combines the information into a prompt for predicting the data domain and downstream ML task via the LLM.
Target feature name; target feature name + dataset description; target feature name + dataset description + target-specific statistics
Table 1: Hyperparameter space evaluated in the LLM experiments from January to March 2026, including model choice, reasoning mode, prompting strategy, and dataset information configuration.
Figure 3: F1-macro distributions across validation runs for Experiments 1–3. Each violin aggregates all runs for a given hyperparameter across the remaining ones. BP denotes the base prompt, DD dataset descriptions, ST target-specific statistics, and DD+ST their combination.
Few-shot
Zero-shot
Model
Reasoning
Info.
F1 Macro
Balanced Acc.
F1 Macro
Balanced Acc.
AutoGluon
–
–
–
–
0.93 ± 0.01
0.93 ± 0.01
NaiveAutoML
–
–
–
–
0.80 ± 0.01
0.81 ± 0.01
H2O
–
–
–
–
0.59 ± 0.01
0.58 ± 0.01
Qwen3-14B
✗
DD+ST
0.98 ± 0.01
0.98 ± 0.01
0.98 ± 0.01
0.98 ± 0.00
Qwen3-14B
✓
DD+ST
0.98 ± 0.00
0.98 ± 0.00
0.98 ± 0.01
0.98 ± 0.01
Table 2: Test-set results of the best-performing models in Experiment 1 for tabular ML task identification compared to the AutoGluon baseline. Performance metrics for the LLM-based methods are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Table 3: Test-set results of the best-performing models in Experiment 2 for cross-domain ML task identification across tabular and time series datasets. Performance metrics are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Figure 5: Confusion matrices of the best-performing models in the cross-domain setting. 5(a) : best state-of-the-art model in Experiment 2. 5(b) : best locally deployable model in Experiment 3. TS denotes time series, and TB denotes tabular.
Few-shot
Zero-shot
Model
Reasoning
Info.
F1 Macro
Balanced Acc.
F1 Macro
Balanced Acc.
Qwen3-4B-Instruct
✗
DD+ST
0.75 ± 0.02
0.74 ± 0.02
0.73 ± 0.02
0.73 ± 0.02
Qwen3-4B-Thinking
✓
DD+ST
0.62 ± 0.03
0.64 ± 0.03
0.45 ± 0.04
0.45 ± 0.04
Qwen3-4B-Instruct
✗
ST
0.50 ± 0.01
0.58 ± 0.02
0.48 ± 0.03
0.56 ± 0.02
Qwen3-4B-Thinking
✓
ST
0.46 ± 0.02
0.55 ± 0.02
0.45 ± 0.04
0.48 ± 0.04
Table 4: Test-set results of the best-performing locally deployable models in Experiment 3 for cross-domain task identification under resource-constrained settings. Performance metrics are reported as the mean and standard deviation over 1,000 bootstrap resamples. DD denotes textual dataset descriptions and ST denotes target-specific statistics.
Table 5: Repository links for the benchmark dataset, implementation, and LM Contamination Index.
Fee type
Model
In-context learning
Avg. latency per prompt
Cost per prompt / total cost
Paid API
GPT-5.3
Zero-shot
–
0.003/1.13
Few-shot
–
0.009/3.38
Free
AutoGluon
–
0.0004s
-
Free
NaiveAutoML
–
0.01s
-
Free
H2O
–
0.009s
-
Free
Qwen3-14B
Zero-shot
6.03s
-
Appendix
Table 6: Latency and cost by model, baseline, and in-context learning strategy. Latency denotes the average per-prompt inference time for locally deployed models and baselines. GPT-5.3 was evaluated using API batch processing to reduce costs, preventing comparable per-prompt latency measurement. Reported costs are the API charges incurred for each experimental setting between December 2025 and March 2026. GPT-5.3 is a commercial API model, whereas Qwen and AutoGluon incurred no API costs.
Method
Total Tokens
System Prompt Tokens
User Prompt Tokens
Min.
Max.
Avg.
Min.
Max.
Avg.
Min.
Max.
Avg.
Zero-shot
537
15585
2502
238
238
238
287
15335
2252
Few-shot
5435
20483
7400
5136
5136
5136
287
15335
2252
Appendix
Table 7: Token usage statistics per prompt across all datasets.
Figure 6: High-level overview of a representative dataset in the collection.
Figure 7: Scaling behavior of model performance in Dataset Phase 2.
Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.
Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Existing retrieval methods prioritize similarity to the query, but similar demonstrations often reinforce the model's likely prediction rather than reveal the distinctions needed for difficult decisions. We propose EdgeLM, a retrieval framework that instead selects edge evidence, demonstrations that are both relevant to the query and informative about the decision boundary. EdgeLM retrieves two complementary forms of edge evidence by selecting data edges, nearby examples with different ground-truth labels, and model edges, similar examples previously misclassified by the deployed model. EdgeLM requires neither model retraining nor task-specific engineering. Across five data wrangling tasks, fifteen datasets, and five open-weight and proprietary LLMs, EdgeLM consistently achieves the best or near-best performance in every setting, while ablations show that the two forms of edge evidence provide complementary benefits. Our code and datasets are publicly available at https://github.com/soroushomidvar/EdgeLM.
Supervised classification on tabular data remains a central machine learning task, but its dependence on large labeled datasets limits its applicability in data-scarce settings. Few-shot methods such as TabPFN achieve strong performance through large-scale synthetic pretraining, yet still require labeled context examples. Large Language Models (LLMs) offer a more flexible alternative through zero- and few-shot in-context learning from task descriptions, but their behavior on tabular data remains inconsistent. We introduce LLMTabBench, a benchmark for evaluating LLMs on tabular classification under low-data conditions. The benchmark studies how LLM prior knowledge interacts with task descriptions and few-shot examples, and how performance changes with increasing data complexity across real-world and controlled synthetic datasets. We find that LLMs can be highly competitive in zero-shot settings, sometimes outperforming models given few-shot examples. However, additional examples may conflict with prior knowledge, thereby degrading performance. We also observe a complexity threshold at which LLM performance declines and few-shot examples become less useful. These results clarify key limits of in-context learning for tabular data and inform the deployment of LLMs in low-data regimes.