From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification
Authors: Yue Qiu, Zekang Du, Yiqun Diao, Bingsheng He, Qinbin Li
Organizations: School of Computer Science and Technology, Huazhong University of Science and Technology · School of Computing, National University of Singapore
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.
Figures & tables
Features
LLMT
Direct-ZSDT
Step-ZSDT
CoT-Tree
ToT-Tree
IO-Tree
TabLLM
InsightTab
SumBoost
FeatLLM
DeLTa
GPTree
Few-shot ready (no finetune)
✓
✗
✗
✓
✓
✓
✗
✓
✓
✓
✓
✓
Token-efficient build
✓
✓
✗
✓
✗
✓
✗
✗
✗
✗
✗
✗
Tree generation
✓
✓
✓
✓
✓
✓
✗
✗
✗
✗
✗
✓
No LLM at inference
✓
✓
✓
✓
✓
✓
✗
✗
✗
✓
✓
✓
Rule-first (vs path-first)
✓
✗
✗
✗
✗
✗
✗
✗
✗
✓
✗
✓
Sample-level privacy preservation
✓
✓
✓
✗
✗
✗
✗
✗
✗
✗
✓
✗
Table 1: Comparison of our LLM-guided tree method ( LLMT ) with other related studies
Figure 1: The test accuracy of LLMs with or without the metadata about the dataset.
Figure 2: Comparison of Utility and Accuracy between generated rule sets and path sets
Figure 3: Comparison between our method and prompting-based methods for generating reasoning trees. Instead of (1) asking the LLM to directly generate long reasoning paths (CoT-Tree), or (2) asking the LLM to generate the next thought based on a long reasoning path, our method prompts the LLM to generate discrete rules, organizes them into a tree using statistical methods, and then uses the LLM for final refinement.
Dataset
DirectZSDT
StepZSDT
LogReg
SVM
CART
XGBoost
LightGBM
IO-Tree
CoT-Tree
ToT-Tree
FeatLLM
DeLTa
GPTree
LLMT
Nursery
0.485
0.450
0.558
0.564
0.602
0.434
0.578
0.439
0.461
0.353
0.571
0.448
0.479
0.762
Diabetes
0.729
0.641
0.618
0.629
0.546
0.690
0.543
0.539
0.701
0.561
0.700
0.541
0.616
0.764
Spambase
0.645
0.732
0.790
0.783
0.723
0.600
0.733
0.723
0.739
0.727
0.802
0.747
0.661
0.811
Abalone
0.550
0.526
0.684
0.672
0.689
0.670
0.658
0.647
0.694
0.570
0.721
0.656
0.667
0.694
Blood
0.597
0.371
0.604
0.635
0.565
0.605
0.508
0.585
0.500
0.551
0.642
0.550
0.572
0.675
Iris
0.946
0.547
0.793
0.756
0.882
0.739
0.685
0.870
0.903
0.867
0.903
0.823
0.814
0.940
Table 2: Accuracy comparison of different methods when the number of training examples per class is two. Best performances are bolded, and our method’s performances, when second-best, are underlined.
Datasets
Nursery
Diabetes
Spambase
#Shots
3
6
12
24
48
2
4
8
16
32
2
4
8
16
32
DirectZSDT
0.485
0.485
0.485
0.485
0.485
0.729
0.729
0.729
0.729
0.729
0.645
0.645
0.645
0.645
0.645
StepZSDT
0.450
0.450
0.450
0.450
0.450
0.641
0.641
0.641
0.641
0.641
0.732
0.732
0.732
0.732
0.732
LogReg
0.505
0.558
0.670
0.778
0.855
0.542
0.618
0.674
0.730
0.748
0.685
0.790
0.785
0.782
0.835
SVM
0.500
0.564
0.661
0.781
0.833
0.542
0.629
0.660
0.671
0.710
0.685
0.783
0.754
0.771
0.807
CART
0.351
0.602
0.786
0.803
0.801
0.519
0.546
0.624
0.663
0.686
0.639
0.723
0.743
0.705
0.757
Table 3: Mean accuracy over ten runs across varying #shots (standard deviation is available at Appendix G ). DirectZSDT and StepZSDT are constant across #shots since they cannot utilize any training examples.
Time (s)
Token (#Input token, #Output token, #Total)
Diabetes
Nursery
Spambase
Diabetes
Nursery
Spambase
DirectZSDT
43.64
31.75
55.33
(2.9k, 1.1k, 4k)
(3.8k, 1.3k, 4.1k)
(3.4k, 1.4k, 4.8k)
StepZSDT
1006.67
1191.40
4036.18
(85.6k, 22.1k, 107.7k)
(128.3k, 27.2k, 155.5k)
(358.9k, 95.4k, 454.3k)
CoT-Tree
17.00
9.58
12.05
(2.7k, 1.3k, 4k)
(2.7k, 1.0k, 3.7k)
(5.6k, 0.9k, 6.5k)
ToT-Tree
33.39
33.22
64.56
(32.9k, 0.5k, 33.4k)
(35.4k, 0.6k, 36k)
(140k, 1.3k, 141.3k)
FeatLLM
259.07
416.16
324.06
(22.2k, 10.3k, 32.5k)
(24.4k, 16.3k, 40.7k)
(48.5k, 12.3k, 60.8k)
Table 4: The training time (s) and #token in the #shots=32 setting. The savings of LLMT are computed against ToT-Tree in terms of #total tokens.
Figure 4: Average number of nodes in trees generated by each method. Fewer nodes indicate a more concise tree.
Dataset
Direct-ZSDT
Step-ZSDT
CART
IO-Tree
CoT-Tree
ToT-Tree
GPTree
LLMT
Diabetes
0.422
0.420
0.414
0.397
0.397
0.403
0.425
0.396
Nursery
0.665
0.634
0.516
0.582
0.534
0.629
0.552
0.456
Spambase
0.378
0.478
0.360
0.341
0.361
0.351
0.396
0.352
Average
0.488
0.511
0.430
0.440
0.431
0.461
0.458
0.401
Table 5: Average Gini impurity of rules across different methods. Lower values indicate better rule quality.
Figure 5: The accuracy of LLMT with different τ
Figure 6: The accuracy of LLMT with/without leaf refinement
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: LLMT Example Tree (Step-by-Step)
Dataset
#Instances
#Features
#Classes
Nursery
12,960
8
3
Diabetes
768
8
2
Spambase
4,601
18
2
Abalone
4,177
8
2
Blood
748
4
2
Iris
150
4
3
Appendix
Table 6: Statistics of datasets used in our experiments.
Dataset
Model
Original
Masked
Δ
Dataset
Model
Original
Masked
Δ
Breast
DeepSeek-V3
56.0
50.0
6.0
Paddy
DeepSeek-V3
34.5
31.5
3.0
Breast
Llama-4
56.9
49.1
7.8
Paddy
Llama-4
40.0
37.5
2.5
Breast
Gemma-2
55.2
52.6
2.6
Paddy
Gemma-2
37.0
32.5
4.5
Breast
Qwen2.5
54.3
54.3
0.0
Paddy
Qwen2.5
32.5
30.5
2.0
Glioma
DeepSeek-V3
89.0
34.0
55.0
Personality
DeepSeek-V3
93.0
59.0
34.0
Glioma
Llama-4
89.0
35.0
54.0
Personality
Llama-4
93.0
64.0
29.0
Appendix
Table 7: Original-versus-masked zero-shot accuracy (%) on six newer datasets. Δ is Original minus Masked in percentage points. Llama-4 denotes Llama-4 Maverick, Gemma-2 denotes Gemma-2-27B-IT, and Qwen2.5 denotes Qwen2.5-72B-Instruct.
Condition budget n=3
Condition budget n=4
Dataset
Rule U.
Path U.
Rule A.
Path A.
Rule U.
Path U.
Rule A.
Path A.
Nursery
31.87
28.32
24.75
25.86
41.34
41.35
31.20
33.32
Diabetes
35.28
32.21
32.60
29.48
64.52
65.14
50.62
50.29
Spambase
54.55
41.06
45.44
36.64
60.42
53.63
52.04
47.74
Abalone
33.41
29.68
33.05
29.45
53.11
32.23
51.42
32.05
Blood
40.01
35.45
33.92
32.03
64.24
31.66
43.48
24.49
Appendix
Table 8: Rule-set and path-set results under matched condition budgets (%). U. and A. denote Set Utility and Set Accuracy. Each entry is averaged over 100 generations.
Dataset
Max Depth
n_estimators
Diabetes
3
3
Nursery
3
3
Spambase
4
6
Abalone
3
3
Blood
3
3
Iris
3
3
Appendix
Table 9: Dataset-specific parameters for all experiments.
Parameter
Value
Model
Qwen2.5-72B-Instruct
API Provider
TogetherAI
Temperature
0.0
Max Tokens
2048
Parallel Batch Size
6
Presence Penalty
0.0
Appendix
Table 10: LLM API configuration settings used for all LLM-based methods.
Method
Parameter
Value
Tree methods
Maximum depth
Table 9
Tree ensembles
Ensemble size
Table 9
Step-ZSDT
Leaf-stopping threshold τ
0.9
XGBoost
learning_rate
0.1
LogReg
max_iter
1000
Linear SVM
Estimator
LinearSVC
Appendix
Table 11: Method-specific settings.
Parameter
Value
Number of Trials per Setting
10
Random Seed
0
Train Batch Size
8
Test Batch Size
8
Appendix
Table 12: General experimental running parameters.
Figure 8: The accuracy and standard deviation of different methods on three datasets.
Dataset
Random Forest
CatBoost
LLMT
Nursery
0.397
0.473
0.762
Diabetes
0.584
0.590
0.764
Spambase
0.725
0.743
0.811
Abalone
0.688
0.700
0.694
Blood
0.549
0.558
0.675
Iris
0.769
0.829
0.940
Appendix
Table 13: Accuracy of Random Forest and CatBoost with two training examples per class, compared with LLMT.
Dataset
#Shots
Random Forest
CatBoost
LLMT
Nursery
3
0.349
0.464
0.762
6
0.397
0.473
0.762
12
0.707
0.411
0.784
24
0.754
0.721
0.818
48
0.798
0.740
0.840
Diabetes
2
0.486
0.523
0.662
Appendix
Table 14: Accuracy of Random Forest and CatBoost across shot counts, compared with LLMT.
Dataset
TabPFN
TabLLM
LLMT
LLMT Forest
Nursery
0.750
0.759
0.762
–
Diabetes
0.710
0.510
0.764
–
Spambase
0.802
0.673
0.811
–
Abalone
0.754
0.569
0.694
–
Blood
0.658
0.496
0.675
–
Iris
0.918
0.327
0.940
–
Appendix
Table 15: Comparison with pretrained black-box few-shot baselines. TabLLM runs out of memory on Communities and Ecom and is prohibitively slow on Myocardial under the T0pp protocol; these cases are marked N/A.
Dataset
LogReg
SVM
CART
Direct-ZSDT
IO-Tree
CoT-Tree
ToT-Tree
FeatLLM
GPTree
LLMT (Ours)
Communities
0.424
0.439
0.404
0.418
0.336
0.411
0.402
0.432
0.407
0.438
Ecom
0.509
0.508
0.554
0.561
0.511
0.524
0.534
0.579
0.490
0.582
Myocardial
0.526
0.522
0.540
0.510
0.535
0.503
0.498
0.596
0.556
0.618
Average
0.486
0.490
0.499
0.496
0.461
0.479
0.478
0.536
0.484
0.546
Appendix
Table 16: Accuracy comparison of non-ensemble methods on three high-dimensional datasets.
Dataset
XGBoost
RandomForest
LightGBM
CatBoost
DeLTa
LLMT Forest (Ours)
Communities
0.428
0.440
0.434
0.437
0.473
0.490
Ecom
0.695
0.504
0.412
0.503
0.619
0.757
Myocardial
0.702
0.545
0.581
0.535
0.565
0.773
Average
0.608
0.496
0.476
0.492
0.552
0.673
Appendix
Table 17: Accuracy comparison of ensemble methods on three high-dimensional datasets.
Dataset
Diabetes
Iris
Spambase
Nursery
Abalone
Blood
Breast
Glioma
#Rules ( K )
10
10
15
10
10
10
10
15
#Features
8
4
18
8
8
4
9
23
Avg. Features Used
8.0
4.0
15.0
8.0
8.0
4.0
9.0
14.9
Feature Coverage
100%
100%
83.3%
100%
100%
100%
100%
64.8%
Appendix
Table 18: Rule diversity analysis: Feature coverage (Avg. Features Used / #Features) measures the proportion of features utilized in the generated rules, with higher values indicating better diversity.
Figure 9: An Example of Direct-ZSDT
Figure 10: An Example of Step-ZSDT
Figure 11: An Example of CART Tree
Figure 12: An Example of IO-Tree
Figure 13: An Example of CoT-Tree
Figure 14: An Example of ToT-Tree
Figure 15: An Example of GPTree
Figure 16: An Example of LLMT
Datasets
#Shots
2
3
4
5
Nursery
3
0.740 ± 0.000
0.762 ± 0.100
0.762 ± 0.100
0.762 ± 0.100
6
0.740 ± 0.000
0.762 ± 0.100
0.707 ± 0.063
0.731 ± 0.077
12
0.740 ± 0.000
0.784 ± 0.091
0.742 ± 0.074
0.737 ± 0.054
24
0.740 ± 0.000
0.818 ± 0.066
0.713 ± 0.084
0.735 ± 0.045
48
0.740 ± 0.000
0.840 ± 0.000
0.788 ± 0.042
0.742 ± 0.039
Diabetes
2
0.662 ± 0.095
0.662 ± 0.095
0.662 ± 0.095
0.669 ± 0.083
Appendix
Table 19: Experimental results across different maximum tree depths.
Figure 17: Meta Rule Generation Prompting Template.
Figure 18: Meta Rule Generation Prompting Template (continued).
Figure 19: Leaf Refinement Prompting Template.
Figure 20: IO Tree Generation Prompting Template.
Figure 21: IO Tree Generation Prompting Template (continued).
Figure 22: CoT Tree Generation Prompting Template.
Figure 23: CoT Tree Generation Prompting Template (continued).
Figure 24: ToT Tree Generation Prompting Template.
Figure 25: ToT Tree Generation Prompting Template (continued).
Supervised classification on tabular data remains a central machine learning task, but its dependence on large labeled datasets limits its applicability in data-scarce settings. Few-shot methods such as TabPFN achieve strong performance through large-scale synthetic pretraining, yet still require labeled context examples. Large Language Models (LLMs) offer a more flexible alternative through zero- and few-shot in-context learning from task descriptions, but their behavior on tabular data remains inconsistent. We introduce LLMTabBench, a benchmark for evaluating LLMs on tabular classification under low-data conditions. The benchmark studies how LLM prior knowledge interacts with task descriptions and few-shot examples, and how performance changes with increasing data complexity across real-world and controlled synthetic datasets. We find that LLMs can be highly competitive in zero-shot settings, sometimes outperforming models given few-shot examples. However, additional examples may conflict with prior knowledge, thereby degrading performance. We also observe a complexity threshold at which LLM performance declines and few-shot examples become less useful. These results clarify key limits of in-context learning for tabular data and inform the deployment of LLMs in low-data regimes.
Large language models (LLMs) have recently been adapted to tabular prediction by serializing structured features into natural language, but their performance in low-data regimes remains limited compared to gradient-boosted decision trees (GBDTs). In this work, we revisit the boosting paradigm, traditionally associated with tree ensembles, and ask whether it can be applied as a general training principle for LLM fine-tuning. We propose BoostLLM, a framework that transforms parameter-efficient fine-tuning into a multi-round residual optimization process by training sequential PEFT adapters as weak learners. To incorporate tabular inductive bias, BoostLLM integrates decision-tree paths as a second input view alongside raw features; analysis reveals that the path view acts as a structured teacher in early training steps before the model shifts toward feature-driven representations. Empirically, BoostLLM achieves consistent improvements over standard fine-tuning across multiple LLM backbones and datasets, matching or surpassing XGBoost across a wide range of shot counts and outperforming GPT-4o-based methods with a 4B model. We further show that the framework scales: pairing with stronger tree models and extended boosting horizons yields additional gains under appropriate stabilization. These results suggest that boosting can serve as a general training principle for LLM fine-tuning, particularly in low-data regimes for structured data.
Large Language Models (LLMs) fine-tuned on serialized tabular data are emerging as powerful alternatives to traditional tree-based models, particularly for heterogeneous or context-rich datasets. However, their deployment in high-stakes domains is hindered by a lack of faithful interpretability; existing methods often rely on global linear proxies or scalar probability shifts that fail to capture the model's full probabilistic uncertainty. In this work, we introduce TabSHAP, a model-agnostic interpretability framework designed to directly attribute local query decision logic in LLM-based tabular classifiers. By adapting a Shapley-style sampled-coalition estimator with Jensen-Shannon divergence between full-input and masked-input class distributions, TabSHAP quantifies the distributional impact of each feature rather than simple prediction flips. To align with tabular semantics, we mask at the level of serialized key:value fields (atomic in the prompt string), not individual subword tokens. Experimental validation on the Adult Income and Heart Disease benchmarks demonstrates that TabSHAP isolates critical diagnostic features, achieving significantly higher faithfulness than random baselines and XGBoost proxies. We further run a distance-metric ablation on the same test instances and TabSHAP settings: attributions are recomputed with KL or L1 replacing JSD in the similarity step (results cached per metric), and we compare deletion faithfulness across all three.
Aryan Chaudhary, Prateek Agarwal, Tejasvi Alladi
Department of Computer Science Birla Institute of Technology and Science India