From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification
Organizations: School of Computer Science and Technology, Huazhong University of Science and Technology · School of Computing, National University of Singapore
Abstract
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.
Figures & tables
| Features | LLMT | Direct-ZSDT | Step-ZSDT | CoT-Tree | ToT-Tree | IO-Tree | TabLLM | InsightTab | SumBoost | FeatLLM | DeLTa | GPTree |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Few-shot ready (no finetune) | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Token-efficient build | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Tree generation | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| No LLM at inference | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Rule-first (vs path-first) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ |
| Sample-level privacy preservation | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ |
| Dataset | DirectZSDT | StepZSDT | LogReg | SVM | CART | XGBoost | LightGBM | IO-Tree | CoT-Tree | ToT-Tree | FeatLLM | DeLTa | GPTree | LLMT |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Nursery | 0.485 | 0.450 | 0.558 | 0.564 | 0.602 | 0.434 | 0.578 | 0.439 | 0.461 | 0.353 | 0.571 | 0.448 | 0.479 | 0.762 |
| Diabetes | 0.729 | 0.641 | 0.618 | 0.629 | 0.546 | 0.690 | 0.543 | 0.539 | 0.701 | 0.561 | 0.700 | 0.541 | 0.616 | 0.764 |
| Spambase | 0.645 | 0.732 | 0.790 | 0.783 | 0.723 | 0.600 | 0.733 | 0.723 | 0.739 | 0.727 | 0.802 | 0.747 | 0.661 | 0.811 |
| Abalone | 0.550 | 0.526 | 0.684 | 0.672 | 0.689 | 0.670 | 0.658 | 0.647 | 0.694 | 0.570 | 0.721 | 0.656 | 0.667 | 0.694 |
| Blood | 0.597 | 0.371 | 0.604 | 0.635 | 0.565 | 0.605 | 0.508 | 0.585 | 0.500 | 0.551 | 0.642 | 0.550 | 0.572 | 0.675 |
| Iris | 0.946 | 0.547 | 0.793 | 0.756 | 0.882 | 0.739 | 0.685 | 0.870 | 0.903 | 0.867 | 0.903 | 0.823 | 0.814 | 0.940 |
| Datasets | Nursery | Diabetes | Spambase | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #Shots | 3 | 6 | 12 | 24 | 48 | 2 | 4 | 8 | 16 | 32 | 2 | 4 | 8 | 16 | 32 |
| DirectZSDT | 0.485 | 0.485 | 0.485 | 0.485 | 0.485 | 0.729 | 0.729 | 0.729 | 0.729 | 0.729 | 0.645 | 0.645 | 0.645 | 0.645 | 0.645 |
| StepZSDT | 0.450 | 0.450 | 0.450 | 0.450 | 0.450 | 0.641 | 0.641 | 0.641 | 0.641 | 0.641 | 0.732 | 0.732 | 0.732 | 0.732 | 0.732 |
| LogReg | 0.505 | 0.558 | 0.670 | 0.778 | 0.855 | 0.542 | 0.618 | 0.674 | 0.730 | 0.748 | 0.685 | 0.790 | 0.785 | 0.782 | 0.835 |
| SVM | 0.500 | 0.564 | 0.661 | 0.781 | 0.833 | 0.542 | 0.629 | 0.660 | 0.671 | 0.710 | 0.685 | 0.783 | 0.754 | 0.771 | 0.807 |
| CART | 0.351 | 0.602 | 0.786 | 0.803 | 0.801 | 0.519 | 0.546 | 0.624 | 0.663 | 0.686 | 0.639 | 0.723 | 0.743 | 0.705 | 0.757 |
| Time (s) | Token (#Input token, #Output token, #Total) | |||||
|---|---|---|---|---|---|---|
| Diabetes | Nursery | Spambase | Diabetes | Nursery | Spambase | |
| DirectZSDT | 43.64 | 31.75 | 55.33 | (2.9k, 1.1k, 4k) | (3.8k, 1.3k, 4.1k) | (3.4k, 1.4k, 4.8k) |
| StepZSDT | 1006.67 | 1191.40 | 4036.18 | (85.6k, 22.1k, 107.7k) | (128.3k, 27.2k, 155.5k) | (358.9k, 95.4k, 454.3k) |
| CoT-Tree | 17.00 | 9.58 | 12.05 | (2.7k, 1.3k, 4k) | (2.7k, 1.0k, 3.7k) | (5.6k, 0.9k, 6.5k) |
| ToT-Tree | 33.39 | 33.22 | 64.56 | (32.9k, 0.5k, 33.4k) | (35.4k, 0.6k, 36k) | (140k, 1.3k, 141.3k) |
| FeatLLM | 259.07 | 416.16 | 324.06 | (22.2k, 10.3k, 32.5k) | (24.4k, 16.3k, 40.7k) | (48.5k, 12.3k, 60.8k) |
| Dataset | Direct-ZSDT | Step-ZSDT | CART | IO-Tree | CoT-Tree | ToT-Tree | GPTree | LLMT |
|---|---|---|---|---|---|---|---|---|
| Diabetes | 0.422 | 0.420 | 0.414 | 0.397 | 0.397 | 0.403 | 0.425 | 0.396 |
| Nursery | 0.665 | 0.634 | 0.516 | 0.582 | 0.534 | 0.629 | 0.552 | 0.456 |
| Spambase | 0.378 | 0.478 | 0.360 | 0.341 | 0.361 | 0.351 | 0.396 | 0.352 |
| Average | 0.488 | 0.511 | 0.430 | 0.440 | 0.431 | 0.461 | 0.458 | 0.401 |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | #Instances | #Features | #Classes |
|---|---|---|---|
| Nursery | 12,960 | 8 | 3 |
| Diabetes | 768 | 8 | 2 |
| Spambase | 4,601 | 18 | 2 |
| Abalone | 4,177 | 8 | 2 |
| Blood | 748 | 4 | 2 |
| Iris | 150 | 4 | 3 |
| Dataset | Model | Original | Masked | Dataset | Model | Original | Masked | ||
|---|---|---|---|---|---|---|---|---|---|
| Breast | DeepSeek-V3 | 56.0 | 50.0 | 6.0 | Paddy | DeepSeek-V3 | 34.5 | 31.5 | 3.0 |
| Breast | Llama-4 | 56.9 | 49.1 | 7.8 | Paddy | Llama-4 | 40.0 | 37.5 | 2.5 |
| Breast | Gemma-2 | 55.2 | 52.6 | 2.6 | Paddy | Gemma-2 | 37.0 | 32.5 | 4.5 |
| Breast | Qwen2.5 | 54.3 | 54.3 | 0.0 | Paddy | Qwen2.5 | 32.5 | 30.5 | 2.0 |
| Glioma | DeepSeek-V3 | 89.0 | 34.0 | 55.0 | Personality | DeepSeek-V3 | 93.0 | 59.0 | 34.0 |
| Glioma | Llama-4 | 89.0 | 35.0 | 54.0 | Personality | Llama-4 | 93.0 | 64.0 | 29.0 |
| Condition budget | Condition budget | |||||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Rule U. | Path U. | Rule A. | Path A. | Rule U. | Path U. | Rule A. | Path A. |
| Nursery | 31.87 | 28.32 | 24.75 | 25.86 | 41.34 | 41.35 | 31.20 | 33.32 |
| Diabetes | 35.28 | 32.21 | 32.60 | 29.48 | 64.52 | 65.14 | 50.62 | 50.29 |
| Spambase | 54.55 | 41.06 | 45.44 | 36.64 | 60.42 | 53.63 | 52.04 | 47.74 |
| Abalone | 33.41 | 29.68 | 33.05 | 29.45 | 53.11 | 32.23 | 51.42 | 32.05 |
| Blood | 40.01 | 35.45 | 33.92 | 32.03 | 64.24 | 31.66 | 43.48 | 24.49 |
| Dataset | Max Depth | n_estimators |
|---|---|---|
| Diabetes | 3 | 3 |
| Nursery | 3 | 3 |
| Spambase | 4 | 6 |
| Abalone | 3 | 3 |
| Blood | 3 | 3 |
| Iris | 3 | 3 |
| Parameter | Value |
|---|---|
| Model | Qwen2.5-72B-Instruct |
| API Provider | TogetherAI |
| Temperature | 0.0 |
| Max Tokens | 2048 |
| Parallel Batch Size | 6 |
| Presence Penalty | 0.0 |
| Method | Parameter | Value |
| Tree methods | Maximum depth | Table 9 |
| Tree ensembles | Ensemble size | Table 9 |
| Step-ZSDT | Leaf-stopping threshold | 0.9 |
| XGBoost | learning_rate | 0.1 |
| LogReg | max_iter | 1000 |
| Linear SVM | Estimator | LinearSVC |
| Parameter | Value |
|---|---|
| Number of Trials per Setting | 10 |
| Random Seed | 0 |
| Train Batch Size | 8 |
| Test Batch Size | 8 |
| Dataset | Random Forest | CatBoost | LLMT |
|---|---|---|---|
| Nursery | 0.397 | 0.473 | 0.762 |
| Diabetes | 0.584 | 0.590 | 0.764 |
| Spambase | 0.725 | 0.743 | 0.811 |
| Abalone | 0.688 | 0.700 | 0.694 |
| Blood | 0.549 | 0.558 | 0.675 |
| Iris | 0.769 | 0.829 | 0.940 |
| Dataset | #Shots | Random Forest | CatBoost | LLMT |
|---|---|---|---|---|
| Nursery | 3 | 0.349 | 0.464 | 0.762 |
| 6 | 0.397 | 0.473 | 0.762 | |
| 12 | 0.707 | 0.411 | 0.784 | |
| 24 | 0.754 | 0.721 | 0.818 | |
| 48 | 0.798 | 0.740 | 0.840 | |
| Diabetes | 2 | 0.486 | 0.523 | 0.662 |
| Dataset | TabPFN | TabLLM | LLMT | LLMT Forest |
|---|---|---|---|---|
| Nursery | 0.750 | 0.759 | 0.762 | – |
| Diabetes | 0.710 | 0.510 | 0.764 | – |
| Spambase | 0.802 | 0.673 | 0.811 | – |
| Abalone | 0.754 | 0.569 | 0.694 | – |
| Blood | 0.658 | 0.496 | 0.675 | – |
| Iris | 0.918 | 0.327 | 0.940 | – |
| Dataset | LogReg | SVM | CART | Direct-ZSDT | IO-Tree | CoT-Tree | ToT-Tree | FeatLLM | GPTree | LLMT (Ours) |
|---|---|---|---|---|---|---|---|---|---|---|
| Communities | 0.424 | 0.439 | 0.404 | 0.418 | 0.336 | 0.411 | 0.402 | 0.432 | 0.407 | 0.438 |
| Ecom | 0.509 | 0.508 | 0.554 | 0.561 | 0.511 | 0.524 | 0.534 | 0.579 | 0.490 | 0.582 |
| Myocardial | 0.526 | 0.522 | 0.540 | 0.510 | 0.535 | 0.503 | 0.498 | 0.596 | 0.556 | 0.618 |
| Average | 0.486 | 0.490 | 0.499 | 0.496 | 0.461 | 0.479 | 0.478 | 0.536 | 0.484 | 0.546 |
| Dataset | XGBoost | RandomForest | LightGBM | CatBoost | DeLTa | LLMT Forest (Ours) |
|---|---|---|---|---|---|---|
| Communities | 0.428 | 0.440 | 0.434 | 0.437 | 0.473 | 0.490 |
| Ecom | 0.695 | 0.504 | 0.412 | 0.503 | 0.619 | 0.757 |
| Myocardial | 0.702 | 0.545 | 0.581 | 0.535 | 0.565 | 0.773 |
| Average | 0.608 | 0.496 | 0.476 | 0.492 | 0.552 | 0.673 |
| Dataset | Diabetes | Iris | Spambase | Nursery | Abalone | Blood | Breast | Glioma |
|---|---|---|---|---|---|---|---|---|
| #Rules ( ) | 10 | 10 | 15 | 10 | 10 | 10 | 10 | 15 |
| #Features | 8 | 4 | 18 | 8 | 8 | 4 | 9 | 23 |
| Avg. Features Used | 8.0 | 4.0 | 15.0 | 8.0 | 8.0 | 4.0 | 9.0 | 14.9 |
| Feature Coverage | 100% | 100% | 83.3% | 100% | 100% | 100% | 100% | 64.8% |
| Datasets | #Shots | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Nursery | 3 | 0.740 0.000 | 0.762 0.100 | 0.762 0.100 | 0.762 0.100 |
| 6 | 0.740 0.000 | 0.762 0.100 | 0.707 0.063 | 0.731 0.077 | |
| 12 | 0.740 0.000 | 0.784 0.091 | 0.742 0.074 | 0.737 0.054 | |
| 24 | 0.740 0.000 | 0.818 0.066 | 0.713 0.084 | 0.735 0.045 | |
| 48 | 0.740 0.000 | 0.840 0.000 | 0.788 0.042 | 0.742 0.039 | |
| Diabetes | 2 | 0.662 0.095 | 0.662 0.095 | 0.662 0.095 | 0.669 0.083 |