Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.
Figures & tables
Figure 1: Meta-learning workflow for generating look-ahead trees using synthetic data. The workflow can be divided into two parts: (1) meta-learning step where labeled synthetic datasets are fed into MetaTree along with the optimal decision trees for each dataset as the training targets, and (2) inference step where the pre-trained model is used to predict the look-ahead trees on an unseen, real-world datasets
Figure 2: Workflow for pre-training with synthetic data
Figure 3: Time complexity as a function of: (a) the depth of the tree, and (b) the number of binary features for generating pre-training targets
# of trees
MetaTree Original
MetaTree Synthetic Data
CART
GOSDT
1
0.6508 (0.0068)
0.6443 (0.0070)
0.6502 (0.0072)
0.6524 (0.0072)
5
0.6783 (0.0063)
0.6755 (0.0064)
0.6814 (0.0063)
0.6670 (0.0069)
10
0.6769 (0.0061)
0.6707 (0.0063)
0.6806 (0.0061)
0.6646 (0.0069)
30
0.7047 (0.0059)
0.6956 (0.0061)
0.7053 (0.0060)
0.6943 (0.0066)
Table 1: Average accuracy and standard error of the mean (in parentheses)
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: MetaTree scaling laws as a function of: (a) dataset size, (b) label noise, and (c) model complexity
Figure 5: Class imbalance distribution comparison between MetaTree benchmarks and the synthetic datasets generated using the proposed workflow. The synthetic dataset display more uniform class distribution between the selected range [0,0.3] compared to hand-curated MetaTree benchmarks. Class count distribution comparison between MetaTree benchmarks and the synthetic datasets generated using the proposed workflow. Compared to MetaTree benchmarks, which were manually hand-picked and hence display irregular distribution, the class count distribution in synthetic datasets can be well-explained by the quality filters we enforced. Both accuracy and class imbalance filters will favor datasets with smaller number of classes.
Conventional decision tree algorithms produce effective, transparent models that can be audited, communicated, and deployed independently of the training data, but require learning every new task from scratch. In contrast, tabular foundation models demonstrate that meta-learning from a synthetic prior distribution enables strong in-context prediction for previously unseen tasks, especially in small-sample regimes. However, this approach does not produce a standalone model that can be inspected in isolation. We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass. MotherTree is pre-trained on a synthetic prior using stochastic gradient descent without requiring reference trees for supervision. On established benchmarks with controlled sample size, the approach is competitive with size-matched trees from common algorithms: recursive partitioning, gradient-based tree learning, globally optimal trees, and distillation from tabular foundation models. Notably, MotherTree consistently improves over from-scratch gradient-based learning and acts as a strong initializer: task-specific tuning of the generated tree outperforms the corresponding from-scratch learner on all benchmarks and sample sizes. These results show that meta-learning can provide effective inductive biases for learning stand-alone, small decision tree classifiers.
Ziyuan Wang, Fredrik D. Johansson
Chalmers University of Technology University of Gothenburg
Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains. For the remaining settings, however, complex ensembled compositions of trees often achieve higher accuracy at the cost of interpretability, leaving practitioners with difficult modeling decisions along an accuracy-interpretability tradeoff. Ideally, we would like to classify as much of the data as possible with one or a small number of trees, achieving interpretability for most samples while maintaining state-of-the-art accuracy. We introduce Multistage Defer Trees: a sequence of sparse decision trees that each make predictions for most samples, while deferring a small proportion to the next tree in the sequence or, ultimately, to a black box. We demonstrate that we can train this model class to match the performance of complex tree-based ensembles while routing most samples through only one or a small number of sparse decision trees. We discuss a range of techniques for training these models while maintaining simplicity. Our method expands the accuracy--interpretability frontier in settings where single-tree methods remain insufficient, demonstrating that even when complex models are necessary, they need not be fully opaque.
Zakk Heile, Hayden McTavish, Margo Seltzer +1
Department of Computer Science · Duke University · Durham, USA +2
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.
Yue Qiu, Zekang Du, Yiqun Diao +2
School of Computer Science and Technology, Huazhong University of Science and Technology · School of Computing, National University of Singapore