Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
Organizations: Capital One
Abstract
Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.
Figures & tables
| # of trees | MetaTree Original | MetaTree Synthetic Data | CART | GOSDT |
|---|---|---|---|---|
| 1 | 0.6508 (0.0068) | 0.6443 (0.0070) | 0.6502 (0.0072) | 0.6524 (0.0072) |
| 5 | 0.6783 (0.0063) | 0.6755 (0.0064) | 0.6814 (0.0063) | 0.6670 (0.0069) |
| 10 | 0.6769 (0.0061) | 0.6707 (0.0063) | 0.6806 (0.0061) | 0.6646 (0.0069) |
| 30 | 0.7047 (0.0059) | 0.6956 (0.0061) | 0.7053 (0.0060) | 0.6943 (0.0066) |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.