Gradient-Boosted Decision Trees

Also known as GBDT

Momentum

2 papers in the last four weeks, down 33% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 35

Apr 21, 2026physics.ao-ph

Improvements to the post-processing of weather forecasts using machine learning and feature selection

This study aims to develop and improve machine learning-based post-processing models for precipitation, temperature, and wind speed predictions using the Mesoscale Model (MSM) dataset provided by the Japan Meteorological Agency (JMA) for 18 locations across Japan, including plains, mountainous regions, and islands. By incorporating meteorological variables from grid points surrounding the target locations as input features and applying feature selection based on correlation analysis, we found that, in our experimental setting, the LightGBM-based models achieved lower RMSE than the specific neural-network baselines tested in this study, including a reproduced CNN baseline, and also generally achieved lower RMSE than both the raw MSM forecasts and the JMA post-processing product, MSM Guidance (MSMG), across many locations and forecast lead times. Because precipitation has a highly skewed distribution with many zero cases, we additionally examined Tweedie-based loss functions and event-weighted training strategies for precipitation forecasting. These improved event-oriented performance relative to the original LightGBM model, especially at higher rainfall thresholds, although the gains were site dependent and overall performance remained slightly below MSMG.
Feb 27, 2026cs.LG

Vectorized Dynamic Histograms for Sparse Oblique Forests

Sparse oblique (SPO), part of the top-ranked configuration of Google's Yggdrasil Decision Forests (YDF), improve the accuracy while maintaining interpretability of Random Forests (RF) and Gradient Boosted Trees (GBT) by scanning a sparse linear combination of features instead of a single feature. Because projections are sampled at runtime, training is significantly slower, as pre-run optimizations such as presorting cannot be used. We overhaul SPO training in YDF, first by fixing inefficiencies in the released version, achieving 2-5x speedup, then by devising two novel methods, one aimed at GBTs and the other at RFs: (1) Hierarchical AVX2 and AVX-512-vectorized histogram filling speeds up typical GBT-depth trees by up to 1.6x in SPO-GBT and 2x for SPO-RF, and (2) runtime-dynamic histograms between vectorized histograms and exact splits, speeds up deeper, RF depth-typical trees (depths 16 to purity) by up to 1.6x with no detectable effect on accuracy. Our optimizations bring SPO-RF training time on par with YDF's axis-aligned RF. We extensively evaluate on 8 natural and 11 synthetic datasets of up to 10.5 million rows and 1.6 million features, from depth 6 to purity, and open-source the implementation.
Feb 13, 2026cs.LG

Probabilistic Wind Power Forecasting with Tree-Based Machine Learning and Weather Ensembles

Accurate production forecasts are essential for the integration of renewable energy sources into the power grid. This paper illustrates how to obtain probabilistic forecasts of wind power generation using gradient boosting trees and an ensemble of weather forecasts. To this end, we perform a comparative analysis across three state-of-the-art probabilistic prediction methods-conformalized quantile regression, natural gradient boosting and conditional diffusion models-all of which can be combined with tree-based machine learning. The methods are validated using four years of data for all Belgian offshore wind farms. We benchmark the models against the power curve and a calibrated wake model as well as a probabilistic method using stochastic variational Gaussian process regression. The tree-based models significantly reduce the mean absolute error in comparison to the deterministic baselines. Additionally, all three methods outperform the Gaussian process baseline in probabilistic skill, while two out of the three also improve point forecast accuracy. The conditional diffusion model attains the best performance, with improvements of 5% in mean absolute error and 12% in continuous rank probability score compared to the probabilistic baseline. Last, the results indicate an average improvement in point forecast accuracy of 17% by using an ensemble of weather forecasts instead of a single provider.
Feb 5, 2026cs.LG

Selecting Hyperparameters for Tree-Boosting

Tree-boosting is a widely used machine learning technique for tabular data. However, its out-of-sample accuracy is critically dependent on multiple hyperparameters. In this article, we empirically compare several popular methods for hyperparameter optimization for tree-boosting including random grid search, the tree-structured Parzen estimator (TPE), Gaussian-process-based Bayesian optimization (GP-BO), Hyperband, the sequential model-based algorithm configuration (SMAC) method, and deterministic full grid search using 5959 regression and binary classification data sets. We find that the SMAC method clearly outperforms all the other considered methods on average, and it gives stable performance across a diverse collection of tabular data sets under a fixed tuning budget, which is relevant for users who cannot afford extensive manual trial-and-error tuning. We further observe that (i) a relatively large number of trials larger than 100100 is typically required for accurate tuning, (ii) using default values for hyperparameters or a full search over a small grid often yields very inaccurate models, (iii) all considered hyperparameters can have a material effect on the accuracy of tree-boosting, i.e., there is no small set of hyperparameters that is more important than others, and (iv) choosing the number of boosting iterations using early stopping yields more accurate results compared to including it in the search space for regression tasks.
Oct 31, 2025stat.ML

Gradient Boosted Mixed Models: Flexible Estimation of Mean and Variance Components for Clustered Data

We introduce Gradient Boosted Mixed Models (GBMixed), a framework which extends boosting to clustered data by jointly modeling the mean and variance components in a linear mixed model via likelihood-based gradients. GBMixed estimates a nonparametric fixed effects function characterizing the overall mean of the response, while also allowing the random effects covariance matrix along with the residual variance to depend on covariates in a flexible manner. We demonstrate how GBMixed facilitates covariate-dependent random effect predictions, and subsequently point predictions and prediction intervals for individual treatment effects, that can adapt between population-level and cluster-level information. Simulations and applications to two real-world datasets demonstrate that GBMixed can accurately recover complex nonlinear fixed effect functions and covariate-dependent covariances in a linear mixed model, while also improving point and probabilistic predictive performance compared with several existing approaches such as parametric linear mixed models, Natural Gradient Boosting, and Gaussian Process Boosting.