Bayesian additive regression trees (BART) can require many splits to approximate boundaries misaligned with the predictor axes. RoBART assigns each tree a rotation shared by all internal nodes, retaining axis-aligned splits in rotated coordinates and constant leaves. We jointly propose a Givens rotation sequence and cutpoints on the resulting grid by Metropolis-Hastings and establish reversibility with respect to the conditional posterior with leaf means integrated out. For additive functions with component-specific rotations and anisotropic Hölder smoothness, we prove posterior contraction in empirical L2 distance and for the noise standard deviation. Under the stated prior, design, and grid conditions, with fixed numbers of predictors, trees, and components and no more components than trees, the rate is a sum of componentwise rates determined by smoothness and the number of rotated coordinates used. We also establish a posterior contraction lower bound showing that there exist functions for which RoBART adapts to the intrinsic dimension but axis-aligned BART does not.
Figures & tables
Figure 1: True functions and posterior means for (a) the X-shaped signal and (b) a Gaussian spike. Both use n=800 uniform inputs on [−1,1]2 and independent N(0,0.12) errors. All methods use 200 trees. Color scales are shared within rows but differ between rows.
Figure 2: Two-tree illustrations of BART (left) and RoBART (right). Within each half, the two small panels show individual tree contributions, and the larger panel shows their sum. Numbers are illustrative predictions in the original predictor coordinates. Blue and orange identify the two tree partitions.
Figure 3: RMSPE (a) and elapsed time in seconds (b) at n∈{1,500,5,000} . Rows give Scenarios 1–4 and columns give (p,d) . Boxes summarize thirty replications; RMSPE is measured against the true regression function.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Operation
Cost
Copy the rotated training design
O(np)
Apply the Givens path to design columns and frame rows
O{Krot(n+p)}
Recompute touched extrema and materialize their grid columns
O{J(n+b)}
Route observations, collect reference values, select order statistics, and evaluate collapsed likelihoods
O{n(1+D)+N}
Enumerate local cutpoint normalizers and evaluate or draw indices
O(Ib)
Evaluate discrete-prior factors by inspecting labels and ancestors
O{pN(1+D)}
Appendix
Table A.1: Operation bounds for one joint proposal under the stated storage convention.
Learning high-quality oblique decision trees remains a significant challenge due to the discrete and non-convex nature of split optimization. We present the Hinge Regression Tree (HRT) framework, which reframes each oblique split as a nonlinear least-squares problem over two linear predictors whose max/min envelope induces ReLU-like representation capacity. We show that the resulting node-level optimization can be interpreted as a damped Newton method, and we establish the monotonic decrease of the node objective for its backtracking line-search variant. We establish, theoretically, that HRT is a universal approximator with an explicit O(δ2) approximation rate. Building upon this base learner, we propose HRT-Boost, a mathematically synergistic ensemble extension that couples node-level Newton updates with stage-wise functional gradient descent. We show that this ensemble construction admits a stage-wise empirical risk reduction guarantee under the squared loss. Empirical evaluations on synthetic and real-world benchmarks show that HRT is highly competitive with established single-tree baselines, and HRT-Boost compares favorably with strong ensemble baselines and often yields substantially more compact models. The code is publicly available at https://github.com/Hongyi-Li-sz/HRT-Boost.
Hongyi Li, Jun Xu, Hong Yan
School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, 518055, China and with the Shenzhen Key Lab for Advanced Motion Control and Modern Automation Equipments, Shenzhen, 518055, China · Department of Electrical Engineering, City University of Hong Kong, Kowloon, Hong Kong
Bayes-assisted conformal prediction combines the strengths of Bayesian modelling with exact, distribution-free frequentist coverage guarantees. Although conformal validity is preserved even when the Bayesian working model (BWM) is misspecified, the size of the resulting prediction sets can degrade substantially when the prior is poorly aligned with the observed data. We address this limitation by introducing RoBAS (Robust Bayes-Assisted Shrinkage): a Bayes-assisted framework for constructing robust nonconformity scores, with two instantiations: one induced by a heavy-tailed BWM, and a closed-form empirical Bayes shrinkage score. The resulting scores adapt to the quality of the working information encoded in the prior: when this information is reliable, they exploit it to produce efficient prediction sets; when it is weak or inaccurate, they revert to the Distance-To-Average (DTA) score, a robust non-informative baseline. We evaluate the proposed scores on tabular and image regression tasks where the training distribution may differ from the calibration and test distributions, while the calibration and test data themselves remain exchangeable. We find that they are competitive with widely used scores in the absence of such shift, while substantially reducing interval widths in shifted settings.
Kianoosh Ashouritaklimi, Stefano Cortinovis, François Caron
Reliably quantifying predictive uncertainty is difficult for complex, high-dimensional, or misspecified models. Both fully Bayesian and bootstrap resampling methods provide principled uncertainty estimates but are often too expensive for modern machine-learning models because they require posterior sampling or repeated model refitting. We introduce Ribbon, a scalable approximation to Dirichlet-reweighted bootstrap uncertainty. Ribbon replaces repeated refitting with an influence-function linearization around a single fitted model, preserving the first-order data-reweighting structure of the Bayesian bootstrap while requiring only post-hoc linear algebra. Ribbon approximates the Bayesian-bootstrap or weighted-likelihood-bootstrap refitting target. With a general concentration parameter, Ribbon gives a calibrated Dirichlet-reweighting family whose uncertainty scale can be tuned on validation data. We show that Ribbon is asymptotically equivalent to a flat-prior Laplace approximation under correct likelihood specification and recovers the robust sandwich covariance under misspecification. Across synthetic regression, MNIST classification, and California Housing benchmarks, Ribbon provides competitive predictive performance and improved calibration in several settings while avoiding repeated model retraining.