Understanding how vegetation loss alters rainfall remains a major challenge in climate and hydrological science, as deforestation modifies precipitation through heterogeneous, seasonal and nonlinear land-atmosphere feedbacks. Existing models struggle to capture these dynamics: convection is parameterised at coarse scales, tipping behaviour is poorly constrained, and rainfall-deforestation analyses are limited to multi-decadal timescales. Therefore, many approaches resolve correlations rather than causal effects, limiting our ability to anticipate hydrological disruption. Using a neural-network model for hourly rainfall prediction, combined with pathway diagnostics and sensitivity analyses, we examine how vegetation perturbations reorganise rainfall across space, intensity regimes, and timescales under deforestation. We assess whether the model captures physically consistent dependencies linking vegetation, atmospheric state, and precipitation, and whether sustained canopy loss induces threshold behaviour. The model accurately predicts rainfall occurrence and intensity (Spearman = 0.84, F1 = 0.93, ROC-AUC = 0.98) and learns temporally ordered dependencies aligned with ecohydrological theory. Sensitivity analyses reveal rapid, asymmetric responses to vegetation loss: heavy rainfall (20-50 mm/h) declines by up to 7% under sustained deforestation, while light rainfall (0.1-1 mm/h) increases by 4%. Rainfall entropy rises by 1.3%, and dry-season intensity increases by 0.3-0.5% per 0.5% forest-cover loss, with strongest impacts in the north-western Amazon and Andean foothills. Threshold analysis reveals a sharp decline in precipitating area fraction after 2-3 months of sustained vegetation change in sensitive regions. These results demonstrate that data-driven approaches uncover process-relevant land-atmosphere coupling and highlight growing hydrological vulnerability in the Amazon.
Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas. However, identifying reliable machine learning models is complicated by overlapping debris-flow and non-debris-flow events in feature space, the need for model interpretability, and limited training data. This paper addresses these challenges through a systematic evaluation of machine learning models in terms of predictive performance, feature importance, and synthetic data augmentation. Using basin-scale observations of post-wildfire debris-flow events across the western United States, we compare 15 models, including the Tabular Prior-Data Fitted Network (TabPFN). Repeated stratified cross-validation shows that TabPFN achieves the highest unaugmented performance with a threat score of 0.637, closely followed by the best tree-based models. SHapley Additive exPlanations (SHAP) are used to identify the features driving predictions, revealing that short-duration rainfall intensity and storm accumulation consistently rank highest, while burn severity and terrain features contribute less. We further evaluate synthetic data augmentation using TabPFN-generated samples to address the scarcity of debris-flow observations. Synthetic augmentation improves the performance of all models except CNN, with the largest mean threat score increase of +0.041 among the deep learning models. By combining rigorous model benchmarking, interpretable feature analysis, and synthetic data augmentation, this work provides a comprehensive framework for improving post-wildfire debris-flow prediction.
Accurate, spatially explicit characterization of tropical forest structure is essential for carbon accounting and ecosystem monitoring, yet most ML pipelines predict canopy-top height proxies (e.g., RH95/RH98) or AGBD as separate scalar targets, rather than learning the forest vertical structure as an ordered profile. The community lacks a ML-ready multimodal benchmark for predicting the entire GEDI RH profile jointly with AGBD, or for evaluating methods that enforce physically consistent ordering across RH percentiles. We address this with Biomazon, a 20 m multimodal benchmark dataset over the Amazon Basin that pairs GEDI RH and AGBD targets with multi-sensor predictors (Sentinel-1/2, ALOS-2 PALSAR-2, Copernicus DEM, Dynamic World LULC, and AlphaEarth embeddings) under standardized spatial splits and evaluation protocols. Using a shared encoder-decoder with task-specific heads as a baseline framework, we conduct a comprehensive ablation study of (i) backbone/model scale, (ii) modality contributions, and (iii) the use of auxiliary embeddings under standalone and fusion settings, and we report both single-target and joint-target results to quantify tradeoffs under a unified training protocol. Finally, we contextualize baseline performance through regionally aligned comparisons against existing gridded products, including GEDI L4D RH10-RH98 and AGBD, at matching temporal scale. Biomazon, together with the accompanying protocols and baseline results, establishes a reference benchmark for future work on structurally consistent RH-profile prediction and structure-biomass modeling in tropical forests.
Estimating temporal soil-loss change is challenging when physically meaningful input factors are noisy or corrupted, particularly because substantial changes are rare relative to the large number of locations exhibiting little change. We study this problem through the Revised Universal Soil Loss Equation (RUSLE) and introduce PhyRestore, a physics-structured latent-factor restoration framework. Rather than directly predicting soil-loss change or correcting a degraded physical estimate, PhyRestore restores corrupted physical factors and reconstructs temporal change through the known physical relationship. We evaluate PhyRestore in a watershed-scale bitemporal raster setting under isolated and simultaneous corruption of rainfall erosivity and cover management, comparing it with the degraded RUSLE estimate and Direct RF, XGBoost, MLP, and CNN models. Factor restoration improves high-magnitude recovery when the corrupted factors remain identifiable, but its advantage weakens under joint corruption, sparse positive extremes, and factor values outside the training support.