Organizations: School of Computational Science and Engineering, Georgia Institute of Technology, 756 West Peachtree Street Northwest, Atlanta, 30308, GA, USA
Emergency managers need to know where floodwater is, how deep it is, and how it will change over the coming hours across an entire river basin. During a flood, however, real-time measurements come from only a handful of stream gauges, and high-resolution hydrodynamic models are too costly to rerun each time new data arrive or to run as large ensembles. We present C-STRIDE, an observation-driven AI digital twin that turns short records from a few stream gauges, together with terrain and rainfall, into basin-wide maps of water depth and extends these predictions up to a day ahead. It is trained on simulations from a calibrated two-dimensional hydrodynamic model and needs no separate data-assimilation step. In the Des Plaines River basin near Chicago, six gauges inform predictions over 4.2 million 30-m grid cells. Terrain improves the predictions most, rainfall keeps errors from growing over longer horizons, and together they reduce errors by about 40% compared with gauge records alone. When future rainfall is known, errors remain near 15% one day ahead, compared with nearly 40% without rainfall. Given real instead of simulated gauge records, the model shifts its predictions toward the observed hydrographs at three of six gauges without retraining, and it runs about 150 times faster than the hydrodynamic model. These results show how sparse gauges, terrain, and rainfall can be combined into fast, continuously updated flood predictions, a step toward operational flood digital twins that still requires testing with real-time data and rainfall forecasts.
Figures & tables
Approach
Run-time inputs
Output
Spatial representation
DTE Hydrology ( Brocca et al., 2024 )
Earth observations, hydrological forcing
Water-cycle states and scenarios
Gridded
Pipedream ( Bartos and Kerkez, 2021 )
Hydraulic model, sensor data
Drainage-network states and forecasts
1-D network
Inundation data assimilation ( Neal et al., 2007 ; Hostache et al., 2018 )
Hydraulic model, gauge or satellite data
Updated water levels and extent
Model grid
U-FLOOD ( Löwe et al., 2021 )
Rainfall, terrain
Event-maximum depth
Raster
SWE-GNN ( Bentivoglio et al., 2023 )
Current hydraulic state, terrain
Time-evolving hydraulic state
Mesh graph
SHRED, Senseiver ( Williams et al., 2024 ; Santos et al., 2023 )
Sparse sensor histories or values
Full field
Grid or continuous queries
Table 1: Representative hydrological digital-twin, data-assimilation, flood-surrogate, and sparse-sensing approaches, summarized by the information they require at run time, the output they produce, and how they represent space. Entries describe modeling interfaces rather than a common accuracy benchmark.
Property
Value
Domain shape
Non-rectangular ( NaN -masked)
Grid size
5,075×1,661
Active in-domain cells
4,188,840
Spatial resolution
30m
DEM source
USGS 3DEP
Manning coefficient
0.02 (water) / 0.05 (land)
Table 2: Des Plaines River basin dataset and sparse-observation configuration.
Figure 1: Overview of the C-STRIDE architecture. Sparse observation histories yk−K:k and precipitation histories pk−K:k are encoded by G into a latent state zk , which conditions the decoder F . The decoder is queried at spatial coordinates ξ together with terrain feature vectors ϕ(ξ) , comprising elevation bz(ξ) , slope bg(ξ) , and Manning coefficient n~(ξ) , to produce the next-step depth h~(tk+1,ξ) . The example panels illustrate sensor observations, precipitation snapshots, query locations, terrain features, and a predicted water-depth field.
Component
Setting
Observation window K+1
12 snapshots
Encoder G
Two-layer LSTM, 256 hidden units
Auxiliary model G′
Two-layer LSTM, 128 hidden units; linear 128→6 head
Decoder F
8 FMMNN blocks, width 1024 , rank 128
Decoder inputs
2 coordinates and 3 terrain features
Latent modulation
Affine 256→896 shift projection
Table 3: Architecture and training settings of the full C-STRIDE encoder–decoder and the separately trained gauge forecasters. The four encoder–decoder configurations share the same architecture and optimization settings.
Model
Gauge history
Past rainfall
Future rainfall
Terrain
Vanilla STRIDE
✓
–
–
–
C-STRIDE (terrain)
✓
–
–
✓
C-STRIDE (forcing)
✓
✓
✓
–
C-STRIDE (full)
✓
✓
✓
✓
CLDNet
–
✓
n/a
✓
CLDNet + LD-EnSF
✓
–
n/a
✓
Table 4: Information available to each model at run time. Future rainfall is used for forecasting. n/a: not evaluated.
Model
April 2013 flood
Four-event average
εh (%)
RMSEh (m)
εh (%)
RMSEh (m)
Vanilla STRIDE
13.69
0.0961
18.14
0.0995
C-STRIDE (terrain)
9.49
0.0667
14.23
0.0811
C-STRIDE (forcing)
13.59
0.0960
15.73
0.0879
C-STRIDE (full)
8.38
0.0589
10.63
0.0631
CLDNet
13.46
0.0930
16.13
0.0867
Table 5: Mean snapshot depth errors for the April 2013 flood and the four-event average.
Figure 2: Effect of terrain and forcing conditioning on depth prediction. Each curve shows the mean snapshot-wise relative L2 error for one STRIDE configuration. Shading denotes one standard deviation across events.
Figure 3: April 2013 depth hydrographs at diagnostic locations A–H. Left panels show peak-depth locations (circles), and right panels show high-total-variation locations (squares), selected along alternating northing rows. Curves compare SynxFlow, full C-STRIDE, terrain-only C-STRIDE, and vanilla STRIDE. The center panel shows the DEM, diagnostic locations, and gauges 1–6 (red triangles).
Figure 4: Spatial prediction of the April 2013 flood at peak total reference depth. Panels show SynxFlow depth (left), full C-STRIDE depth (center), and prediction minus reference (right) over the evaluation domain.
NSE ↑
KGE ↑
εhpeak (%) ↓
Model
Mean
Median
Mean
Median
Mean
Median
Vanilla STRIDE
0.949
0.937
0.898
0.876
3.10
3.06
C-STRIDE (terrain)
0.976
0.984
0.936
0.952
2.75
1.66
C-STRIDE (forcing)
0.931
0.955
0.881
0.880
4.90
0.93
C-STRIDE (full)
0.951
0.960
0.915
0.934
6.33
3.29
Table 6: Gauge hydrograph fidelity relative to SynxFlow during the April 2013 flood. Mean and median are over six gauges, while bold values identify the best configuration.
Figure 5: Gauge-scale comparison during the April 2013 flood using simulated histories as C-STRIDE inputs. The left panel shows the reference depth and the six gauge locations. Hydrographs compare mean-aligned WSE from full C-STRIDE (blue dashed), SynxFlow (blue solid), and USGS (black solid).
Figure 6: Response of full C-STRIDE to replacing simulated gauge histories with USGS observations at inference time, without retraining. Dashed curves show predictions with simulated (blue) or USGS (black) inputs. Solid curves show SynxFlow (blue) and USGS (black). All series use the same retrospective mean alignment, allowing comparison of hydrograph timing and shape.
SynxFlow
C-STRIDE (sim. input)
C-STRIDE (USGS input)
Gauge
NSE
Δtpeak (h)
NSE
Δtpeak (h)
NSE
Δtpeak (h)
Russell (05527800)
0.883
+6
0.780
+7
0.731
+3
Gurnee (05528000)
0.893
−25
0.944
+2
0.932
+2
Des Plaines (05529000)
0.889
−10
0.880
−11
0.970
0
Salt Creek (05531500)
0.163
+4
0.162
+6
0.603
+4
Riverside (05532500)
0.182
+1
0.118
0
0.840
0
Table 7: Agreement with USGS mean-aligned WSE. Positive peak-time errors indicate late predictions. The first maximum defines peak time. At Gurnee, the observed crest is nearly flat, so peak-time errors there are sensitive to small fluctuations. The last two rows use gauges withheld from encoder inputs.
Figure 7: Gauge hydrographs from the four-mainstem-input model. Panels marked “Input” supply histories to the encoder, while Salt Creek and DuPage, marked “Held out”, do not supply inputs. Simulated- and USGS-input predictions are compared with SynxFlow and observations using the line styles in Fig. 6 .
Model
CSI ↑
Precision ↑
Recall ↑
Bias ( →1 )
Vanilla STRIDE
0.888
0.947
0.935
0.987
C-STRIDE (terrain)
0.925
0.959
0.963
1.004
C-STRIDE (forcing)
0.894
0.951
0.938
0.986
C-STRIDE (full)
0.933
0.963
0.967
1.004
CLDNet
0.871
0.926
0.936
1.011
CLDNet + LD-EnSF
0.843
0.901
0.929
1.032
Table 8: Flood-extent metrics for the April 2013 flood at a 0.5 m depth threshold. Frequency bias above one indicates over-prediction.
Model
m=1
m=4
m=6
m=12
m=24
Vanilla STRIDE
17.26±8.72
18.15±9.12
21.05±11.35
27.69±13.42
38.14±12.63
C-STRIDE (terrain)
11.72±9.75
13.49±9.84
17.42±12.30
26.88±14.81
38.74±13.68
C-STRIDE (forcing)
14.70±4.29
15.33±4.53
15.76±4.52
16.86±5.07
18.28±4.26
C-STRIDE (full)
8.06±5.09
8.72±4.97
9.38±4.75
11.75±5.88
14.48±4.67
Full, rainfall mismatch α=1
8.06±5.09
8.74±5.01
9.44±4.85
12.05±6.16
14.72±4.63
Full-field persistence
2.73±3.85
9.56±10.43
13.02±12.92
21.74±16.21
36.37±17.79
Table 9: Mean ± standard deviation of relative depth error (%) at forecast horizons m∈{1,4,6,12,24} across 40 forecasts. Persistence uses the complete preceding reference field.
Figure 8: Mean rainfall intensity and mean absolute mismatch for α=1 over the forecast horizon. The solid curve shows true precipitation and the dashed curve shows the imposed mismatch, averaged equally over the 507 rainfall values and then over forecasts. Shading denotes one standard deviation of the mismatch across forecasts.
Figure 9: Relative depth error over a 24-step forecast. “True” and “Noisy” denote full C-STRIDE with observed and perturbed future rainfall, respectively. Terrain-only and vanilla receive no rainfall. Persistence holds the complete preceding reference field fixed. Curves show means over the forecasts, with shading denoting one standard deviation. Forecasts can overlap in time, so this spread does not represent uncertainty across independent training runs.
Figure 10: Effect of observation-history length K+1 on full C-STRIDE prediction. The models are compared at common prediction times, so each curve covers the same part of the flood evolution. Curves show mean snapshot relative L2 depth error, and shading denotes one standard deviation across events.
Configuration
εh (%)
RMSEh (m)
Full, six gauges
10.63
0.0631
50% tributary dropout
11.15
0.0665
Four mainstem gauges
10.67
0.0636
Noise level 0.05
11.77
0.0700
Noise level 0.10
12.17
0.0730
Noise level 0.20
12.43
0.0741
Table 10: Mean relative depth error and RMSE for the sensor configurations.
AI-driven flood digital twins demand fast hydrodynamic surrogates for ensemble forecasting and observation assimilation. Yet even GPU-accelerated two-dimensional shallow water equation (SWE) solvers still require ∼55 minutes per 96-hour run on a ∼4.2-million-active-cell metropolitan basin (the DesPlaines River basin at 30m resolution), making such workloads prohibitive at native resolution. We present the Conditional Latent Dynamics Network (CLDNet): a low-dimensional latent neural ODE driven by rainfall, paired with a coordinate-based decoder conditioned on static terrain (elevation, slope, Manning roughness) that reconstructs depth and discharge at arbitrary query points. Pointwise decoding decouples memory from grid size and handles irregular watersheds natively, enabling metropolitan-scale training on a single compute node and direct queries at exact gauge coordinates without raster snapping. We evaluate CLDNet on a synthetic 250,000-cell Texas benchmark and on a new DesPlaines case study of 114 real-rainfall StageIV storms whose reference simulator we validate against United States Geological Survey (USGS) gauges at the April2013 flood-of-record (Nash--Sutcliffe efficiency 0.57--0.94 on mean-recentered water-surface elevation). CLDNet roughly halves the relative root-mean-squared error of an unconditional baseline, outperforms regular-grid VAE--ConvLSTM and FNO baselines on the Texas benchmark (both presuppose a Cartesian grid and do not apply to the irregular Des~Plaines watershed), reaches a critical success index of ≈86% at the 0.5m inundation threshold, and produces a full 96-hour basin-wide forecast in ∼29 seconds -- a ∼115× speedup.
Phillip Si, Yuan Qiu, Omar Sallam +4
School of Computational Science and Engineering, Georgia Institute of Technology, 756 West Peachtree Street Northwest, Atlanta, 30308, GA, USA · Environmental Science Division, Argonne National Laboratory, 9700 S Cass Ave, Lemont, 60439, IL, USA
Near-real-time flood depth prediction demands surrogate models that are accurate, fast, and transferable across watersheds. Supervised surrogates can match physics-based simulators in accuracy but need millions of training rows per watershed and cannot extrapolate beyond their original mesh. We propose a domain-aware coreset construction pipeline that conditions a tabular foundation model at inference time. The pipeline stratifies storms by return period and most-affected watershed, then samples hexagons with a target-aware spatial selector. With 0.7% of the per-watershed training pool, the model attains a mean R2 of 0.663 across nine Houston-area watersheds, within 98.5% of the supervised reference (R2 = 0.673). It transfers to held-out watersheds without task-specific retraining, staying ahead of a coreset-trained supervised baseline. On real storms it exceeds the supervised reference on a far out-of-distribution case and trails it on a mostly in-distribution one. Domain-aware coreset construction lets tabular foundation models deliver data-efficient, watershed-transferable flood predictions without per-watershed training.
Lipai Huang, Adithi Srinath, Manas Singh +2
Urban Resilience.AI Lab, Zachry Department of Civil and Environmental Engineering, Texas A&M University, College Station, TX, USA · Department of Computer Science and Engineering, Texas A&M University, College Station, TX, USA · Resilitix Intelligence LLC, Houston, TX, USA +1
Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. A central challenge is to produce high-resolution forecasts across continental domains where hydrological behavior varies widely from place to place. We developed Hapi, a U-Net Swin Transformer that uses fine three-dimensional patches and hierarchical shifted-window attention to forecast river discharge, surface runoff, snow water equivalent, and soil wetness index across the contiguous United States. The model produces medium-range forecasts (24--72~h) at 0.05∘ resolution with adaptive task weighting and required only 0.11 seconds for a four-variable 72-h CONUS forecast on one A100 GPU. In a held-out 2024 potential-skill evaluation with ERA5-Land inputs prescribed over the forecast horizon, Hapi achieved the highest F1-score for floods in 20 of 21 comparisons across seven GloFAS return periods and three forecast leads. Independent validation against observed daily discharge at 3{,}881 U.S. Geological Survey gauges showed that Hapi achieved the highest median Nash--Sutcliffe efficiency at every lead, supported by regional-cluster bootstrap intervals. In a matched 24-h comparison of loss formulations, adaptive task balancing produced the lowest discharge errors and the highest F1-score for floods.
Hong Zhang, John K. Hutchison, Rao Kotamarthi +7
Argonne National Laboratory, Lemont, IL 60439, USA