Efficient Continuous DEM Reconstruction under Limited Target-Resolution Supervision
Authors: Zekai Shi, Meng Zhang, Haokun Zhang, Bo Zhang
Organizations: School of Human Settlements and Civil Engineering, Xi’an Jiaotong University, Xi’an, 710049, China · School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University, Xi’an, 710072, China
High-resolution digital elevation models (DEMs) support Earth observation applications, but paired training references are often available only at coarser output resolutions. Reconstructing finer terrain grids therefore requires both effective transfer beyond the supervised scale and control of dense-query computation. To address this problem, SCOPE learns a continuous terrain representation from coarser-resolution pairs. It predicts a latent coefficient field on the low-resolution grid and reuses local Fourier residual functions through basis evaluation and geometry-guided ensemble fusion. This separates high-dimensional coefficient prediction from output-grid construction. Experiments on geographically distributed land--ocean samples assess supervised reconstruction, unseen-scale inference, cross-domain generalization, and theoretical computation. SCOPE leads the compared methods across six metrics in the main supervised-scale evaluation. At an unseen factor three times the training factor, land reconstruction reduces RMSE and MAE by approximately 12% relative to bicubic interpolation, with errors close to target-scale fine-tuning. Ninefold output density increases counted multiply--accumulate operations by only about 2%. Frozen-model validation on held-out external marine regions reduces RMSE relative to the DEM-specific implicit baseline EBCF-CDEM by approximately 19% under self-downsampling and 2% with cross-product inputs, while also yielding lower RMSE than LIIF-MS in both settings. These results demonstrate the value of reusable coefficient fields for accurate reconstruction beyond the supervised resolution with low incremental arithmetic cost.
Figures & tables
Figure 1: Schematic architectural trends in scale-dependent computational burden from 1× to 30× for fixed-grid reconstruction, coordinate-wise INR decoding, and SCOPE. SCOPE concentrates coefficient prediction on the LR grid and uses lightweight local evaluation for denser output. Solid curves extend through 15× ; dashed extensions illustrate conceptual trends to 30× . Markers identify the supervised 5× scale and the maximum accuracy-evaluation factor of 15× . Quantitative GMACs and computation–latency trajectories are reported in Sec. 5 .
Figure 2: Conceptual distinction among representative continuous reconstruction methods and SCOPE. The comparison highlights the predicted object, supervision setting, and decoding mechanism: query-wise implicit value prediction, query-centered frequency-aware reconstruction, DEM-specific elevation-bias prediction, and reusable latent coefficient-field evaluation.
Product
Coverage
Grid spacing
Role in this study
TanDEM-X EDEM
Global land
30 m
Land HR reference
GEBCO_2024
Global ocean and land
∼ 450 m
Ocean background
NOAA CRMs (A–C, E)
U.S. coastal zones
∼ 30 m
Land–sea transition reference
Torres Strait 2023 (F)
139–146 ∘ E, 8–13 ∘ S
30 m
Reef and island reference
Bass Strait 2022 (G)
143–149 ∘ E, 38–41 ∘ S
30 m
Shelf bathymetry reference
EMODnet DTM 2024 (D)
Caribbean tile
∼ 100 m
Fixed- 5× evaluation only
Table 1: DEM and bathymetric source products and their experimental roles. Parenthetical region identifiers refer to Fig. 3 . The first six products support the main study; GEBCO_2023 and the four regional products below it support frozen-model external evaluation. All external reference rasters are prepared on a nominal 90 m evaluation grid. The external subset contains 12 reference rasters and 3,825 valid patches; Great Barrier Reef blocks B–D are used, with block A excluded for spatial overlap with the main-study region.
Figure 3: Geographic distribution and coverage context of the source products. Brown and purple outlines distinguish main-study and external-evaluation product extents, respectively. A, B, C, and E identify NOAA CRMs Vols. 10, 7, 1–5, and 9; D identifies the EMODnet DTM 2024 Caribbean tile; F and G identify Torres Strait 2023 and Bass Strait 2022. H 1 /H 2 , I, J, and K identify CHS NONNA-100, MH370 bathymetry, Northern Australia, and Great Barrier Reef sources, respectively. The evaluated subsets are defined in Table 1 and Sec. 2.3 . Red triangles locate representative regions, insets enlarge selected areas, and gray shading marks latitudes outside the 65∘ N– 65∘ S study limits.
Figure 4: Architecture of SCOPE network. (a) Overall framework. The LR DEM is encoded into an LR-aligned terrain feature field and transformed by NAF into a latent coefficient field. For a continuous query coordinate q , coefficients sampled at neighboring LR locations xi are evaluated using the relative offsets Δqi to produce local residual candidates ri(q) . These candidates are fused by LAE into r(q) and added to the bicubic base surface B(q) . (b) LAE module. Four neighboring residual candidates are fused using geometry-guided compatibility weights.
Category
Variant
RMSE ↓
MAE ↓
Slope ↓
Aspect ↓
PSNR ↑
Corr. ↑
Full model
SCOPE network ∗
4.150
2.816
12.051
46.473
77.48
0.659
Base & Fusion
No base
4.682
3.238
14.718
57.076
73.88
0.521
Bilinear base
4.153
2.817
12.051
46.571
77.43
0.659
Nearest base
4.176
2.830
12.223
47.227
77.19
0.651
Bilinear fusion
4.186
2.830
12.285
47.013
77.16
0.652
Refinement & Loss
No post-refinement
4.885
3.372
17.468
63.935
73.64
0.436
Table 2: Ablation study of the SCOPE network under the fixed 5× reconstruction setting. Slope and Aspect denote slope-based and aspect-based reconstruction errors, respectively. Corr. denotes the slope-consistency score between reconstructed and reference terrain structures.
Category
Method
Params (M)
GMACs
RMSE ↓
MAE ↓
Slope ↓
Aspect ↓
PSNR ↑
Corr. ↑
Interpolation
Bicubic
0
–
4.733
3.205
13.304
47.041
76.16
0.640
Bilinear
0
–
5.551
3.747
14.172
49.389
74.81
0.618
Nearest
0
–
8.059
5.535
36.380
80.548
71.06
0.216
Fixed-scale SR
SRCNN
0.479
19.150
7.844
3.890
13.858
50.629
67.75
0.576
EDSR
21.582
34.577
4.417
2.973
12.981
49.803
75.36
0.612
SwinIR
17.555
41.848
4.456
3.067
13.095
50.560
74.82
0.606
Table 3: Quantitative comparison with interpolation, fixed-scale super-resolution, and implicit neural reconstruction baselines under the fixed 5× setting. All baseline architectures follow their official implementations or recommended model configurations and are trained and evaluated under the unified data and evaluation protocol. Params and GMACs are reported for batch size 1 with a 40×40 LR input and a 200×200 output, counting convolution, linear projection, attention matrix multiplication, and explicit basis-evaluation matrix multiplication. Red and blue indicate the best and second-best results for each metric, respectively. Lower RMSE, MAE, Slope, and Aspect indicate better performance, while higher PSNR and Corr. indicate better performance.
Figure 5: Fixed-scale visual comparison on a representative land patch under the 5× reconstruction setting. Panels (a)–(g) show the LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network, respectively, while panel (h) shows the GT HR DEM. All panels are compared over the same elevation range. The enlarged regions highlight local terrain details, where the SCOPE network better preserves ridge–valley structures and fine-scale relief variations.
Figure 6: Fixed-scale visual comparison on a representative marine and coastal patch under the 5× reconstruction setting. Panels (a)–(g) show the LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network, respectively, while panel (h) shows the GT HR DEM. All panels are shown with a shared elevation range. The enlarged regions illustrate the reconstruction of coastal and bathymetric transition structures, where the SCOPE network produces a sharper and more coherent terrain surface.
Figure 7: Absolute elevation-error comparison on the representative land patch under the fixed 5× reconstruction setting. Panels (a)–(g) show the absolute elevation errors of the LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network, respectively, with respect to the GT HR DEM shown in panel (h). The enlarged regions highlight error distributions around local ridge–valley structures.
Figure 8: Absolute elevation-error comparison on the representative marine and coastal patch under the fixed 5× reconstruction setting. Panels (a)–(g) show the absolute elevation errors of the LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network, respectively, with respect to the GT HR DEM shown in panel (h). The enlarged regions emphasize errors near coastal and bathymetric transition zones.
Figure 9: Elevation profile comparison on representative land and ocean patches under the fixed 5× reconstruction setting. For each case, the HR reference and SCOPE network reconstruction are shown with the same A–B profile line, while the full and zoomed elevation profiles compare HR, LR, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network. The zoomed profiles provide a detailed view of local peak–valley reconstruction fidelity.
Figure 10: Aspect-based directional terrain comparison on the representative land patch under the fixed 5× reconstruction setting. The HR reference, LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network are visualized with the same cyclic aspect color mapping. The enlarged regions highlight local ridge–valley directional structures, where the SCOPE network better preserves fine-scale terrain orientation and directional continuity.
Figure 11: Aspect-based directional terrain comparison on the representative marine and coastal patch under the fixed 5× reconstruction setting. The HR reference, LR input, bicubic interpolation, EDSR, SwinIR, LIIF, EBCF-CDEM, and the SCOPE network are visualized with the same cyclic aspect color mapping. The enlarged regions emphasize coastal and bathymetric directional transitions, where the SCOPE network produces aspect patterns that are more consistent with the HR terrain structure.
Method
RMSE ↓
MAE ↓
Slope ↓
Aspect ↓
PSNR ↑
Corr. ↑
Bicubic
8.452
6.145
15.230
57.081
69.53
0.537
Bilinear
8.989
6.544
15.931
57.565
69.17
0.528
Nearest
10.774
7.733
37.668
83.901
67.16
0.172
SCOPE-Direct
8.164
5.880
15.184
60.453
69.05
0.507
SCOPE-Transfer
8.546
6.223
14.963
61.152
68.78
0.496
Table 4: Matched-product supervision and transfer in the existing LR–HR paired experiment, distinct from the frozen external tests in Table 5 . SCOPE-Direct is trained directly on matched paired data, while SCOPE-Transfer is trained with synthetic downsampling and evaluated on the paired data. Red and blue indicate the best and second-best results for each metric, respectively. Lower RMSE, MAE, Slope, and Aspect indicate better performance, while higher PSNR and Corr. indicate better performance.
Figure 12: Cross-dataset visual comparison on a representative land patch under the matched LR–HR paired setting. The HR reference, LR input, bicubic interpolation, and SCOPE network reconstruction are shown together with local error and aspect insets. Here, the SCOPE network corresponds to the directly trained SCOPE-Direct variant reported in Table 4 .
Figure 13: Cross-dataset visual comparison on a representative marine and coastal patch under the matched LR–HR paired setting. The HR reference, LR input, bicubic interpolation, and SCOPE network reconstruction are shown together with local error and aspect insets. Here, the SCOPE network corresponds to the directly trained SCOPE-Direct variant reported in Table 4 .
Input setting
Method
RMSE ↓
MAE ↓
Slope ↓
Aspect ↓
PSNR ↑
Corr. ↑
Self-downsampled
Bicubic
3.7834
2.3876
15.0834
51.3582
77.4655
0.6045
LIIF-MS
73.6168
73.3150
15.7968
53.3109
44.5991
0.5898
EBCF-CDEM
4.2196
2.8100
20.5145
71.4522
73.6952
0.3776
EDSR
3.5019
2.2343
14.1920
52.4866
77.5683
0.5953
SCOPE-5F
3.4009
2.1734
13.9307
51.0133
78.3732
0.6241
Matched GEBCO_2023
Bicubic
9.0587
5.7358
17.6567
68.0274
68.6858
0.4574
Table 5: Frozen-model external marine evaluation at 5× (nominal 450→90 m), using the same 12 reference rasters and 3,825 valid patches per method in both input settings. All metrics are averaged over patches. Slope, Aspect, and Corr. follow the definitions used throughout the paper and correspond to RMSE-Slope, RMSE-Aspect, and SlopeCorr in the evaluation logs. Red and blue indicate the best and second-best distinct values within each setting, including ties at the displayed precision.
Method
Params (M)
GMACs
RMSE ↓
MAE ↓
Slope ↓
Aspect ↓
PSNR ↑
Corr. ↑
Bicubic
0
–
7.657
5.212
21.468
71.386
69.57
0.447
LIIF-MS
15.375
946.533
8.620
5.909
22.293
73.185
68.61
0.426
EBCF-CDEM
1.450
84.799
10.265
7.215
29.331
96.334
65.01
0.182
SCOPE-5F
15.400
26.208
6.753
4.609
19.824
71.534
70.45
0.455
SCOPE-5MS
15.400
26.208
6.769
4.613
19.915
71.667
70.43
0.454
SCOPE-15D
15.400
26.208
6.789
4.629
19.839
72.028
70.37
0.446
Table 6: Quantitative comparison at the target 15× reconstruction scale. LIIF-MS is trained with 5× as the primary scale and random downsampling across {2×,3×,4×} for multi-scale supervision. SCOPE-5F denotes fixed 5× training only, SCOPE-5MS denotes multiscale training up to 5× , SCOPE-15D denotes direct 15× training, and SCOPE-15FT denotes 5× training followed by 15× fine-tuning. Params and GMACs are reported for batch size 1 with a 40×40 LR input and a 600×600 output, counting convolution, linear projection, attention matrix multiplication, and explicit basis-evaluation matrix multiplication. Red and blue indicate the best and second-best results for each metric, respectively. Lower RMSE, MAE, Slope, and Aspect indicate better performance, while higher PSNR and Corr. indicate better performance.
Figure 14: Scale-transfer behavior of different SCOPE network training protocols from 2× to 15× . The top row reports absolute RMSE, MAE, and slope error, while the bottom row reports metric differences computed as each compared protocol minus SCOPE-5F at the same scale. SCOPE-5F denotes fixed 5× training, SCOPE-5MS denotes multiscale training, SCOPE-15D denotes direct 15× training, and SCOPE-15FT denotes 5× training followed by 15× fine-tuning.
Fusion
RMSE
MAE
Slope
Aspect
MLP-based LAE
4.150
2.816
12.051
46.473
Transformer-style fusion
4.186
2.832
12.302
47.061
Table 7: Comparison between MLP-based LAE and Transformer-style local fusion. Lower values indicate better performance.
Figure 15: Computation–latency trajectories across reconstruction scales. Both axes are logarithmic; lower-left positions indicate lower cost. Marker sizes encode 2× , 3× , 5× , 10× , 15× , 20× , and 30× ; lines connect increasing factors within each model. Dashed lines partition the cost space.
Photogrammetry and Remote Sensing, ETH Zürich, Zürich, Switzerland · Remote Sensing Technology Institute, German Aerospace Center (DLR), Wessling, Germany