Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal building damage assessment pipeline built on a siamese detector derived from YOLOX, designed to compress information at both ends of the ground/space link. On the ground, pre-disaster reference images are encoded into a compact latent space -- compressed by up to a factor of 64 -- and uplinked to the satellite. On board, this reference is compared with a fresh post-disaster acquisition so that the downlink carries only actionable object-level products, bounding boxes and damage classes, instead of full scenes. This cuts the data exchanged in both directions, while on xBD the strongly compressed reference still preserves most of the detection performance. Because on-board acquisitions suffer from residual pre/post co-registration errors, we introduce a latent-space shift estimation and correction module that regresses the global offset from the coarse feature level and realigns the post-disaster features before fusion. It substantially improves robustness to de-registration -- especially under large shifts, where fusion-only variants collapse -- while also raising nominal accuracy and remaining compatible with the strongest compression. We finally port the pipeline to two embedded targets, a Xilinx Versal VCK190 and an NVIDIA Jetson AGX Orin, and report hardware performance (latency, throughput, power efficiency). The core detector and its compression port cleanly to both, but the operators needed for long-range robustness survive only on the Jetson GPU, whereas the Versal DPU does not.
Figures & tables
Figure 1: System concept. The ground segment encodes pre-disaster references into a compressed latent ( ×64 ) uplinked to the satellite; the on-board segment detects and co-registers against the fresh post-disaster image and downlinks only low-rate products (bounding boxes, damage level, optional crops over damaged regions) instead of full scenes.
Configuration
D3
D4
D5
Size
Input image ( 10242×3 )
–
–
–
3.15 MB
Latent (no comp.)
128
256
512
3.67 MB
Latent (comp. r=8 )
16
32
64
0.46 MB
Latent (comp. r=64 )
2
4
8
57 kB
Latent (comp. r=128 )
1
2
4
29 kB
Table 1: Size of the pre-disaster latent representation under channel compression, for a 1024×1024×3 input, reported as an equivalent int8 footprint. Spatial resolutions: 1282 (D3), 642 (D4), 322 (D5).
Figure 2: Proposed ground/on-board architecture (details in Sec. 3.3 ). Compressed pre-disaster latents are uplinked and expanded on board; the offset-estimation module regresses a global shift (δx,δy) from the D3 latent, which the latent offset alignment applies to realign the post D3/D4/D5 features before multi-scale fusion.
Configuration
P
R
F1
mAP@0.5
Early Fusion (baseline)
58.3
55.0
56.3
55.1
Siamese (no shift-aug)
59.0
58.0
58.2
56.4
Siamese
58.5
57.4
57.8
56.6
Siamese + Cross-Attention
61.1
56.0
58.4
56.5
Siamese + Portable Attn (Sp.-Diff)
58.9
57.3
57.9
56.7
Siamese + Shift Est. & Corr.
61.5
57.3
59.2
57.8
Table 2: Nominal detection performance on xBD (no test-time shift), %. All variants use shift augmentation unless stated.
Figure 3: Robustness to pre/post misregistration on xBD: mAP@0.5 vs. applied post-disaster shift (log- x ), averaged over shift directions. The vertical line marks the 150 px training range of the shift augmentation; the proposed module (green) stays flat beyond it, whereas shift-augmentation and attention baselines degrade.
Figure 4: Offset estimation error (mean ± std) vs. applied shift. Compression ×64 matches the uncompressed estimator; ×128 (a single D3 latent channel) roughly doubles the error and breaks alignment. The rise beyond ∼150 px reflects extrapolation outside the training range.
Platform
Model
float32
Quant.
Δ
Jetson (TensorRT)
Siamese
57.90
57.59 (int8)
−0.31
Jetson (TensorRT)
S.E.C. + C × 64
57.63
54.39 (FP16)
−3.24∗
VCK190 (Vitis AI)
Siamese + C × 128
52.66
53.98 (int8)
+1.32
VCK190 (Vitis AI)
Portable Sp.-Diff
58.85
58.03 (int8)
−0.81
Table 3: Impact of quantisation/deployment on porting, per model (mAP@0.5, %). Rows compare each model’s float32 reference against its on-target quantised version, not models across platforms.
Figure 5: Jetson AGX Orin sweep. Top: latency (stacked bars: preprocess / inference / postprocess, left axis) and throughput (line, right axis). Bottom: power (mean bar with a min–max whisker, left axis) and energy efficiency (line, right axis). Five axes: model, precision, NMS mode, batch size, image size.
Figure 6: VCK190 sweep (int8, batch 6), two axes (model, image size). From left to right: latency/throughput by model, latency/throughput by image size, power/efficiency by model, power/efficiency by image size. The DPU batch size is fixed and the NMS/precision knobs of the Jetson path do not apply.
Platform
Model
GFLOPs
Lat.
Thr.
Power
Eff.
(/img)
(ms)
(Mpx/s)
(W)
(GF/s/W)
Jetson Orin
Siamese
69.8
25.6
41.0
12.5
288
VCK190
Siamese
69.8
74.1
85.0
29.8
172
Jetson Orin
S.E.C. + C × 64 (best)
70.4
38.8
27.1
12.0
181
VCK190
Siamese + C × 128 (best)
69.9
67.9
92.7
32.3
171
Table 4: Hardware performance at 10242 . Only noatt (Siamese) is compared across platforms; per-platform best models use different compression and are not cross-compared. Jetson: FP16, batch 1, 30 W; VCK190: int8, batch 6.