Adversarial patches for aerial perception are typically evaluated as digital composites, with printing left as an implementation detail. This study imposes three physical constraints during optimization rather than after it: the color range a particular printer can reproduce, the size of the flat panel a vehicle offers, and the loss of fine detail incurred when the patch is imaged from altitude. The principal comparison isolates the ink set. Two patches share all seventeen recorded optimization settings and differ only in the colors available to them. One is constrained to a uniform color cube; the other to a gamut measured by printing and scanning a 216-patch chart. Each was optimized at three seeds and evaluated against thirteen victim conditions, with every rate reported against a size-matched optimized control. The effect of the measured gamut is victim-dependent rather than uniform. Net attack success rises on three of six closed-set segmentation victims, and for these the seed ranges of the two ink sets are disjoint: +0.120 on DeepLabv3-R101 and +0.041 on SegFormer-B0. The color-cube patch is consistently stronger on the open-vocabulary segmenter and on two of four detectors, though no detector exceeds a net of +0.026 under either ink set. The natural explanation is that a printable palette is simply less chromatic and lower in frequency than a digital one. Eleven further patches test this account and it does not hold. Once cardinality is matched, a palette as chromatic as the cube attacks equally well. Cardinality itself shows no trend from three inks to thirty-two. Palettes matched on cardinality, lightness and chroma, and differing only in hue placement, span 0.035 to 0.136. A physical evaluation with printed decals did not detect transfer; it bounds the transferred rate at 0.133, which does not exclude the simulated value of 0.121.
Figures & tables
Term
wj
Role
Link ( Equation 7 )
0.10
soft distance to nearest printable ink
Lflat ( Equation 6 )
0.05
anchored anisotropic total variation
Lanchor ( Equation 9 )
0.20
hinged distance to the reference design
Lgamut ( Equation 8 )
0.05
penalty for leaving the measured gamut
Lnps
0.02
reported for comparability only
Table 1: Realism term weights wj in Equation 2 . These shape a feasible iterate; they do not create one, since the projection Π already enforces the constraint set.
Level
Constraint set
0
Unconstrained pixels
1
Band limit
2
Band limit + anchored total variation
3
Band limit + anchored total variation + palette
4
Band limit + anchored total variation + palette + anchor design
Table 2: Nested realism constraint sets. Each level adds one constraint to the preceding level.
Figure 1 : Patches produced at each realism level at matched patch size and optimisation budget, with the deployed patch at right. Annotations give the constraint level, the number of colours used out of the palette offered, the high-frequency energy ratio, and, where an anchor design exists, the fraction of its edges retained.
Figure 2 : Seven independent draws from the print-chain augmentation applied to the reference cargo design. Each draw samples substrate reflectance, ink response and illumination.
Figure 3 : The controlled pair. Both patches share all seventeen recorded optimisation settings and differ only in the ink set: a placeholder colour cube (left) and the measured printer gamut (right). Median ΔE00 from patch colours to the nearest reproducible colour falls from 8.0 to 0.0 .
Group
Quantity
Value
Corpus
Frames
12562
Annotated vehicle instances
21767
primary / secondary
12562 / 9205
class car / truck
14662 / 7105
Target vehicles
10
Labelling
Fully labelled / ignored
18315 / 3452
Table 3: GeoPatchCity summary and the subset evaluated in this study. Scale statistics are over primary targets, for which the viewpoint grid is centred. Split sizes are read from the distributed split files.
Victim
Family
Reference
Access
Closed-set semantic segmentation
FCN-R50
CNN
[ 53 , 56 ]
transfer
DeepLabv3-R101
CNN
[ 54 , 56 ]
transfer
UPerNet-ConvNeXt-T
CNN
[ 55 , 41 ]
transfer
UPerNet-Swin-T
transformer
[ 55 , 40 ]
transfer
SegFormer-B0
transformer
[ 39 ]
white box
Table 4: The victim roster: eleven models, giving thirteen victim conditions once the three CLIPSeg prompts are counted separately. White-box access is used for SegFormer-B0 and RetinaNet during optimisation of the compared patches; all other victims are evaluated by transfer only. The two UPerNet models are additionally attacked white box in Section 5.6 .
Group
Setting
Cube run
Measured run
Objective
Task
joint
joint
Segmentation victim
SegFormer-B0
SegFormer-B0
Detection victim
RetinaNet
RetinaNet
Realism level
4
4
Anchor design
cargo
cargo
Print chain
enabled
enabled
Table 5: Every recorded optimisation setting for the controlled pair. The two runs agree on all seventeen axes and differ only in the ink set (final row).
Cube ink set
Measured ink set
Arm
Victim
ASR
ctrl
net
ASR
ctrl
net
Seg.
DeepLabv3-R101
0.667
0.173
+0.494
0.768
0.173
+0.596
FCN-R50
0.187
0.012
+0.175
0.193
0.012
+0.181†
SegFormer-B0
0.101
0.001
+0.100
0.127
0.001
+0.127
UPerNet-ConvNeXt-T
0.046
0.010
+0.036
0.058
0.010
+0.048†
UPerNet-Swin-T
0.037
0.003
+0.034
0.039
0.003
+0.036†
Table 6: Net attack success rate, net=ASR−ASRcontrol ( Equation 13 ), on a single 2534-instance holdout. Both patches share all seventeen recorded optimisation settings ( Table 5 ) and differ only in the ink set. Each is scored against its own size-matched control. Bold marks the higher net per victim in this single run. A dagger marks a victim whose three-seed ranges overlap in Table 7 , where every victim in this table was replicated; those rows are not resolved and should not be read as an ordering, whichever way the single run fell. The daggers are therefore measured rather than extrapolated. Note that on UPerNet-Swin-T the sign reverses under replication, so the bold entry in that row records this run and not a claim. The CLIPSeg prompt “a vehicle seen from above” is omitted because its attackable subset is empty under both patches, making ASR undefined by Equation 12 .
Figure 4 : Semantic segmentation under each ink set, SegFormer-B0. Columns correspond to the same instance under the clean frame, the cube-constrained patch, and the measured-gamut patch; rows alternate input image and predicted vehicle mask. IoU and its change relative to the clean frame are given beneath each prediction.
Cube ink set
Measured ink set
Arm
Victim
mean
range
mean
range
Δ
∣Δ∣/s
Sep.
Closed-set seg.
DeepLabv3-R101
0.655
0.647–0.667
0.775
0.768–0.785
+0.120
12.6
✓
FCN-R50
0.187
0.180–0.195
0.213
0.193–0.228
+0.025
1.9
SegFormer-B0
0.106
0.098–0.118
0.147
0.127–0.172
+0.041
2.3
✓
UPerNet-ConvNeXt-T
0.050
0.045–0.060
0.060
0.058–0.064
+0.010
1.6
UPerNet-Swin-T
0.037
0.034–0.042
0.035
0.030–0.039
−0.003
0.7
Table 7: Seed replication across the full victim roster: three independent optimisation runs per ink set, all other settings fixed. Means are over the three seeds and the range is their minimum and maximum. Δ is the measured-gamut mean minus the cube mean, so a positive value favours the measured gamut. “Sep.” marks victims whose three-seed ranges do not overlap; where they overlap the ordering is not resolved at this number of seeds and is not claimed. Controls are seed-independent and are those of Table 6 .
Figure 5 : The palette ablation of Table 8 , on one axis. The grey band on each panel is the measured-gamut run plus or minus the between-seed standard deviation of 0.0178 for this victim, so a point inside it is not separable from run-to-run variation. (a) Beyond the measured level, chroma does not matter: the two palettes bracketing the cube ink set’s own mean chroma of 54.6 sit inside the band. (b) Cardinality shows no trend across a tenfold range. (c) Three palettes matched to the measured set on cardinality, lightness and chroma, differing only in hue placement, span 0.035 to 0.136 . Swatches beneath each point are the palettes themselves.
Experiment
Ink set
k
C∗
ASR
net
δ/s
Chroma varied, k=8 throughout
half chroma
8
16.4
0.103
0.102
-2.7
measured gamut ∗
8
32.8
0.151
0.150
—
1.5 × chroma
8
47.7
0.146
0.145
-0.3
2 × chroma
8
58.6
0.142
0.141
-0.5
Cardinality varied, C∗≈33 throughout
3 inks
3
31.9
0.110
0.109
-2.3
5 inks
5
31.7
0.119
0.118
-1.8
Table 8: Testing hypothesis H1 on SegFormer-B0. Every run matches the measured-gamut patch of Table 5 in all seventeen recorded settings except the ink set. C∗ is the palette’s mean CIELAB chroma after conversion to sRGB, k its cardinality. The measured gamut appears in the first two blocks as the shared reference point. δ is the difference from it, expressed in units of the between-seed standard deviation of 0.0178 measured for this victim in Table 7 ; a value below about 2 is not separable from run-to-run variation.
Figure 6 : Attack success rate against realism constraint level at matched patch size, optimisation budget, supercell, victim and instance set. The subtitle records the victim and the attackable count, both held fixed.
Figure 7 : Grad-CAM attributions for the vehicle channel, SegFormer-B0, computed at the final 256-channel feature map. Annotations give the fraction of class-activation mass falling inside the decal footprint.
Figure 8 : Semantic segmentation under each ink set, SegFormer-B0, on three instances where the attack succeeded. Columns are the clean frame, the colour-cube patch, and the measured-gamut patch; rows alternate input image and predicted vehicle map, annotated with intersection over union against ground truth and its drop from the clean prediction. Instances were drawn evenly across the nadir angle range from the 181 successes of the cube patch within a pool of 1800 attackable instances. This is a selected sample, included to show the failure mode rather than its frequency.
Figure 9 : Per-instance segmentation detail for the measured-gamut patch on SegFormer-B0. Columns give (a) the clean frame with ground truth outlined, (b) the clean prediction, (c) the patched frame with the decal footprint outlined, (d) the prediction under attack, and (e) the flipped pixels separated into those under the decal and those away from it.
Attack term
ASR
IoU drop
Under
Away
Posterior hinge, Equation 3
0.147
0.047
11.3%
1.9%
Log-odds hinge
0.136
0.037
10.3%
1.9%
Log-odds hinge, per-region ink logits
0.139
0.040
10.5%
2.0%
Log-odds hinge, away from footprint only
0.019
−0.031
3.0%
1.9%
Untargeted, whole frame
0.000
−0.041
0.0%
0.2%
Posterior hinge, level 0, no print chain
0.171
0.080
15.6%
2.2%
Table 9: Alternative segmentation attack terms, SegFormer-B0, level 4 unless stated, 2.0m decal, 1000 optimisation steps, one seed each. Scored on a stride-5 subset of the holdout (507 instances, 368 attackable), so rates are comparable within this table but not with Table 6 . “Under” and “away” are the shares of vehicle pixels removed beneath and outside the decal footprint.
Patch
ASR
IoU drop
Under
Away
Universal, printable (deployed)
0.417
0.195
25.2%
3.0%
Universal, free pixels
0.583
0.296
34.1%
6.5%
Per-frame, printable
1.000
0.567
48.0%
26.2%
Per-frame, free pixels
1.000
0.755
48.4%
47.0%
Table 10: The objective of Equation 3 on one frame at a time and across the fleet, SegFormer-B0, 24 attackable nadir holdout instances, all 2.0m footprints. Per-frame patches are optimised against the instance they are scored on and are an upper bound, not an attack.
Figure 10 : Segmentation of the same holdout instances under universal and per-frame patches, SegFormer-B0. Universal patches were optimised once over 855 training instances; per-frame patches against the frame shown alone, with the same objective, footprint and placement and without EOT or the print chain. Rows alternate input, with the decal footprint outlined, and predicted vehicle map; “away” is the share of vehicle pixels removed outside the footprint. The instances are evenly spaced over the 24 of Table 10 , which were chosen without reference to any attack outcome.
Victim
Run
Attackable
ASR
Control
net
Nadir ASR
UPerNet-ConvNeXt-T
original
2080
0.125
0.010
+0.115
0.164
fixed
2080
0.162
0.010
+0.152
0.184
UPerNet-Swin-T
original
1879
0.005
0.003
+0.002
0.000
fixed
1879
0.059
0.003
+0.056
0.092
Table 11: White-box attacks on the UPerNet pair, full 2534-instance holdout, one run each. “Original” uses Equation 3 and the projected parametrisation; “fixed” uses the log-odds hinge and per-region ink logits. The control is the unoptimised reference decal at the same size.
Figure 11 : Attack loss during white-box optimisation against the UPerNet pair, in 100-step means. The original and fixed runs optimise attack terms on different scales, so each curve is indexed to its own first block; values below 1 are descending. The original Swin-T run diverges.
Figure 12 : The repaired white-box attack on UPerNet-Swin-T. Columns as in Figure 9 . The three instances are successes, drawn evenly across the viewing-angle range from 1879 attackable instances; this is a selected sample, and the rate is in Table 11 .
Figure 13 : Object detection under each ink set, RetinaNet, on three instances where the attack succeeded. Ground-truth boxes are dotted and detections above threshold solid; the best vehicle score and its drop from the clean frame are given beneath each panel. Instances were drawn evenly across the nadir angle range from the 121 successes of the cube patch within a pool of 2286 attackable instances. This is a selected sample, included to show the failure mode rather than its frequency.
Figure 14 : Object detection under each ink set, RetinaNet. The ground-truth box is shown dotted and detections above threshold solid. Confidence and its change relative to the clean frame are given beneath each panel.
Figure 15 : Per-instance detection detail for the measured-gamut patch on RetinaNet. Columns give (a) the clean frame, (b) the frame under attack, and (c) the decal footprint. Ground-truth boxes are dotted and detections above threshold solid; the note beneath each attacked frame states whether any box survived.
Figure 16 : Attack success rate over nadir angle and altitude for the measured-gamut patch. Each cell gives ASR and the number of attackable instances; cells without an attackable instance are marked undefined.
Figure 17 : Attack success against nadir angle for the measured-gamut patch across the full victim roster. The shaded band marks the near-nadir range the patch was optimised at.
Nadir angle
att.
ASR
Altitude
att.
ASR
Illumination
att.
ASR
0\SIUnitSymbolDegree10\SIUnitSymbolDegree
357
0.289
19m30m
68
0.691
morning
367
0.174
10\SIUnitSymbolDegree20\SIUnitSymbolDegree
217
0.226
30m40m
55
0.182
midday
1036
0.118
20\SIUnitSymbolDegree30\SIUnitSymbolDegree
192
0.172
40m50m
60
0.067
evening
397
0.108
30\SIUnitSymbolDegree40\SIUnitSymbolDegree
394
0.074
50m63m
55
0.036
40\SIUnitSymbolDegree50\SIUnitSymbolDegree
328
0.040
50\SIUnitSymbolDegree60\SIUnitSymbolDegree
228
0.009
Table 12: Transfer envelope for the measured-gamut patch on SegFormer-B0. Altitude bins are restricted to near-nadir frames so that viewing angle and altitude are not confounded; “att.” is the number of attackable instances.
Figure 18 : Attack success rate by illumination condition for the measured-gamut patch on SegFormer-B0, with vehicle-level confidence intervals.
Figure 19 : Attack success against each geometric axis for the measured-gamut patch on SegFormer-B0, with the unoptimised size-matched control and vehicle-level confidence bands. Attackable counts are printed per bin.
Figure 20 : Attack success against each geometric axis for the measured-gamut patch on RetinaNet, with the size-matched control and vehicle-level confidence bands. The two curves remain close at every bin.
Figure 21 : Physical captures with RetinaNet detections overlaid. Clean and patched frames are taken from the same pose seconds apart. Best vehicle confidence and the number of retained boxes are given beneath each panel.
Deep neural network (DNN)-based object detectors are widely used for analyzing aerial and satellite imagery in applications such as environmental monitoring and urban analytics. Despite their strong performance, these models are known to be vulnerable to adversarial examples, and physical adversarial attacks using printable patterns pose realistic security threats. In this paper, we evaluate physical adversarial patch attacks against an aerial vehicle detector by bridging digital optimization and real-world deployment. Adversarial patches are optimized in the digital domain using a loss function that minimizes the maximum objectness score while incorporating non-printability score (NPS) and total variation (TV) constraints to ensure both printability and spatial smoothness. The optimized patches are printed and deployed in three configurations: ON, OFF, and OFF-Side. Experiments using a YOLOv3 detector show that while the OFF patch achieves the highest effectiveness in the digital domain (85.51% Average Objectness Reduction Rate (AORR)), the ON patch demonstrates superior robustness in physical environments (0.197-0.343 Objectness Score Ratio (OSR)) due to its consistent visibility. Furthermore, our results indicate that weather-based augmentation does not necessarily improve patch optimization in this domain. These findings provide critical insights into the practical vulnerabilities of aerial object detection systems.
Jung Heum Woo, Eun-Kyu Lee
School of Information Technology Incheon National University Incheon 22012, Republic of Korea
Pixel-wise adversarial patches are computationally heavy and often visually detectable, limiting utility in security-critical systems. We present adversarial Voronoi camouflage that optimizes only seed-point locations under fixed, printable palettes using a soft assignment, producing structured, splinter camouflage-like patterns without additional regularization. Evaluated on person detection with COCO-style AP@[.5:.95], naive placement (Inria -> COCO) performs comparably bad, while garment-level application via segmentation mask (3DPeople) results in a significant AP drop. The attack transfers to out-of-domain backgrounds and across detector families (YOLOv9/10/11/12), indicating robustness in black-box settings. Repainting with different palettes largely nullifies the effect, and single-color tweaks show limited tolerance (<=0.17), highlighting a structure-palette coupling. The parameter-efficient, palette-constrained design improves visual plausibility while degrading real-time detector performance. Physical validation and color calibration are left for future work. Code: https://github.com/JensBayer/Voronoi This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026.
Jens Bayer, Stefan Becker, David Münch +2
Fraunhofer IOSB and Fraunhofer Center for Machine Learning · Karlsruhe Institute of Technology
Although deep neural network-based remote sensing object detectors have achieved strong performance, they remain vulnerable to adversarial perturbations. Existing studies mainly focus on digital or white-box settings, whereas black-box physical attacks remain underexplored. These attacks are often constrained by limited physical feasibility and inefficient optimization in high-dimensional search spaces. To address these challenges, this paper proposes ColorFD, a black-box physical attack based on multiple pure-color patches. The patch positions and color parameters are jointly optimized using Differential Evolution (DE). A target-wise fitness and selection mechanism evaluates the attack state of each target and preserves target-specific improvements during evolution. Two guidance strategies further constrain the patch search space. Key-region localization identifies sensitive regions through finite-difference color probing. Common-feature extraction provides category-level spatial priors and avoids repeated localization. Although evaluated on aircraft, the formulation is not inherently restricted to this category. Experiments on YOLOv3u, YOLOv5u, and Faster R-CNN show that ColorFD outperforms the tested black-box patch method across all evaluated detectors and remains competitive with strong white-box baselines. Physical-world experiments further demonstrate that the optimized pure-color patches can be transferred from the digital domain to real imaging conditions.
Tiannuo Guo, Guhang Qiu, Yuzhen Xie +3
College of Information Science and Technology, Beijing University of Chemical Technology, Beijing 100029, China