An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing the response of a model that preserves finer timing. We characterize this dependence as observer-relative stealth. For recorded event streams, retiming an event within its protected accumulation window leaves the accumulated integer tensor exactly unchanged. We use this exact blind space to construct Null, a gradient-guided timestamp-retiming attack, and define SC-ASR_A(tau) to measure attack success while bounding the change visible to observer A. On DVS Gesture at a 10% event budget, Null reaches 81.56 +/- 5.81% ASR on ConvSNN and 98.67 +/- 0.45% on a GRU while preserving the protected tensor exactly. On DailyDVS-200, a protocol-scale Multi-View Fusion Network variant reaches 99.28 +/- 0.11% exact-null ASR, compared with 9.70 +/- 1.06% for its matched control. In a five-attack comparison, Null is the only method with nonzero attack success at exact observer equality, reaching 81.4% on DVS Gesture and 87.35% on DailyDVS-200. We also search the same exact blind space with an independently implemented constrained projected-gradient optimizer, C-PGD. At matched victim-gradient evaluations, C-PGD reaches 84.50 +/- 2.89% ASR on DVS Gesture and 89.55 +/- 4.39% on DailyDVS-200, again with exact protected equality. Perturbations that are exactly hidden from the protected observer become visible under shifted, finer, overlapping, and randomized temporal views. Adding observer constraints reduces the real-valued blind-space fraction from 87.5% to 75.0% to 62.5%, while DVS ConvSNN ASR falls from 74.9% to 61.9% to 37.2%. These results show that stealth is not a property of the perturbation alone.
Figures & tables
Dataset
ConvSNN
SEW
Transformer
GRU
DVS Gesture
81.6 / 0.6
34.1 / 0.8
51.0 / 1.1
98.7 / 2.2
CIFAR10-DVS
93.4 / 8.6
65.0 / 2.1
86.0 / 5.1
99.0 / 22.0
N-Caltech101
44.5 / 1.7
36.2 / 1.3
70.7 / 2.3
89.5 / 8.8
DailyDVS-200
84.3 / 9.8
34.2 / 9.3
98.3 / 13.6
98.9 / 16.2
Table 1: Exact-null attack across four aligned event-vision datasets. Optimized ASR / random-control ASR (%). N-MNIST is reported separately below because its original and aligned protocols use different victim sets. DVS ConvSNN/SEW use displacement-conditioned random controls. DVS Transformer/GRU and the later dataset protocols use exact signed-displacement matching where reported. All optimized rows satisfy DA=D∞=0 .
Same attack, different observer. Relative representation deviation
Dataset
Canonical
Shift-1
Shift-2
Shift-4
Half-width
Random mean
DVS Gesture
0.000
0.081
0.116
0.137
0.144
0.103
DailyDVS-200
0.000
0.078
0.111
0.131
0.143
0.099
DVS observer intersections. Blind-space size and attack success
Observer constraints
Phase 0
Phases 0+2
Phases 0+2+4
Blind fraction (%)
87.5
75.0
62.5
Table 2: Changing the observer changes visibility and the common blind space. The top reports mean relative deviation of the same Null attacks under alternative observers over three seeds. The bottom reports DVS attacks re-optimized while enforcing equality to additional fixed phase observers.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Work
Perturbation
Count-preserving
Protected observation
Visibility / defense analysis
DVS-Attacks ( Marchisio et al., 2021 )
event-sequence perturbation
no constraint
–
sensor-filter study
Lee–Myung ( Lee and Myung, 2022 )
time shifts + added events
no
–
–
Yao et al. ( Yao et al., 2024 )
direct raw-event attack
no constraint
–
–
Du et al. ( Du et al., 2026 )
raw-event, configurable-latency attack
no exact count constraint
–
latency-robust optimization
Yu et al. ( Yu et al., 2026 )
timing-only retiming
yes
–
–
Temporal Poisoning ( Riaño et al., 2026 )
timestamp redistribution
yes
collapsed counts
detector study
Appendix
Table 3: Comparison with event-security work. “Protected observation” denotes a guarantee on the realized representation available to a stated observer. Collapsed counts are one specific observation. “Exact chosen R ” denotes the general invariance condition R(X′)=R(X) , which reduces to count equality when R is frame accumulation.
Victim
Attack
5%
10%
20%
30%
LIF
gradient retiming
8.66 ± 0.15
23.12 ± 0.98
54.81 ± 1.77
75.92 ± 1.71
LIF
random retiming
0.20 ± 0.10
0.17 ± 0.12
0.34 ± 0.21
0.55 ± 0.26
GRU
gradient retiming
66.08 ± 2.45
90.70 ± 1.64
98.44 ± 1.05
99.45 ± 0.26
GRU
random retiming
0.17 ± 0.12
0.24 ± 0.15
0.37 ± 0.06
0.41 ± 0.20
Appendix
Table 4: N-MNIST ASR (%) over three seeds.
Victim
Attack
2%
5%
10%
20%
ConvSNN
gradient retiming
24.98 ± 3.43
56.45 ± 5.26
81.56 ± 5.81
90.98 ± 1.27
ConvSNN
disp.-cond. random
0.14 ± 0.25
0.14 ± 0.24
0.56 ± 0.49
1.70 ± 0.75
SEW-ResNet
gradient retiming
10.44 ± 3.75
23.22 ± 3.74
34.06 ± 3.69
48.17 ± 1.40
SEW-ResNet
disp.-cond. random
0.27 ± 0.23
0.55 ± 0.23
0.82 ± 0.70
1.77 ± 0.92
Appendix
Table 5: DVS Gesture ASR (%) over three seeds. “Disp.-cond. random” samples attack-derived absolute shift magnitudes, then chooses random legal event identities/directions.
Victim
Attack
2%
5%
10%
20%
Event Transformer
gradient
16.00 ± 6.30
36.23 ± 6.22
50.96 ± 3.27
62.01 ± 3.21
Event Transformer
exact random
0.27 ± 0.23
0.81 ± 0.40
1.08 ± 0.45
2.01 ± 1.37
Temporal GRU
gradient
68.61 ± 6.07
96.63 ± 2.01
98.67 ± 0.45
99.11 ± 0.90
Temporal GRU
exact random
0.15 ± 0.26
1.18 ± 0.66
2.21 ± 0.84
6.46 ± 3.47
Appendix
Table 6: DVS Gesture ASR (%, mean ± std over 3 seeds) for the recurrent and attention-based victims. “Exact random” reproduces the adversarial signed-shift multiset sample by sample.
Victim
1 bin
2 bins
4 bins
full window
ConvSNN
34.73
51.88
69.87
75.73
SEW-ResNet
17.30
27.85
34.60
37.55
Event Transformer
57.96
56.73
59.59
53.47
Temporal GRU
98.70
98.27
99.13
98.70
Appendix
Table 7: DVS Gesture shift ablation, seed 0. ASR (%) under the main progressive path.
Victim
Clean acc.
5% ASR
5% random
10% ASR
10% random
ConvSNN
69.20 ± 0.39
82.50 ± 3.36
4.32 ± 0.92
93.42 ± 2.07
8.59 ± 1.66
SEW-ResNet
72.02 ± 0.33
46.23 ± 1.73
1.70 ± 0.10
64.98 ± 3.60
2.11 ± 0.11
Temporal GRU
55.70 ± 0.35
97.72 ± 0.47
11.67 ± 0.51
98.98 ± 0.22
22.04 ± 1.46
Event Transformer
69.83 ± 1.26
78.64 ± 0.28
3.71 ± 0.65
86.00 ± 0.39
5.10 ± 1.19
Appendix
Table 8: CIFAR10-DVS clean accuracy and ASR (%, mean ± std over three seeds). Random columns use the exact signed-displacement-matched control.
Victim
Attack
1%
2%
5%
10%
20%
ConvSNN
optimized
12.08 ± 0.63
19.81 ± 1.69
30.89 ± 2.09
44.48 ± 1.57
60.01 ± 1.28
ConvSNN
exact
0.42 ± 0.09
0.73 ± 0.24
1.25 ± 0.32
1.67 ± 0.60
1.57 ± 0.16
SEW-ResNet
optimized
7.45 ± 1.85
14.58 ± 1.03
26.89 ± 1.59
36.23 ± 3.39
46.71 ± 3.56
SEW-ResNet
exact
0.71 ± 0.62
0.52 ± 0.54
0.80 ± 0.54
1.27 ± 0.71
2.46 ± 1.08
Transformer
optimized
26.89 ± 2.26
45.96 ± 3.36
60.10 ± 2.73
70.66 ± 2.83
80.01 ± 2.18
Transformer
exact
0.77 ± 0.08
0.88 ± 0.09
1.70 ± 0.37
2.25 ± 0.35
5.33 ± 0.50
Appendix
Table 9: N-Caltech101 ASR (%, mean ± std over three seeds). “Exact” is the signed-displacement-matched random control.
Victim
Attack
1%
2%
5%
10%
20%
ConvSNN
optimized
34.39 ± 5.21
50.30 ± 8.38
70.37 ± 5.20
84.29 ± 6.23
94.89 ± 2.43
ConvSNN
exact
4.75 ± 1.60
5.75 ± 0.50
7.15 ± 1.31
9.81 ± 2.01
15.81 ± 5.14
SEW-ResNet
optimized
11.35 ± 0.99
17.46 ± 1.32
27.11 ± 2.50
34.23 ± 3.26
46.14 ± 1.83
SEW-ResNet
exact
3.64 ± 1.84
5.19 ± 1.34
5.62 ± 0.92
9.30 ± 1.81
11.27 ± 0.77
Transformer
optimized
65.24 ± 1.18
85.50 ± 1.07
94.34 ± 1.13
98.27 ± 0.96
99.38 ± 0.84
Transformer
exact
3.00 ± 1.04
5.12 ± 1.52
8.20 ± 1.35
13.65 ± 3.18
24.75 ± 1.96
Appendix
Table 10: DailyDVS-200 ASR (%, mean ± std over three seeds). “Exact” is the signed-displacement-matched random control.
τ
Yu PIL- L0
Free
Null
0
0.0 ± 0.0
0.0 ± 0.0
81.56 ± 5.81
0.18
5.65 ± 1.79
0.0 ± 0.0
81.56 ± 5.81
0.19
23.42 ± 1.12
0.99 ± 0.49
81.56 ± 5.81
0.195
41.20 ± 2.06
59.93 ± 2.03
81.56 ± 5.81
0.20
62.34 ± 1.49
95.06 ± 0.23
81.56 ± 5.81
0.205
80.11 ± 4.08
95.06 ± 0.23
81.56 ± 5.81
Appendix
Table 11: Three-seed DVS direct comparison: SC-ASR (%, mean ± std). The raw row is unconstrained ASR.
τ
Yu PIL- L0
Free retiming
Null-space retiming
0
0.00 [0.00,1.58]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
0.01
0.00 [0.00,1.58]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
0.05
0.00 [0.00,1.58]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
0.1
0.00 [0.00,1.58]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
0.15
0.00 [0.00,1.58]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
0.17
0.84 [0.23,3.00]
0.00 [0.00,1.58]
74.90 [69.03,79.97]
Appendix
Table 12: SC-ASRA(τ) (%) over the full measured post-hoc representation-budget grid for the complete direct-comparison seed. Values in brackets are pointwise 95% Wilson intervals over 239 clean-correct samples.
ϵ (ms)
Yu PIL- L0
Free retiming
Null-space retiming
100
0.42
0.00
1.67
200
7.11
0.00
51.88
300
68.20
0.00
74.48
404
96.65
0.00
74.48
600
99.58
0.00
74.90
1000
100.00
0.42
74.90
Appendix
Table 13: Temporal-radius success for the three timestamp-retiming methods in the unified DVS seed-0 run at the 10% operating point. Values are A(ϵ) in percent.
Method
timestamp-only
x/y/p fixed
exact A
event budget
displacement rule
role
Null
yes
yes
yes
10% upper
within window
exact-null optimizer
C-PGD
yes
yes
yes
10% upper
within window
exact-null optimizer
Free
yes
yes
no
≈ 10%
native/free
post-hoc evaluation
Yu PIL- L0
yes
yes
no chosen- A constraint
≈ 10% realized
native PIL constraints
post-hoc evaluation
SDA / Yao
native attack
native attack
no chosen- A constraint
native realized
native
post-hoc evaluation
Appendix
Table 14: Threat-model comparison. “Same” means identical between Null and C-PGD in the direct exact-null experiment. Prior attacks retain their native constraints and are therefore evaluated post hoc rather than treated as a common feasible-set contest.
Attack
LIF acc. (%)
frame-CNN acc. (%)
δF
flips
none
96.90±0.08
97.80±0.43
–
–
null tone, a=2
96.50±0.29
97.80±0.43
2×10−7
0
null carrier, a=2
96.57±0.33
97.80±0.43
2×10−7
0
null-PGD, ϵ=.10
50.73±0.53
97.80±0.43
3×10−7
0
null-PGD, ϵ=.25
0.33±0.12
97.80±0.43
3×10−7
0
null-PGD, ϵ=.50
0.00±0.00
97.80±0.43
3×10−7
0
Appendix
Table 15: Controlled attack study (mean ± std over 3 seeds, 1,000 streams/seed). δF is the relative change of the accumulated frame representation. “Flips” is the percentage of frame-CNN labels that change. Rows labeled null lie in the protected aggregation null space.
Victim
10%
20%
30%
ConvSNN
4.11 ± 0.32
10.00 ± 0.51
18.71 ± 1.67
SEW-ResNet18
1.45 ± 0.51
3.50 ± 0.24
4.81 ± 0.32
Event Transformer
2.89 ± 0.31
8.68 ± 0.98
19.61 ± 4.42
Temporal GRU
61.23 ± 4.16
90.59 ± 1.74
97.87 ± 1.19
Appendix
Table 16: Aligned N-MNIST null-space ASR (%, mean ± std over three seeds).
Victim
Clean Top-1
ncc (s0/s1/s2)
Null ASR
matched control
ActionNet
42.18%
412/421/434
97.55 ± 0.28
12.27 ± 2.26
MVFNet
50.21%
506/522/507
99.28 ± 0.11
9.70 ± 1.06
Swin-T
22.74%
211/247/194
99.66 ± 0.60
20.01 ± 2.94
TimeSformer
37.63%
367/330/361
97.94 ± 0.61
9.74 ± 2.61
Appendix
Table 17: DailyDVS protocol-scale architecture results at the 10% exact-null operating point. These are models trained under our common preprocessing, not official benchmark checkpoints. Clean Top-1 is measured on our exact evaluation split. ncc lists clean-correct attacked clips for seeds 0/1/2. All optimized and matched-control rows preserve the protected tensor exactly and give bit-identical protected-consumer logits.
Victim
10% budget
20% budget
ConvSNN
53.9±3.4
74.2±3.7
SEW
12.9±1.1
23.9±1.2
Transformer
49.5±4.3
67.1±2.6
GRU
59.0±2.6
73.3±1.6
Appendix
Table 18: Cross-modality extension to SHD. Three-seed ASR (%, mean ± std) under exact protected equality.
Table 20: DVS observer-family ablations. Top: mean relative representation change for the same canonical-null ConvSNN attacks ( n=239 ). Bottom: blind-space size and ASR when the attack is constrained to multiple phase observers.
Event cameras emit asynchronous events in response to environmental appearance changes. The scarcity of real-world event datasets makes simulation essential. However, most simulators infer event timestamps from frame sequences, forcing many threshold crossings to share a small set of discrete times; a failure mode we term timestamp batching that worsens under fast motion and occlusion. We present TIDES, a continuous-time event simulator built on dynamic Gaussian splatting. Because TIDES operates on an explicit 3D scene representation with learnt geometry and motion, it can derive per-pixel intensity dynamics directly from the scene, rather than by differencing rendered frames. This enables accurate threshold-crossing prediction, including multiple crossings per rendering step, without temporal upsampling or frame interpolation. The same 3D scene model reveals where objects partially occlude one another; TIDES uses this to guide adaptive time stepping, concentrating computation only in regions where occlusion dynamics make simple models of brightness change unreliable. Finally, we model finite sensor bandwidth using a tile-level arbiter whose throughput, jitter, and event drops reproduce realistic sensor artifacts. Across paired RGB-event benchmarks, TIDES attains state-of-the-art event-stream fidelity. We also show that events simulated by TIDES transfer more effectively to real downstream tasks than competitors'.
Christopher Thirgood, Dipon Kumar Ghosh, Simon Hadfield
Event cameras provide low-latency, high temporal resolution perception for real-time vision tasks such as robotics.The novel circuitry (i.e. asynchronous, independent pixels) that enables these advantages also introduces new algorithmic challenges. Velocity-invariant representations alleviate missing observations under slow motion and motion blur under fast motion, but most discard temporal information by converting events into image-like representations. We propose Set of Centre Active Receptive Fields (SCARF), a real-time velocity-invariant representation that preserves raw events while consistently handling fast motion, stationary scenes, and independently moving objects. SCARF achieves state-of-the-art performance in both computational efficiency and representation quality.
Mikihiro Ikura, Luna Gava, Jiahang Wu +2
Istituto Italiano di Tecnologia, Via Morego 30 16163 Genova, Italy
We propose tokenization of events and present a tokenizer, Spiking Patches, specifically designed for event cameras. Given a stream of asynchronous and spatially sparse events, our goal is to discover an event representation that preserves these properties. Prior works have represented events as frames or as voxels. However, while these representations yield high accuracy, both frames and voxels are synchronous and decrease the spatial sparsity. Spiking Patches gives the means to preserve the unique properties of event cameras and we show in our experiments that this comes without sacrificing accuracy. We evaluate our tokenizer using a GNN, PCN, and a Transformer on gesture recognition and object detection. Tokens from Spiking Patches yield inference times that are up to 3.4x faster than voxel-based tokens and up to 10.4x faster than frames. We achieve this while matching their accuracy and even surpassing in some cases with absolute improvements up to 3.8 for gesture recognition and up to 1.4 for object detection. Thus, tokenization constitutes a novel direction in event-based vision and marks a step towards methods that preserve the properties of event cameras.