SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
Authors: Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela
Organizations: Universidad de Alcal´a, Alcal´a de Henares, Spain · Universidad Polit´ecnica de Madrid, Madrid, Spain · Universidad Rey Juan Carlos, M´ostoles, Spain
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules. While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed. Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters. This forces an expensive dense reconstruction, nullifying much of the potential speedup. We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets. We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3. When scaled to InternVideoNext-L, our approach surpasses the previous state of the art at roughly half its computational cost. Moreover, the sparse formulation naturally supports auxiliary task tokens, which improves fine-grained assembly detection on ATTACH.
Figures & tables
Figure 1 : TFLOPs vs. accuracy on THUMOS-14. SV-TAD (red squares) matches or exceeds AdaTAD [ 21 ] (blue circles) at substantially lower compute across three backbone scales: 43% fewer TFLOPs with VMAEv2-B, 64% with VMAEv2-L, and 51% when comparing our InternVideoNext-L against AdaTAD’s VMAEv2-G.
Figure 2 : SV-TAD architecture. Input frames and auxiliary tokens ( X ) pass through a frozen pretrained ViT backbone with interleaved token selection (TS-ON) blocks. The first three blocks use standard 2D convolutions, with the rest using SparseConv2D. The expanded view shows our modified ViT block: frozen attention computes importance scores, token selection prunes uninformative patches, and the trainable sparse adapter operates directly on the sparse sequence. Timing annotations show per-block latency decreasing from 60ms to 7.6ms as the sequence progressively shrinks.
Figure 3 : Kernel Timing comparison on THUMOS-14. DenseConv2D (blue) keeps latency nearly constant across different keep rates, while our SparseConv2D (orange) scales linearly with Nkept .
Table 1 : Effect of keep rate kr on THUMOS-14 (InternVideoNext-L backbone).
Figure 4 : Sparse adapter architecture. Our bottleneck adapter (orange, trainable; the surrounding ViT block is frozen) is inserted after each Transformer block’s attention layer. Visual ( Xvis ) and auxiliary ( Xaux ) tokens are down-projected ( DownProj1 ) before SparseConv2D gathers each kept token’s spatial neighbors via the precomputed neighbor index table N ( −1 marks pruned/out-of-bounds neighbors, top-left inset). CrossAttn then lets auxiliary tokens query the sparse visual features, updating only Xaux at no extra cost to the visual path, before UpProj1 restores the channel dimension and a learnable scalar γ modulates the adapter’s residual contribution.
Figure 5 : Sparse convolution efficiency scaling at kr=0.3 . (a) Our sparse convolution (orange) achieves 1.75 × speedup over dense (blue) at 6,144 frames (40ms vs. 70ms). (b) Dense convolution memory grows to ∼ 6GB at 6,144 frames regardless of sparsity. Our sparse convolution maintains constant memory ( ∼ 800MB) via chunked processing, achieving 7.5 × reduction.
Table 7
Method
TFLOPs ↓
0.1
0.2
0.3
0.4
0.5
Avg.
SV-TAD
10.29
20.17
18.69
16.90
14.39
11.25
16.28
SV-TAD + Kp
11.17
22.58
21.10
19.21
16.80
13.76
18.69
Table 4 : Ablation of auxiliary landmark supervision on ATTACH (Person Split) using VideoMAEv2-B.
ActivityNet-1.3
THUMOS14
Method
Backbone
TFLOPs ↓
0.5
0.75
0.95
Avg.
TFLOPs ↓
0.3
0.4
0.5
0.6
0.7
Avg.
AFSD [ 19 ]
I3D
—
52.40
35.30
6.50
34.40
—
67.3
62.4
55.5
43.7
31.1
52.0
E2E-TAD [ 23 ]
SlowFast-R50
82.48
50.47
35.99
10.33
34.10
82.48
69.4
64.3
56.0
46.4
34.9
54.2
BasicTAD [ 33 ]
SlowOnly-R50
—
51.20
33.41
7.57
33.12
—
75.5
70.8
63.5
50.9
37.4
59.2
TALLFormer [ 5 ]
VideoSwin-B
—
54.10
36.20
7.90
35.60
—
76.0
70.0
63.2
—
34.5
59.4
Re 2 TAD [ 38 ]
Re 2 VSwin-T
—
54.75
37.81
9.03
36.80
—
77.0
71.5
62.4
49.7
36.3
59.4
Table 5 : Results comparison on ActivityNet-1.3 and THUMOS14 (* indicates mAP obtained from official logs). Size-matched comparisons are highlighted.
Hyperparameter
THUMOS-14
ActivityNet
ATTACH
Epochs
100
10
50
Batch size
2
16
4
Number of GPUs
2
4
2
Optimizer
AdamW
Base learning rate
1×10−4
Weight decay
5×10−2
Table S1 : Complete training hyperparameters across all datasets.
THUMOS-14 (mAP@IoU)
Method
Source
0.3
0.4
0.5
0.6
0.7
Avg.
AdaTAD++ (VMAE-B)
Reported
88.3
83.7
76.6
65.0
51.8
73.3
AdaTAD++ (VMAE-B)
Ours
84.64
80.33
72.68
62.44
49.67
69.95
Table S2 : AdaTAD++ reproduction comparison. Reported values from original paper vs. our reproduction using OpenTAD.