SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
Authors: Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela
Organizations: Universidad de Alcal´a, Alcal´a de Henares, Spain · Universidad Polit´ecnica de Madrid, Madrid, Spain · Universidad Rey Juan Carlos, M´ostoles, Spain
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules. While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed. Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters. This forces an expensive dense reconstruction, nullifying much of the potential speedup. We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets. We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3. When scaled to InternVideoNext-L, our approach surpasses the previous state of the art at roughly half its computational cost. Moreover, the sparse formulation naturally supports auxiliary task tokens, which improves fine-grained assembly detection on ATTACH.
Figures & tables
Figure 1 : TFLOPs vs. accuracy on THUMOS-14. SV-TAD (red squares) matches or exceeds AdaTAD [ 21 ] (blue circles) at substantially lower compute across three backbone scales: 43% fewer TFLOPs with VMAEv2-B, 64% with VMAEv2-L, and 51% when comparing our InternVideoNext-L against AdaTAD’s VMAEv2-G.
Figure 2 : SV-TAD architecture. Input frames and auxiliary tokens ( X ) pass through a frozen pretrained ViT backbone with interleaved token selection (TS-ON) blocks. The first three blocks use standard 2D convolutions, with the rest using SparseConv2D. The expanded view shows our modified ViT block: frozen attention computes importance scores, token selection prunes uninformative patches, and the trainable sparse adapter operates directly on the sparse sequence. Timing annotations show per-block latency decreasing from 60ms to 7.6ms as the sequence progressively shrinks.
Figure 3 : Kernel Timing comparison on THUMOS-14. DenseConv2D (blue) keeps latency nearly constant across different keep rates, while our SparseConv2D (orange) scales linearly with Nkept .
Table 1 : Effect of keep rate kr on THUMOS-14 (InternVideoNext-L backbone).
Figure 4 : Sparse adapter architecture. Our bottleneck adapter (orange, trainable; the surrounding ViT block is frozen) is inserted after each Transformer block’s attention layer. Visual ( Xvis ) and auxiliary ( Xaux ) tokens are down-projected ( DownProj1 ) before SparseConv2D gathers each kept token’s spatial neighbors via the precomputed neighbor index table N ( −1 marks pruned/out-of-bounds neighbors, top-left inset). CrossAttn then lets auxiliary tokens query the sparse visual features, updating only Xaux at no extra cost to the visual path, before UpProj1 restores the channel dimension and a learnable scalar γ modulates the adapter’s residual contribution.
Figure 5 : Sparse convolution efficiency scaling at kr=0.3 . (a) Our sparse convolution (orange) achieves 1.75 × speedup over dense (blue) at 6,144 frames (40ms vs. 70ms). (b) Dense convolution memory grows to ∼ 6GB at 6,144 frames regardless of sparsity. Our sparse convolution maintains constant memory ( ∼ 800MB) via chunked processing, achieving 7.5 × reduction.
Table 7
Method
TFLOPs ↓
0.1
0.2
0.3
0.4
0.5
Avg.
SV-TAD
10.29
20.17
18.69
16.90
14.39
11.25
16.28
SV-TAD + Kp
11.17
22.58
21.10
19.21
16.80
13.76
18.69
Table 4 : Ablation of auxiliary landmark supervision on ATTACH (Person Split) using VideoMAEv2-B.
ActivityNet-1.3
THUMOS14
Method
Backbone
TFLOPs ↓
0.5
0.75
0.95
Avg.
TFLOPs ↓
0.3
0.4
0.5
0.6
0.7
Avg.
AFSD [ 19 ]
I3D
—
52.40
35.30
6.50
34.40
—
67.3
62.4
55.5
43.7
31.1
52.0
E2E-TAD [ 23 ]
SlowFast-R50
82.48
50.47
35.99
10.33
34.10
82.48
69.4
64.3
56.0
46.4
34.9
54.2
BasicTAD [ 33 ]
SlowOnly-R50
—
51.20
33.41
7.57
33.12
—
75.5
70.8
63.5
50.9
37.4
59.2
TALLFormer [ 5 ]
VideoSwin-B
—
54.10
36.20
7.90
35.60
—
76.0
70.0
63.2
—
34.5
59.4
Re 2 TAD [ 38 ]
Re 2 VSwin-T
—
54.75
37.81
9.03
36.80
—
77.0
71.5
62.4
49.7
36.3
59.4
Table 5 : Results comparison on ActivityNet-1.3 and THUMOS14 (* indicates mAP obtained from official logs). Size-matched comparisons are highlighted.
Hyperparameter
THUMOS-14
ActivityNet
ATTACH
Epochs
100
10
50
Batch size
2
16
4
Number of GPUs
2
4
2
Optimizer
AdamW
Base learning rate
1×10−4
Weight decay
5×10−2
Table S1 : Complete training hyperparameters across all datasets.
THUMOS-14 (mAP@IoU)
Method
Source
0.3
0.4
0.5
0.6
0.7
Avg.
AdaTAD++ (VMAE-B)
Reported
88.3
83.7
76.6
65.0
51.8
73.3
AdaTAD++ (VMAE-B)
Ours
84.64
80.33
72.68
62.44
49.67
69.95
Table S2 : AdaTAD++ reproduction comparison. Reported values from original paper vs. our reproduction using OpenTAD.
Temporal Action Detection (TAD) requires precise localization of action boundaries within long, untrimmed video sequences. While current high-performing methods achieve strong accuracy, they are often characterized by excessive parameter counts, substantial computational overhead, and a reliance on specialized operators that hinder deployment across diverse hardware platforms. This paper presents LiquidTAD, a framework that distills the exponential relaxation prior of liquid neural dynamics into a parallel temporal operator, rather than reproducing full Liquid Neural Network (LNN) dynamics. By introducing a Parallel Liquid-inspired Relaxation mechanism, sequential ODE solving is avoided through a fully vectorized, non-recursive formulation built entirely upon standard neural operations, enabling hardware-agnostic deployment with linear complexity with respect to the temporal length. A complementary Hierarchical Decay-Rate Sharing Strategy further adapts this relaxation prior across feature pyramid levels, stabilizing optimization and implicitly compensating for temporal compression in deeper layers. Experimental evaluations on THUMOS-14 and ActivityNet-1.3 demonstrate that LiquidTAD achieves accuracy competitive with strong baselines while substantially lowering the model footprint. Specifically, on THUMOS-14, LiquidTAD achieves 69.46% average mAP with only 10.82M parameters and 27.17G FLOPs, reducing the parameter count by over 60% compared with ActionFormer.
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with negligible overhead and executing them for end-to-end gains. We present SPADE, a training-free sparse-attention engine of three parts: (i) vDiT-SSR, a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions; (ii) runtime scheme generation using SICS and a head-wise policy; and (iii) an executor with low-overhead index search, flash block-sparse attention, and kernel grouping. Across Hunyuan-Video and Wan 2.1/2.2 for text-to-video and image-to-video generation, SPADE raises sparsity and speed while preserving quality, accelerating attention by 2.26x-3.40x and end-to-end inference by 1.49x-1.80x. Our code is open-sourced at https://github.com/6somehow/DAC-SPADE.
Video understanding is a crucial part of computer vision, with numerous application scenarios. With the increasing popularity of mobile devices, an increasing number of efforts are trying to deploy video understanding models on them. However, existing video understanding models are difficult to deploy due to their large size and prohibitive power consumption. Spiking Neural Networks (SNNs) have shown bioplausibility and low power advantages over Artificial Neural Networks (ANNs), especially on neuromorphic chips which are regarded as essential components of future mobile devices. However, excessively long conversion time-steps and severe performance degradation problems limit their application. To solve the problems above, we explore the application of SNNs on temporal action detection (TAD), which is an important task in video understanding, and propose the first SNN-based end-to-end TAD architecture coined as SpikeTAD. While maintaining extremely low power consumption, SpikeTAD achieves an average mAP of 67.2% in THUMOS14 and 37.42% in ActivityNet-1.3, demonstrating the feasibility of a low-power TAD model. Our code is available at https://github.com/MCG-NJU/SpikeTAD.
Min Yang, Mi Zhou, Limin Wang
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, Jiangsu, China