Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary F1 scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s
Figures & tables
Fig. 1: The proposed architecture aligns feature subspaces into isolated temporal group chains ( P1 and P2 ) processed by independent Mamba blocks without cross-group state sharing. At each time step, their recurrent outputs are concatenated along the feature dimension, projected via learned linear fusion ( 2048→64 ), and processed by a dilated TCN, MLP boundary head, and penalized dynamic programming (DP) decoder.
Method
Params
Breakfast
GTEA
50Salads
@1 s
@0.4 s
@1 s
MS-TCN [ 1 ]
0.80M
0.349
0.557
0.370
ASFormer [ 3 ]
1.13M
N/C
0.575
0.455
BaFormer [ 4 ]
1.63M
0.313
0.606
0.524
DiffAct [ 5 ]
1.21M
0.182
0.285
0.312
ASRF [ 2 ]
1.30M
0.370
0.592
0.607
TABLE I: Primary action-boundary F1 under the common one-to-one protocol. Bold and underlined entries are best and second best. N/C: reproduction not completed within the available compute budget.
Width
Params
F1@1
2048
13.359M
0.4687±0.0735
64
0.658M
0.4818±0.0653
32
0.453M
0.4736±0.0460
TABLE II: Stateful streaming results on 50Salads: five folds, two seeds, zero neural look-ahead, and one-sample peak confirmation. Width denotes channels per recurrent group.
Configuration
Primary F1@1
Full pipeline †
0.611±0.046
w/o nonlinear head (linear head)
0.603±0.032
w/o input normalization
0.593±0.040
w/o learned fusion (mean pooling)
0.578±0.056
w/o TCN
0.559±0.044
TABLE III: Component study on 50Salads (five folds, two seeds). † : full-pipeline result from the main experiment; component variants were trained in a separate batch.
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Department of Electronic Engineering, Tsinghua University, Beijing, China · Amazon, Seattle, America