Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary F1 scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s
Figures & tables
Fig. 1: The proposed architecture aligns feature subspaces into isolated temporal group chains ( P1 and P2 ) processed by independent Mamba blocks without cross-group state sharing. At each time step, their recurrent outputs are concatenated along the feature dimension, projected via learned linear fusion ( 2048→64 ), and processed by a dilated TCN, MLP boundary head, and penalized dynamic programming (DP) decoder.
Method
Params
Breakfast
GTEA
50Salads
@1 s
@0.4 s
@1 s
MS-TCN [ 1 ]
0.80M
0.349
0.557
0.370
ASFormer [ 3 ]
1.13M
N/C
0.575
0.455
BaFormer [ 4 ]
1.63M
0.313
0.606
0.524
DiffAct [ 5 ]
1.21M
0.182
0.285
0.312
ASRF [ 2 ]
1.30M
0.370
0.592
0.607
TABLE I: Primary action-boundary F1 under the common one-to-one protocol. Bold and underlined entries are best and second best. N/C: reproduction not completed within the available compute budget.
Width
Params
F1@1
2048
13.359M
0.4687±0.0735
64
0.658M
0.4818±0.0653
32
0.453M
0.4736±0.0460
TABLE II: Stateful streaming results on 50Salads: five folds, two seeds, zero neural look-ahead, and one-sample peak confirmation. Width denotes channels per recurrent group.
Configuration
Primary F1@1
Full pipeline †
0.611±0.046
w/o nonlinear head (linear head)
0.603±0.032
w/o input normalization
0.593±0.040
w/o learned fusion (mean pooling)
0.578±0.056
w/o TCN
0.559±0.044
TABLE III: Component study on 50Salads (five folds, two seeds). † : full-pipeline result from the main experiment; component variants were trained in a separate batch.
Temporal action segmentation (TAS) in untrimmed videos requires dense temporal supervision. However, most of the annotation cost is spent identifying action transitions where segmentation errors concentrate and small temporal shifts can disproportionately degrade segment-level metrics. We introduce B-ACT, a clip-budgeted active learning framework that explicitly allocates supervision to these error-prone boundary regions. B-ACT operates in a hierarchical two-stage loop: (i) it ranks and queries unlabeled videos using predictive uncertainty, and (ii) within each selected video, it detects candidate transitions from the current model predictions and selects the top-K boundaries via a novel boundary score. The boundary score fuses neighborhood uncertainty, class ambiguity, and temporal prediction dynamics to reveal the underlying importance of each frame. Importantly, our annotation protocol requests labels only at the boundary frames while still training on boundary-centered clips to exploit temporal context through the model's receptive field. Extensive experiments on GTEA, 50Salads, and Breakfast demonstrate that boundary-centric supervision delivers strong label efficiency and consistently surpasses representative TAS active learning baselines and prior state of the art under sparse budgets. Gains are largest on datasets where performance is highly sensitive to boundary placement, as measured by edit and overlap-based F1 metrics.
Halil Ismail Helvaci, Sen-ching Samson Cheung
Department of Electrical and Computer Engineering, University of Kentucky, Lexington, KY, USA
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.
Runzhong Zhang, Yueqi Duan, Yang Chen +4
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Department of Electronic Engineering, Tsinghua University, Beijing, China · Amazon, Seattle, America
Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed videos. Existing appearance-based video detection methods often struggle with limited viewpoint diversity during training, while motion-based detection approaches frequently fail to model fine-grained temporal relationships across consecutive motion windows. This paper introduces a novel two-stage action detection approach designed to improve both view-invariance and global temporal coherence properties. In the first stage, we extract motion features from augmented virtual viewpoints, solely used at training. Then, the second stage introduces a new view-invariant, multi-scale temporal encoder based on selective state-space sequence modelling to aggregate information across viewpoints and time scales. Experiments on PKU-MMD and BABEL benchmarks demonstrate that this approach significantly outperforms state-of-the-art methods in all considered splits. Code and trained models are available at: https://icb-vision-ai.github.io/HydraView-TAD
Yannick Porto, Renato Martins, Thomas Chalumeau +1