Continual Action Quality Assessment via Adaptive Manifold-Aligned Graph Regularization
Authors: Kanglei Zhou, Qingyi Pan, Xingxing Zhang, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang, Liyuan Wang
Organizations: Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing 100084, China · Department of Statistics and Data Science, Tsinghua University, Beijing 100084, China · Department of Computer Science and Technology, Institute for AI, BNRist Center, Tsinghua-Bosch Joint ML Center, THBI Lab, Tsinghua University · Department of Computer Science, Durham University, DH1 3LE Durham, U.K. · State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing 100191, China · Zhongguancun Laboratory, Beijing 100190, China
Action Quality Assessment (AQA) quantifies human actions in videos, supporting applications in sports scoring, rehabilitation, and skill evaluation. A major challenge lies in the non-stationary nature of quality distributions in real-world scenarios, which limits the generalization of conventional methods. We introduce Continual AQA (CAQA), which equips AQA with Continual Learning (CL) capabilities to handle evolving distributions while mitigating catastrophic forgetting. Although parameter-efficient fine-tuning of pretrained models has shown promise in continual learning, our empirical study shows that the evaluated adapter-based PEFT setting provides less effective downstream adaptation than FPFT for fine-grained AQA. Our empirical and theoretical analyses reveal two insights: (i) sufficiently expressive backbone adaptation is important for bridging the upstream--downstream representation gap; yet (ii) uncontrolled FPFT may induce overfitting and feature manifold shift, thereby aggravating forgetting. To address this, we propose Adaptive Manifold-Aligned Graph Regularization (MAGR++), which couples backbone fine-tuning that stabilizes shallow layers while adapting deeper ones with a two-step feature rectification pipeline: a manifold projector to translate deviated historical features into the current representation space, and a graph regularizer to align local and global distributions. We construct four CAQA benchmarks from three datasets with tailored evaluation protocols and strong baselines, enabling systematic cross-dataset comparison. Extensive experiments show that MAGR++ achieves state-of-the-art performance, with average correlation gains of 3.6% offline and 12.2% online over the strongest baseline, confirming its robustness and effectiveness.
Figures & tables
Fig. 1 : Empirical motivation and challenges of CAQA. 1 : Real-world AQA exhibits non-stationary quality distributions across action types and athlete groups. 1 : Such distribution shifts lead to severe forgetting when moving from joint training to sequential fine-tuning. 1 and 1 : Recent strong CL baselines still show clear performance gaps under offline and online CAQA settings, respectively.
Fig. 2 : SRCC and rMSE comparison of fixed backbone, PEFT (I3D-Adapters), and FPFT in representative AQA tasks.
Fig. 3 : Illustration of fine-tuning choices under different upstream-downstream discrepancies. Parameter-constrained adaptation can be effective when limited downstream adjustment is sufficient, whereas larger discrepancies may benefit from more expressive full-parameter adaptation, as observed in our AQA experiments.
Fig. 4 : Core idea of MAGR++: (a) Old features ( blue circles ) deviate from the current manifold ( orange curve ) due to manifold shift; (b) Mixing old and new features ( green circles ) leads to confusion in score regression; (c) The manifold projector translates old features from the previous manifold ( yellow curve ) to the current one; (d) The feature space is further aligned with the quality score space.
Fig. 5 : Overview of MAGR++. At the end of the session t−1 5 , representative features are selected via Ordered Uniform Sampling (OUS, 5 ) and stored in the memory bank M . At the start of the session t 5 , the backbone is adapted with layer-adaptive FPFT 5 to balance stability and plasticity. A Manifold Projector (MP) is then trained 5 to align old features with the evolving feature space 5 , enabling effective replay and adaptation 5 . Finally, the memory bank is refreshed with rectified old features and newly sampled prototypes 5 .
Fig. 6 : Illustration of the proposed layer-adaptive FPFT strategy. (a) Adaptive layer selection determines the optimal layer boundary lopt . (b) Constrained full tuning stabilizes shallow layers while preserving the adaptability of deeper quality-aware representations.
Fig. 7 : Illustrations of IIJ-GR: 7 Euclidean distance, 7 Angular distance, and 7 Distance Matrix Partitioning (DMP).
TABLE I : Offline performance comparison. The primary metric is ρavg , with Joint Training (JT) as the Upper Bound (UB) and Sequential Fine-Tuning (SFT) as the Lower Bound (LB). Best results are highlighted in bold. ↑ / ↓ indicate higher/lower is better.
TABLE II : Online performance comparison. The primary metric is ρavg , with Sequential Fine-Tuning (SFT) as the Lower Bound (LB). Best results are in bold, excluding UB and LB. ↑ / ↓ indicate higher/lower is better.
Setting
FineDiving
MTL-AQA
UNLV-Dive
Deviation Strength (MSE)
26.85
35.28
51.75
FS-Aug [ 47 ]
−0.09
−4.34
+14.18
MAGR [ 24 ]
+5.66
+6.56
+15.64
Δρavg
MAGR++
+9.50
+9.25
+26.43
TABLE III : Feature deviations and correlation gains.
Method
Publisher
Memory
ρavg ( ↑ )
ρaft ( ↓ )
ρfwt ( ↑ )
RL2 ( ↓ )
w/ MUSDL [ 69 ]
JT (UB)
–
None
0.9211
–
–
0.0035
SFT (LB)
–
None
0.7771
0.0821
0.7337
0.0104
LwF [ 51 ]
TPAMI’17
None
0.7697
0.0675
0.7200
0.0152
DER++ [ 53 ]
NeurIPS’20
Raw Data
0.7846
0.0985
0.6618
0.0109
STAR [ 56 ]
ICLR’25
Raw Data
0.7700
0.0529
0.6619
0.0108
TABLE IV : Plug-and-play performance comparison on CD-AQA. The primary metric is ρavg , with Joint Training (JT) as the Upper Bound (UB) and Sequential Fine-Tuning (SFT) as the Lower Bound (LB). Best results are in bold, excluding UB and LB. ↑ / ↓ indicate higher/lower is better.
Method
Params. (M)
Training Time (h)
Inference FPS
Δρavg
Δρaft
Δρfwt
Feature MER
12.62
2.22
85.83
+0.1825
+0.0731
−0.0003
SLCA [ 57 ]
13.62
2.27
85.83
+0.1765
−0.0672
+0.1127
NC-FSCIL [ 58 ]
12.62
2.33
85.83
+0.2968
−0.0378
+0.0180
MAGR [ 24 ]
12.63
2.23
85.83
+0.3521
−0.1301
+0.1376
ProNC [ 59 ]
12.62
2.35
85.83
+0.3518
−0.0198
+0.1246
MAGR++ (Ours)
12.63
2.32
85.83
+0.3925
−0.1421
+0.0736
TABLE V : Computational performance on MTL-AQA. All metrics are reported as improvements over the offline LB.
ID
Setting
ρavg ( ↑ )
ρaft ( ↓ )
ρfwt ( ↑ )
LA-FPFT
MP
IIJ-GR
1
✓
✓
✓
0.9205
0.0103
0.1274
2
✕
✓
✓
0.9135 -1%
0.0233 +126%
0.1204 -5%
3
✓
✕
✓
0.8418 -9%
0.0995 +866%
0.1082 -15%
4
✓
✓
✕
0.8548 -7%
0.0143 +39%
0.1211 -5%
5
✕
✕
✓
0.8016 -13%
0.1427 +1285%
0.0965 -24%
TABLE VI : Ablation study of core components on MTL-AQA. Performance variations are reported relative to ID 1.
TABLE VII : Ablation study of design choices for core components in MTL-AQA. Performance variations are reported relative to the full model (ID 1).
Fig. 8 : Comparison of different memory sizes on MTL-AQA.
Fig. 9 : Cluster quality ratio and overall correlation.
Fig. 10 : Impact of layer selection threshold ϵ .
Fig. 11 : Performance vs. parameter changes across task orders.
Fig. 12 : Performance comparison under label scarcity and labeling noise. 12 varies the number of training samples per session, and 12 evaluates robustness under noisy annotations. 12 , 12 , 12 , and 12 show correlation plots at noise level 7, with fitted regression lines.
Fig. 13 : Loss landscapes on MTL-AQA. 13 - 13 show loss curves for five sessions, while 13 presents the average curve.
Fig. 14 : Visualization of t-SNE feature distributions (first three columns) and overall correlation plots (last column). The feature space is projected into two dimensions and normalized to [0,1] .
Fig. 15 : Error analysis on MTL-AQA. 15 : Boxplots of absolute errors with mean, median, and standard deviation. 15 : Cumulative error accuracy curves with area under the curve (AUC).
TABLE A8 : Comparison results on UNLV-Vault and JIGSAWS. The primary metric is ρavg , with Sequential Fine-Tuning (SFT) as the Lower Bound (LB). Best results are in bold, excluding UB and LB. ↑ / ↓ indicate higher/lower is better.
Fig. A16 : Representative samples from MTL-AQA covering high-, mid-, and low-score cases. The first five columns show sampled frames, and the last column reports assessment results with errors. A16 and A16 show successful cases, while A16 depicts a failure case.
ID
Setting
ρavg ( ↑ )
ρaft ( ↓ )
ρfwt ( ↑ )
1
MAGR++ (Ours)
0.9205
0.0103
0.1274
2
MP w/o Residual Link
0.8933 -3%
0.0389 +278%
0.0693 -46%
3
Eq. 11 w/ KL Loss
0.9173 -0.3%
0.0155 +51%
0.1029 -19%
Appendix
TABLE A9 : Additional ablation results on MTL-AQA. Reported percentages denote relative changes compared to ID 1.
Method
Optimizer
LR
Epochs
Memory
Replay Ratio
Replay Mini-Batch
Batch Size
Source
EWC
Adam
1e−4
1/50
–
–
–
5
Repo
SI
Adam
1e−4
1/50
–
–
–
5
Repo
LwF
Adam
1e−4
1/50
–
–
–
5
Repo
MER
Adam
1e−4
1/50
200
1:1
3
5
Repo
DER++
Adam
1e−4
1/50
200
1:1
3
5
Repo
TOPIC
Adam
1e−4
1/50
200
1:1
3
5
Re-implemented from the paper ( λ1=0.5 , λ2=0.55 )
Appendix
TABLE A10 : Implementation details of compared methods. Benchmark-level settings, such as memory size, replay mini-batch, batch size, and online/offline training epochs, are fixed by the CAQA protocol and shared across methods. Method-specific hyper-parameters, such as regularization weights and learning rates when applicable, follow the original papers or official implementations.
Fig. A17 : Robustness under parameter mutation. (a) Illustration of the mutation process, where Gaussian noise is injected into the backbone parameters at the beginning of each session to simulate violations of the model continuity assumption. (b) Performance under increasing mutation strength σ . MAGR++ maintains a stable average correlation under mild and moderate mutations, and degrades more gracefully than prior methods, demonstrating strong robustness beyond the theoretical assumptions.
Fig. A18 : Sensitivity analysis of MAGR++ with respect to the loss weights λtune , λproj , and λreg . Each parameter is varied independently while fixing the others to 1.
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type. This setting limits real-world deployment, as it requires prior action-type knowledge to select the corresponding model and suffers from poor generalization across diverse actions. To address these limitations, we study the challenging task of all-in-one AQA, which aims to assess heterogeneous actions within a single unified model. We propose a novel Mixture of Action Knowledge Experts (MoAKE) framework, designed to mitigate negative knowledge transfer caused by large semantic discrepancies among actions. MoAKE learns complementary experts that capture diverse action patterns within a shared semantic space and dynamically aggregates their knowledge to adapt the assessment to the input action. Each expert is tailored with segment-aware prototypes to handle varying temporal lengths, together with an Adaptive Intra- and Inter-Segment Relationship Modeling (AIISRM) module to model multi-granularity temporal dynamics. Furthermore, we establish comprehensive benchmarks for all-in-one as well as zero/few-shot AQA. Extensive experiments on three long-term datasets demonstrate that MoAKE significantly outperforms existing methods in the all-in-one setting, while also achieving consistent generalization on three short-term datasets under zero/few-shot evaluation. Code is available at https://github.com/XuHuangbiao/MoAKE.
Huangbiao Xu, Huanqi Wu, Xiao Ke +3
Fujian Provincial Key Laboratory of Networking Computing and Intelligent Information Processing, College of Computer and Data Science, Fuzhou University, Fuzhou 350108, China · Engineering Research Center of Big Data Intelligence, Ministry of Education, Fuzhou 350108, China · School of Intelligence Science and Technology, University of Science and Technology Beijing, Beijing 100083, China
Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by updating a largely shared parameter set. This often leads to cross-level task interference, hindering accurate adaptation to the current task and object. To address this limitation, we propose HyLoVQA. It maintains a drift-resilient memory bank of anchors. The bank stores the content of visual objects and textual tasks, and they are updated using current input features. Conditioned on retrieved anchors, a hypernetwork generates lightweight Low-Rank Adaptation (LoRA) adapters. This ensures parameter efficiency, allowing the model to adapt to each task and object dynamically. Additionally, we formulate an alignment loss that aligns semantic discrepancies in the feature space with functional changes in the parameter space, thereby constraining LoRA adapters to remain focused on the current task and object. Extensive experiments on VQA v2 and NExT-QA under both standard and compositional settings demonstrate the superiority of HyLoVQA over prior state-of-the-art methods.
Yiran Wang, Chenyi Xiong, Ziyue Qin +3
School of Computer Science, Hubei University, Wuhan 430062, China · Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University), Wuhan 430062, China · Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
In continual visual question answering (VQA), existing Continual Learning (CL) methods are mostly built for symmetric, unimodal architectures. However, modern Vision-Language Models (VLMs) violate this assumption, as their trainable components are inherently asymmetric. This structural mismatch renders VLMs highly prone to catastrophic forgetting when learning from continuous data streams. Specifically, the asymmetry causes standard global regularization to favor the massive language decoder during optimization, leaving the smaller but critical visual projection layers highly vulnerable to interference. Consequently, this localized degradation leads to a severe loss of compositional reasoning capabilities. To address this, we propose Asymmetric Information Masking (AIM), which balances stability and plasticity by applying targeted masks based on modality-specific sensitivity. Experiments on VQA v2 and GQA under continual VQA settings show that AIM achieves state-of-the-art performance in both Average Performance (AP) and Average Forgetting (AF), while better preserving generalization to novel skill-concept compositions.
Peifeng Zhang, Zice Qiu, Donghua Yu +4
Sun Yat-Sen University · National Supercomputing Center in Shenzhen · Tsinghua University