Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
Organizations: Graduate School of Information Science and Engineering, Ritsumeikan University, Osaka, Japan · College of Information Science and Engineering, Ritsumeikan University, Osaka, Japan · Alumni Association of Carnegie Mellon University, Pittsburgh, PA, USA
Abstract
This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation (), while updating only 8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.
Figures & tables
| Game | Abbrev. | Genre | Inc. | Dec. | Unchg. |
|---|---|---|---|---|---|
| ApexSpeed | apex | Racing | 6629 | 4446 | 2375 |
| Solid | solid | Racing | 6520 | 4333 | 1947 |
| TinyCars | tiny | Racing | 5944 | 4402 | 2279 |
| Heist | fps | Shooter | 6561 | 3986 | 2337 |
| Shootout | gallery | Shooter | 7049 | 3121 | 2236 |
| TopDown | topdown | Shooter | 6914 | 3963 | 2724 |
| Method | Apex W-F1 | Apex M-F1 | Apex Acc | Endless W-F1 | Endless M-F1 | Endless Acc | FPS W-F1 | FPS M-F1 | FPS Acc |
|---|---|---|---|---|---|---|---|---|---|
| Random Forest | 51.96 (1.16) | 43.00 (1.26) | 56.68 (0.78) | 53.91 (2.18) | 46.86 (2.32) | 59.43 (1.59) | 56.03 (2.85) | 46.38 (2.92) | 60.82 (2.45) |
| ResNet+LSTM Full | 57.27 (1.43) | 53.22 (1.47) | 58.07 (1.35) | 58.27 (2.40) | 54.89 (2.79) | 58.47 (1.69) | 62.38 (1.79) | 57.73 (2.06) | 63.15 (1.60) |
| ResNet+LSTM SALFT | 56.27 (1.65) | 52.23 (1.09) | 56.90 (2.04) | 56.39 (2.71) | 52.08 (3.52) | 57.75 (1.29) | 61.21 (2.23) | 55.98 (2.55) | 61.77 (2.64) |
| ViViT Full | 58.05 (1.84) | 54.44 (2.23) | 58.20 (2.16) | 54.54 (1.26) | 50.58 (0.73) | 54.93 (1.79) | 62.37 (2.90) | 57.46 (3.10) | 63.03 (2.86) |
| ViViT SALFT(Ours) | 58.45 (0.97) | 54.00 (1.69) | 59.13 (0.99) | 55.58 (1.90) † | 51.49 (2.48) | 56.34 (2.04) | 62.75 (1.93) | 57.99 (2.20) | 63.58 (1.84) |
| Method | Gallery W-F1 | Gallery M-F1 | Gallery Acc | Gun W-F1 | Gun M-F1 | Gun Acc | Platform W-F1 | Platform M-F1 | Platform Acc |
| Method | Trainable Params ↓ | FLOPs per Step (G) ↓ |
|---|---|---|
| Full fine-tuning | 88.65M (100%) | 6402.90 |
| Second-Stage Selective Fine-Tuning | 7.09M (7.998%) | 4449.00 |
| Type | Class | Rollout | Grad-CAM | Grad-SAM | Our Method |
|---|---|---|---|---|---|
| Positive ↓ | Predicted | 46.96 | 47.47 | 49.82 | 24.01 |
| Target | 45.81 | 44.54 | 48.69 | 23.13 | |
| Negative ↑ | Predicted | 86.86 | 87.91 | 84.67 | 93.91 |
| Target | 80.75 | 84.64 | 78.65 | 87.72 |