Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
Authors: Yi Xia, Ibrahim Khan, Mury Fajar Dewantoro, Wenwen Ouyang, Ruck Thawonmas
Organizations: Graduate School of Information Science and Engineering, Ritsumeikan University, Osaka, Japan · College of Information Science and Engineering, Ritsumeikan University, Osaka, Japan · Alumni Association of Carnegie Mellon University, Pittsburgh, PA, USA
This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation (p>0.05), while updating only ≈8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.
Figures & tables
Fig. 1 : Representative screenshots of each game in the AGAIN dataset.
Game
Abbrev.
Genre
Inc.
Dec.
Unchg.
ApexSpeed
apex
Racing
6629
4446
2375
Solid
solid
Racing
6520
4333
1947
TinyCars
tiny
Racing
5944
4402
2279
Heist
fps
Shooter
6561
3986
2337
Shootout
gallery
Shooter
7049
3121
2236
TopDown
topdown
Shooter
6914
3963
2724
TABLE I : Games in the AGAIN dataset with number of samples per arousal change class. The column “Abbrev.” stands for the abbreviated game identifier; “Inc.”, “Dec.”, and “Unchg.” denote arousal Increase, Decrease, and Unchanging, respectively.
Method
Apex W-F1
Apex M-F1
Apex Acc
Endless W-F1
Endless M-F1
Endless Acc
FPS W-F1
FPS M-F1
FPS Acc
Random Forest
51.96 (1.16)
43.00 (1.26)
56.68 (0.78)
53.91 (2.18)
46.86 (2.32)
59.43 (1.59)
56.03 (2.85)
46.38 (2.92)
60.82 (2.45)
ResNet+LSTM Full
57.27 (1.43)
53.22 (1.47)
58.07 (1.35)
58.27 (2.40)
54.89 (2.79)
58.47 (1.69)
62.38 (1.79)
57.73 (2.06)
63.15 (1.60)
ResNet+LSTM SALFT
56.27 (1.65)
52.23 (1.09)
56.90 (2.04)
56.39 (2.71)
52.08 (3.52)
57.75 (1.29)
61.21 (2.23)
55.98 (2.55)
61.77 (2.64)
ViViT Full
58.05 (1.84)
54.44 (2.23)
58.20 (2.16)
54.54 (1.26)
50.58 (0.73)
54.93 (1.79)
62.37 (2.90)
57.46 (3.10)
63.03 (2.86)
ViViT SALFT(Ours)
58.45 (0.97)
54.00 (1.69)
59.13 (0.99)
55.58 (1.90) †
51.49 (2.48)
56.34 (2.04)
62.75 (1.93)
57.99 (2.20)
63.58 (1.84)
Method
Gallery W-F1
Gallery M-F1
Gallery Acc
Gun W-F1
Gun M-F1
Gun Acc
Platform W-F1
Platform M-F1
Platform Acc
TABLE II : Values are percentages with 95% confidence intervals in parentheses. Best result per game per metric is highlighted in bold . Cases where ViViT SALFT yields empirical improvements reaching the theoretical minimum p -value ( 0.0625 for N=5 , exact two-sided Wilcoxon signed-rank test) are denoted by † (vs. ViViT Full) and ‡ (vs. the best competing baseline).
Method
Trainable Params ↓
FLOPs per Step (G) ↓
Full fine-tuning
88.65M (100%)
6402.90
Second-Stage Selective Fine-Tuning
7.09M (7.998%)
4449.00
TABLE III : Trainable parameters and per-step computational cost (FLOPs) of different fine-tuning methods (batch size of 4). Lower values (↓) indicate better efficiency.
Fig. 2 : Weighted F1 against cumulative fine-tuning FLOPs for gallery as a representative case (full fine-tuning vs. SALFT). Analysis is based on the first fold of 5-fold cross-validation. Complete results for all nine games are provided in the GitHub repo.
Fig. 3 : Interpretation example for arousal-increasing class in solid (SALFT-fine-tuned model, first fold).
Fig. 4 : Interpretation example for arousal-decreasing class in solid (SALFT-fine-tuned model, first fold).
Type
Class
Rollout
Grad-CAM
Grad-SAM
Our Method
Positive ↓
Predicted
46.96
47.47
49.82
24.01
Target
45.81
44.54
48.69
23.13
Negative ↑
Predicted
86.86
87.91
84.67
93.91
Target
80.75
84.64
78.65
87.72
TABLE IV : Positive and negative perturbation AUC results (shown as percentages) for the predicted and target classes on the solid test set. Lower (↓) is better for positive perturbation and higher (↑) for negative. Best results are in bold .
Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psychological fact: emotions change as a result of cumulative reactions to consecutive causal events. To bridge this gap, we introduce Dynamic Affective Reasoning, the first large-scale benchmark for viewer-centric affect transitions and causal reasoning over consecutive video events. DAR contains 15,087 videos and 36,908 event-aligned affective segments annotated with 27 emotion categories. Unlike existing video-based emotion datasets, DAR presents a new viewer-centric perspective on fine-grained emotional expressions and transitions, and provides dense, temporally grounded, and causally explicit reasoning chains. Based on DAR, we formally define three challenging tasks: affective segmentation, fine-grained emotion classification, and affective reasoning. Complementing this benchmark, we propose DAR-R1, a two-stage framework that combines supervised fine-tuning with Group Relative Policy Optimization. Experiments across 10+ MLLMs show that DAR-R1 sets a new state-of-the-art for dynamic affective reasoning, in terms of both emotional localization and affective reasoning. Project page: https://github.com/Zhang-Zhiyan/DAR.
Zhiyan Zhang, Peipei Song, Jinpeng Hu +3
University of Science and Technology of China, Hefei, China · Hefei University of Technology, Hefei, China
Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PEFT, it does not imply \textit{compute efficiency}: parameter-sparse methods can still incur full-model training cost per step, and typically need long schedules to reach peak accuracy. We introduce Circuit Fine-Tuning (CFT), a compute-efficient framework that uses circuit discovery---conventionally used to explain trained models---to select modules for fine-tuning before training. Whereas attribution is conventionally formulated against a trained task head, we formulate it against a near-zero-initialized probe head, which isolates the response of the backbone to the target distribution rather than the preferences of a particular classifier. CFT then fine-tunes only the recovered subgraph. CFT needs no learning-rate warmup and reaches peak accuracy in ∼20 epochs on average---versus 44--96 for strong PEFT baselines---yielding 2.3--6.6× fewer training FLOPs and up to 16× less wall-clock time, while adding zero parameters and no inference operations. Experiments across a standard visual transfer benchmark (VTAB-1k), hierarchical backbones (Swin), domain-shifted medical imaging (CBIS-DDSM), and a vision-language model (Gemma-3 on CUB-200) demonstrate the effectiveness of CFT. Code is available at https://github.com/UriKialy/CFT
Uri Z. Kialy, Gil Ben-Artzi
School of Computer Science, Ariel University, Israel
We propose a pre-fine-tuning probing method for Parameter-Efficient Fine-Tuning (PEFT) layer selection, aiming to obtain more stable and higher gains with fewer trainable parameters when adapting large vision--language models (VLMs). Unlike the common practice of applying LoRA and other adapters to all layers at once---where layer selection often relies on heuristic rules---we focus on the vision encoder and directly evaluate the "adaptability'' of each Transformer layer. Specifically, we characterize each layer from two perspectives: (i) the statistical properties of its Q/K/V projection weights (e.g., norms and condition numbers); (ii) robustness under controlled parameter perturbations. We then systematically compare these indicators with the downstream performance gains brought by applying PEFT to a single layer. Across experiments covering seven benchmarks and five PEFT variants, we observe a consistent correlation: layers (or matrices) with larger weight norms and higher condition numbers are usually more robust to perturbations and are more likely to yield larger fine-tuning gains. These results show that distribution-statistics analysis and perturbation tests before fine-tuning can provide practical signals for adaptation-layer selection, thereby maintaining or improving performance while reducing trainable parameters.
Qingtao Xia, Jiahua Bao, Siyao Cheng +1
Faculty of Computing Harbin Institute of Technology Harbin, China