Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition
Organizations: The University of Tokyo Tokyo, Japan
Abstract
We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training performers, an equal-weight logit-mean ensemble reaches 36.80 +/- 4.00% per-fold Macro-F1, a protocol-matched +11.07 pp (+43% relative) over the same-split reproduced baseline. Our central contribution is a tested explanation suite: for a strong ensemble member, part-masking and counterfactual edits show (rather than assert) that its decisions depend on motion-grounded body-region evidence, and this region saliency aligns with rule-based Laban Movement Analysis (LMA) attributes far more than with classical kinematics: region-level saliency-LMA Spearman rho = +0.500 versus +0.033, roughly 15x, and the alignment holds for the submitted 11-way ensemble itself at rho = +0.517; the audit is post hoc and needs no retraining. The same suite faithfully reports a negative: within-window temporal saliency is diffuse rather than localized. On the hidden challenge test set the submitted ensemble scored 37.23 % Macro-F1 and received the Best Performance Award of the MMAC Challenge 2026 (Human score is 39 %). Code is available at https://github.com/nawta/diema-challenge and the presentation at https://nawta.github.io/mmac2026/.
Figures & tables
| Approach | Configuration | Outcome ( F1) | Methodological lesson |
|---|---|---|---|
| Supervised contrastive head | pairwise SupCon, =0.05, batch 128 | -11.53 pp, significant | 12 classes at batch 128 give too few positives per class, suggesting a per-class center loss may fit better at this scale |
| Mixture-of-experts fusion | K=4, hand-crafted gradient-free routing | collapses to 6.55% F1 | a shared head with random init and 21-D routing is associated with backbone collapse, suggesting a feature adapter, learned routing, and zero-init are needed |
| Scenario-text alignment | global InfoNCE, 3-fold | -1.42 pp | the projection collapses to performer identity (cross-group retrieval 10.6%) |
| Part-rationale alignment | cosine text alignment, 3-fold | -0.64 pp, CI [-1.30, -0.09] | the text bottleneck is associated with an F1 ceiling of 9–12%; shortcut-breaking works but yields no F1 lift |
| VLM zero-shot distillation | 12-class, a recent VLM, preliminary | failed in preliminary zero-shot ranking | the true class ranked last (12/12); a vision-language model did not infer emotion reliably from faceless, scene-free skeleton renderings |
| Contact-marker late fusion | C3D 22-D contact stats, 3-fold | -0.59 pp | the contact features carry a 69.8% country signal, so some performers go out of distribution under LPO, suggesting a domain-invariance objective is needed |
| System | Trainable params (M) ‡ | Macro-F1 (mean SD) | Macro-F1 95% CI | Accuracy (mean SD) |
|---|---|---|---|---|
| STGCN++ official baseline | 1.40 | 25.21 4.49 | n/a | 27.11 3.67 |
| STGCN++ reproduced | 1.41 | 25.73 4.03 | [25.37, 27.23] | 27.54 3.92 |
| Best single (Region-Aware) | 1.03 | 30.05 3.48 | [29.32, 31.15] | 30.87 3.43 |
| 7-way ensemble | 12.98 | 33.86 2.92 | [33.01, 35.04] | 34.68 3.00 |
| 11-way logit-mean (submitted) | 12.98 | 36.80 4.00 | [35.90, 37.94] | 37.40 4.06 |
| Component | Metric | Value | 95% CI / null | What it shows |
|---|---|---|---|---|
| Part-masking | Faithfulness AUC gap (important-reverse), zero-mask | +0.124 0.031 | 10-fold SD; positive in all folds | Masking important parts lowers F1: attributions are faithful |
| Temporal saliency | AUC gap (salient-reverse), negative result | +0.002 ( 0) | entropy 98.7% of log T ( uniform) † | Frame importance is diffuse (a reported negative) |
| Stability | Spearman vs =0 ref at =0.02 | +0.983 0.039 | 10-fold SD; top part never flips | Attribution ranking is stable under input noise |
| Counterfactual | Most-disruptive edit p(true): head+freeze | -0.0164 | mean over 864 val 10 folds | Perturbing motion changes predictions: motion-dependent |
| Narrator grounding | Unsupported claims in audited cards | 0/50 cards | deterministic audit | No unsupported claims in the 50 audited cards |
| F1 cost | F1: baseline with post-hoc saliency/LMA export | 0.0 pp | post-hoc; no retraining | Explanations add no accuracy cost |
| Member | Family | Macro-F1 (pp) |
|---|---|---|
| MAMP-NTU60 | External (frozen) | |
| SkateFormer | Attention | |
| MAMP-NTU120 | External (frozen) | |
| Region-Aware ConvTr | Graph-conv. | |
| C3D-marker-stats | External (frozen) | |
| MotionBERT-Lite | External (frozen) |
| Member | Family | All | JP | TW | Excl. earliest JP |
|---|---|---|---|---|---|
| MAMP-NTU60 | External (frozen) | ||||
| SkateFormer | Attention | ||||
| MAMP-NTU120 | External (frozen) | ||||
| Region-Aware ConvTr | Graph-conv. | ||||
| C3D-marker-stats | External (frozen) | ||||
| MotionBERT-Lite | External (frozen) |
| Laban axis | What it captures | Attributes (8 per axis) |
|---|---|---|
| Body | posture and bilateral form | head bow, trunk lean, arm openness, left–right speed asymmetry, left–right position asymmetry, stillness ratio, head lateral tilt, shoulder drop |
| Effort | motion quality (time/weight/flow) | suddenness, sustainedness, strong energy, light energy, bound flow, free flow, intensity peak, intensity variance |
| Shape | body volume and its change | body volume, shoulder width, head height, contraction change, rise/sink, spread change, arm enclosure, advance/recede |
| Space | trajectory and locomotion | directness, path curvature, horizontal root displacement, vertical root displacement, lateral sway, cumulative turn, dominant direction, locomotion ratio |
| Approach | Configuration | Outcome ( F1) | Methodological lesson |
|---|---|---|---|
| Supervised contrastive head | pairwise SupCon, =0.05, batch 128 | -11.53 pp, significant | 12 classes at batch 128 give too few positives per class, suggesting a per-class center loss may fit better at this scale |
| Mixture-of-experts fusion | K=4, hand-crafted gradient-free routing | collapses to 6.55% F1 | a shared head with random init and 21-D routing is associated with backbone collapse, suggesting a feature adapter, learned routing, and zero-init are needed |
| Scenario-text alignment | global InfoNCE, 3-fold | -1.42 pp | the projection collapses to performer identity (cross-group retrieval 10.6%) |
| Part-rationale alignment | cosine text alignment, 3-fold | -0.64 pp, CI [-1.30, -0.09] | the text bottleneck is associated with an F1 ceiling of 9–12%; shortcut-breaking works but yields no F1 lift |
| VLM zero-shot distillation | 12-class, a recent VLM, preliminary | failed in preliminary zero-shot ranking | the true class ranked last (12/12); a vision-language model did not infer emotion reliably from faceless, scene-free skeleton renderings |
| Contact-marker late fusion | C3D 22-D contact stats, 3-fold | -0.59 pp | the contact features carry a 69.8% country signal, so some performers go out of distribution under LPO, suggesting a domain-invariance objective is needed |