cs.LGOct 7, 2026

MovieSTAGE: Scene, Transition, and Global Encoding for Movie-fMRI ADHD Classification

Authors: Boseong Kim, Haejun Chung, Ikbeom Jang

Organizations: Hanyang University, Seoul, Republic of Korea · Hankuk University of Foreign Studies, Yongin, Republic of Korea

Abstract

Naturalistic movie-fMRI provides a shared, temporally structured probe of brain dynamics, yet predictive models commonly rely on whole-run functional connectivity (FC) or temporally generic representations that are not aligned with narrative events. We introduce MovieSTAGE (Scene, Transition, and Global Encoding), a multiscale framework that combines hypergraph-structured FC-profile organization within scenes, unsigned FC-profile differences across adjacent scenes, and whole-movie FC. We evaluated 260 participants from the CMI-HBN Despicable Me cohort on case-control, ADHD-subtype, and three-class classification using 10 repetitions of stratified five-fold cross-validation, complete out-of-fold (OOF) predictions, and paired subject-cluster bootstrap and permutation tests. MovieSTAGE achieved AUROCs of 0.69, 0.73, and 0.75 and balanced accuracies of 67.6%, 69.8%, and 58.3%, respectively, yielding the highest mean point estimates among the evaluated methods. On the three-class task, the full model outperformed all two-branch variants, the HGNN scene encoder outperformed MLP, GAT, and BNT alternatives under matched settings, and the human-annotated partition outperformed duration-matched random and fixed-count GSBS controls. These controlled results support incremental predictive value from event-aligned scene and transition representations when combined with whole-movie FC in this cohort. Post-hoc model-derived analyses generated network-level hypotheses involving frontoparietal and default-mode systems.

Figures & tables

Explore similar work

Apr 14, 2026cs.CV

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn generalized representations across diverse brain states. We present Brain-DiT, a universal multi-state fMRI foundation model pretrained on 349,898 sessions from 24 datasets spanning resting, task, naturalistic, disease, and sleep states. Unlike prior fMRI foundation models that rely on masked reconstruction in the raw-signal space or a latent space, Brain-DiT adopts metadata-conditioned diffusion pretraining with a Diffusion Transformer (DiT), enabling the model to learn multi-scale representations that capture both fine-grained functional structure and global semantics. Across extensive evaluations and ablations on 7 downstream tasks, we find consistent evidence that diffusion-based generative pretraining is a stronger proxy than reconstruction or alignment, with metadata-conditioned pretraining further improving downstream performance by disentangling intrinsic neural dynamics from population-level variability. We also observe that downstream tasks exhibit distinct preferences for representational scale: ADNI classification benefits more from global semantic representations, whereas age/sex prediction comparatively relies more on fine-grained local structure.
Oct 5, 2026cs.LG

CoHyFuse: Condition-wise Hypergraph Fusion with Global Connectome in Task-fMRI

Task-fMRI connectomes reveal state-dependent neural reconfigurations, yet conventional methods marginalize these signals by aggregating distinct conditions into static pairwise graphs, thereby obscuring condition-specific multi-ROI organization. We introduce CoHyFuse, a condition-aware ROI-centered hypergraph framework that constructs a task-state-specific incidence matrix from condition-wise functional connectivity (FC)-profile embeddings, allowing the same ROI to form different multi-ROI hyperedges across task phases. Condition-specific neighborhood sizes KqK_q further adapt the hyperedge scale to each task state, and the resulting condition embeddings are fused with a complementary whole-session FC branch for prediction. In the AABC cohort (N=1,074), CoHyFuse achieved the best mean out-of-fold predictive performance among evaluated baselines on FACENAME Fluid Cognition Composite (FCC) prediction (7.83±\pm0.10 MAE, 0.439±\pm0.026 R2R^2) and VISMOTOR age prediction (7.52±\pm0.37 MAE, 0.592±\pm0.022 R2R^2). In an auxiliary CMI-HBN attention-deficit/hyperactivity disorder (ADHD) classification benchmark (N=223), CoHyFuse obtained 72.0±\pm2.1% macro-AUC and 74.2±\pm2.9% accuracy. Ablation studies support the contributions of condition-wise incidence construction and dual-view fusion, suggesting that state-resolved ROI-set structure provides complementary predictive information beyond whole-session FC alone. Occlusion analysis identifies the Distraction condition as the primary driver of model prediction, pointing toward the Salience/Ventral Attention Network (SAN)--FrontoParietal Network (FPN) and within-SAN hyperedge-defined ROI-set motifs as candidate model-relevant patterns. This framework provides an interpretable, state-resolved view of the connectome for downstream cohort analysis.
May 28, 2026cs.LG

MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework for predicting whole-brain fMRI responses to naturalistic audiovisual stimuli. MIRAGE achieves state-of-the-art performance via a native multimodal backbone and adaptive feature gating across layers. These representations are then combined with a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. Controlled comparisons show that natively multimodal features consistently outperform post-hoc aggregation of independent unimodal features, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable to interpret the modality-specific gating profile over the backbone, and each modality traces a distinct anatomical pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a generalizable, interpretable, and accurate approach for whole-brain encoding.