Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.
Figures & tables
Figure 1: Overview of the proposed framework. Multimodal stimulus features are extracted with pretrained models (VideoMAE, HuBERT, Qwen, and BridgeTower), concatenated, and linearly projected to a 512-dimensional sequence. A shared causal Transformer encoder maps the stimulus sequence to an encoded representation, and a masked autoregressive Transformer decoder predicts parcel-wise fMRI time series via cross-attention to this representation. To accommodate inter-subject variability, the decoder includes partially subject-specific components. The model is trained by minimizing L=LMSE+λLcorr . See Methods 2.3 for more details.
Figure 2: Performance of ridge, seq2one, and seq2seq across cortical networks and whole brain. (A–B) Mean Pearson correlations r for ridge regression, seq2one, and seq2seq on Friends (A) and Movies (B), averaged across subjects. (C) Cortical parcellation into seven canonical networks used for network-level analyses. (D–E) Whole-brain prediction maps for seq2seq on Friends and Movies. Network abbreviations: Vis, visual; SomMot, somatomotor; DorsAttn, dorsal attention; SalVentAttn, salience/ventral attention; Limbic, limbic; Cont, control; Default, default mode.
Figure 3: Performance of fully shared, fully individual, and hybrid models across cortical networks and whole brain. (A–B) Mean Pearson correlations r for fully shared, fully individual, and hybrid (shared encoder + partial subject-specific decoder) models on Friends (A) and Movies (B), averaged across subjects. (C–D) Difference maps (hybrid − individual) show advantages in early sensory and attention-related cortices, including visual, somatomotor, and dorsal attention regions. (E–F) Difference maps (hybrid − shared) highlight advantages in higher-order association networks, including default-mode, control, and temporoparietal regions.
Figure 4: Data-efficient adaptation: comparison of individual and hybrid (fine-tuned) models across data scales and cortical networks. (A–B) Whole-brain learning curves showing prediction performance as a function of available target-subject training data. Models were trained or fine-tuned on Bourne (12–96 min) and evaluated on either the same movie (Bourne, in-distribution; A) or a different movie (Wolf, out-of-distribution; B). Solid lines denote hybrid models fine-tuned from multi-subject pretraining, while dashed lines denote individual models trained from scratch. (C–D) Network-level learning curves for individual model evaluated on Bourne (C) and Wolf (D), showing modest gains and high variability across networks. (E–F) Network-level learning curves for hybrid fine-tuned model evaluated on Bourne (E) and Wolf (F), revealing consistently higher accuracy and clear, monotonic improvements across all networks.
Figure 5: Transfer performance of individual, shared, and hybrid models under three representative data-scarcity scenarios. Bars show the mean whole-brain Pearson correlation across subjects, and red dots indicate subject-wise performance (one dot per subject). The hybrid model consistently achieved the highest whole-brain prediction accuracy across all transfer scenarios.
Figure S1: Temporal model selection on the validation split. (A) Whole-brain prediction accuracy as a function of the stimulus context length k for ridge (orange), transformer seq2one (blue), and transformer seq2seq (green). Solid lines denote subject means (n=4) and shaded bands indicate ± s.e.m. The sweep identifies k=15 (ridge), k=30 (seq2one), and k=45 (seq2seq). (B–C) With the stimulus window length fixed at k , ridge (B) and seq2one (C) perform a target sweep over the predicted fMRI frame in 5 TR increments from TR 5 to TR k with a 1 TR sliding-window stride. The optima occur at frame 15 and frame 25, respectively. (D–E) For seq2seq with k=45 , we evaluated output subsequences with a minimum prediction start of ≥5 TRs: (D) the start a varies with the end b fixed. (E) the end b varies with the start a fixed. Both sweeps exhibit a broad performance plateau when the output segment includes late TRs near the window end. Balancing accuracy and robustness, the seq2seq output is fixed to the contiguous sequence TRs 11-40.
Figure S2: (A–B) Difference maps (seq2one − ridge) showing widespread improvements across cortex. (C–D) Difference maps (seq2seq − seq2one) showing further gains in association regions, with only minor negative effects (blue) in limited parcels.
Figure S3: (A–B) Whole-brain prediction maps for hybrid on Friends and Movies.
Friends (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Ridge
0.229
0.261
0.286
0.223
0.250 ± 0.015
Seq2one
0.252
0.296
0.312
0.250
0.278 ± 0.016
Seq2seq
0.268
0.309
0.331
0.264
0.293 ± 0.016
Movies (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Table S1: Per-subject test correlations for temporal models on held-out data.
Friends (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Shared
0.257
0.286
0.289
0.253
0.271 ± 0.009
Individual
0.268
0.309
0.330
0.264
0.293 ± 0.016
Hybrid
0.303
0.313
0.337
0.287
0.310 ± 0.011
Movies (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Table S2: Per-subject test correlations for shared, individual, and hybrid parameterization strategies on held-out data.
Film
Exp
Overall
Sub1
Sub2
Sub3
Sub5
Bourne
Individual
0.1363
0.1574±0.0282
0.1038±0.0293
0.1680±0.0186
0.1156±0.0198
Hybrid
0.1696
0.1909±0.0339
0.1422±0.0395
0.1934±0.0210
0.1515±0.0315
Shared
0.1638
0.1839±0.0347
0.1361±0.0339
0.1875±0.0157
0.1474±0.0326
Figures
Individual
0.1721
0.1878±0.0187
0.1800±0.0193
0.1770±0.0161
0.1435±0.0270
Hybrid
0.2013
0.2260±0.0227
0.2064±0.0138
0.1981±0.0110
0.1745±0.0295
Shared
0.1888
0.2096±0.0173
0.1934±0.0145
0.1867±0.0153
0.1655±0.0258
Table S3: Film-wise transfer performance in the small-in scenario.
Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework for predicting whole-brain fMRI responses to naturalistic audiovisual stimuli. MIRAGE achieves state-of-the-art performance via a native multimodal backbone and adaptive feature gating across layers. These representations are then combined with a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. Controlled comparisons show that natively multimodal features consistently outperform post-hoc aggregation of independent unimodal features, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable to interpret the modality-specific gating profile over the backbone, and each modality traces a distinct anatomical pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a generalizable, interpretable, and accurate approach for whole-brain encoding.
Abdulkadir Gokce, Badr AlKhamissi, Martin Schrimpf
Cognitive neuroscience is fragmented into specialized models, each tailored to specific experimental paradigms, hence preventing a unified model of cognition in the human brain. Here, we introduce TRIBE v2, a tri-modal (video, audio and language) foundation model capable of predicting human brain activity in a variety of naturalistic and experimental conditions. Leveraging a unified dataset of over 1,000 hours of fMRI across 720 subjects, we demonstrate that our model accurately predicts high-resolution brain responses for novel stimuli, tasks and subjects, superseding traditional linear encoding models, delivering several-fold improvements in accuracy. Critically, TRIBE v2 enables in silico experimentation: tested on seminal visual and neuro-linguistic paradigms, it recovers a variety of results established by decades of empirical research. Finally, by extracting interpretable latent features, TRIBE v2 reveals the fine-grained topography of multisensory integration. These results establish artificial intelligence as a unifying framework for exploring the functional organization of the human brain.
Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal processing. Inspired by the brain's associative mechanisms, where viewing an image can evoke related sounds and linguistic representations, we propose a unified framework that leverages Multimodal Large Language Models (MLLMs) to align brain signals with a shared semantic space encompassing text, images, and audio. A router module dynamically selects and fuses modality-specific brain features according to the characteristics of each stimulus. Experiments on various fMRI datasets with textual, visual, and auditory stimuli demonstrate state-of-the-art performance, achieving an 8.48% improvement on the most commonly used benchmark. We further extend our framework to EEG and MEG data, demonstrating flexibility and robustness across varying temporal and spatial resolutions. To our knowledge, this is the first unified BCI architecture capable of robustly decoding multimodal brain activity across diverse brain signals and stimulus types, offering a flexible solution for real-world applications.
Chunyu Ye, Yunhao Zhang, Jingyuan Sun +3
1State Key Laboratory of Multimodal Artificial Intelligence System, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Department of Computer Science, The University of Manchester +1