Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.
Figures & tables
Figure 1: Overview of the proposed framework. Multimodal stimulus features are extracted with pretrained models (VideoMAE, HuBERT, Qwen, and BridgeTower), concatenated, and linearly projected to a 512-dimensional sequence. A shared causal Transformer encoder maps the stimulus sequence to an encoded representation, and a masked autoregressive Transformer decoder predicts parcel-wise fMRI time series via cross-attention to this representation. To accommodate inter-subject variability, the decoder includes partially subject-specific components. The model is trained by minimizing L=LMSE+λLcorr . See Methods 2.3 for more details.
Figure 2: Performance of ridge, seq2one, and seq2seq across cortical networks and whole brain. (A–B) Mean Pearson correlations r for ridge regression, seq2one, and seq2seq on Friends (A) and Movies (B), averaged across subjects. (C) Cortical parcellation into seven canonical networks used for network-level analyses. (D–E) Whole-brain prediction maps for seq2seq on Friends and Movies. Network abbreviations: Vis, visual; SomMot, somatomotor; DorsAttn, dorsal attention; SalVentAttn, salience/ventral attention; Limbic, limbic; Cont, control; Default, default mode.
Figure 3: Performance of fully shared, fully individual, and hybrid models across cortical networks and whole brain. (A–B) Mean Pearson correlations r for fully shared, fully individual, and hybrid (shared encoder + partial subject-specific decoder) models on Friends (A) and Movies (B), averaged across subjects. (C–D) Difference maps (hybrid − individual) show advantages in early sensory and attention-related cortices, including visual, somatomotor, and dorsal attention regions. (E–F) Difference maps (hybrid − shared) highlight advantages in higher-order association networks, including default-mode, control, and temporoparietal regions.
Figure 4: Data-efficient adaptation: comparison of individual and hybrid (fine-tuned) models across data scales and cortical networks. (A–B) Whole-brain learning curves showing prediction performance as a function of available target-subject training data. Models were trained or fine-tuned on Bourne (12–96 min) and evaluated on either the same movie (Bourne, in-distribution; A) or a different movie (Wolf, out-of-distribution; B). Solid lines denote hybrid models fine-tuned from multi-subject pretraining, while dashed lines denote individual models trained from scratch. (C–D) Network-level learning curves for individual model evaluated on Bourne (C) and Wolf (D), showing modest gains and high variability across networks. (E–F) Network-level learning curves for hybrid fine-tuned model evaluated on Bourne (E) and Wolf (F), revealing consistently higher accuracy and clear, monotonic improvements across all networks.
Figure 5: Transfer performance of individual, shared, and hybrid models under three representative data-scarcity scenarios. Bars show the mean whole-brain Pearson correlation across subjects, and red dots indicate subject-wise performance (one dot per subject). The hybrid model consistently achieved the highest whole-brain prediction accuracy across all transfer scenarios.
Figure S1: Temporal model selection on the validation split. (A) Whole-brain prediction accuracy as a function of the stimulus context length k for ridge (orange), transformer seq2one (blue), and transformer seq2seq (green). Solid lines denote subject means (n=4) and shaded bands indicate ± s.e.m. The sweep identifies k=15 (ridge), k=30 (seq2one), and k=45 (seq2seq). (B–C) With the stimulus window length fixed at k , ridge (B) and seq2one (C) perform a target sweep over the predicted fMRI frame in 5 TR increments from TR 5 to TR k with a 1 TR sliding-window stride. The optima occur at frame 15 and frame 25, respectively. (D–E) For seq2seq with k=45 , we evaluated output subsequences with a minimum prediction start of ≥5 TRs: (D) the start a varies with the end b fixed. (E) the end b varies with the start a fixed. Both sweeps exhibit a broad performance plateau when the output segment includes late TRs near the window end. Balancing accuracy and robustness, the seq2seq output is fixed to the contiguous sequence TRs 11-40.
Figure S2: (A–B) Difference maps (seq2one − ridge) showing widespread improvements across cortex. (C–D) Difference maps (seq2seq − seq2one) showing further gains in association regions, with only minor negative effects (blue) in limited parcels.
Figure S3: (A–B) Whole-brain prediction maps for hybrid on Friends and Movies.
Friends (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Ridge
0.229
0.261
0.286
0.223
0.250 ± 0.015
Seq2one
0.252
0.296
0.312
0.250
0.278 ± 0.016
Seq2seq
0.268
0.309
0.331
0.264
0.293 ± 0.016
Movies (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Table S1: Per-subject test correlations for temporal models on held-out data.
Friends (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Shared
0.257
0.286
0.289
0.253
0.271 ± 0.009
Individual
0.268
0.309
0.330
0.264
0.293 ± 0.016
Hybrid
0.303
0.313
0.337
0.287
0.310 ± 0.011
Movies (test)
Mean ± s.e.m.
Model
sub-01
sub-02
sub-03
sub-05
Table S2: Per-subject test correlations for shared, individual, and hybrid parameterization strategies on held-out data.
Film
Exp
Overall
Sub1
Sub2
Sub3
Sub5
Bourne
Individual
0.1363
0.1574±0.0282
0.1038±0.0293
0.1680±0.0186
0.1156±0.0198
Hybrid
0.1696
0.1909±0.0339
0.1422±0.0395
0.1934±0.0210
0.1515±0.0315
Shared
0.1638
0.1839±0.0347
0.1361±0.0339
0.1875±0.0157
0.1474±0.0326
Figures
Individual
0.1721
0.1878±0.0187
0.1800±0.0193
0.1770±0.0161
0.1435±0.0270
Hybrid
0.2013
0.2260±0.0227
0.2064±0.0138
0.1981±0.0110
0.1745±0.0295
Shared
0.1888
0.2096±0.0173
0.1934±0.0145
0.1867±0.0153
0.1655±0.0258
Table S3: Film-wise transfer performance in the small-in scenario.
1State Key Laboratory of Multimodal Artificial Intelligence System, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Department of Computer Science, The University of Manchester +1