cs.CVJul 24, 2025

A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli

Authors: Qianyi He, Monica D. Rosenberg, Yuan Chang Leong

Organizations: Data Science Institute University of Chicago, USA · Department of Psychology, Neuroscience Institute University of Chicago, USA

Abstract

Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.

Figures & tables

Explore similar work

CardsList
  1. MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

    May 28, 2026Abdulkadir Gokce, Badr AlKhamissi, Martin SchrimpfFunctional Magnetic Resonance ImagingUnimodal Metrics

  2. A foundation model of vision, audition, and language for in-silico neuroscience

    May 5, 2026Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit +5NeuroscienceMultimodal Foundation Model

  3. Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing

    May 15, 2025Chunyu Ye, Yunhao Zhang, Jingyuan Sun +3Brain-Computer InterfaceMultimodal Alignment