cs.CVJul 24, 2025

A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli

Authors: Qianyi He, Monica D. Rosenberg, Yuan Chang Leong

Organizations: Data Science Institute University of Chicago, USA · Department of Psychology, Neuroscience Institute University of Chicago, USA

Abstract

Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.

Figures & tables

Explore similar work

May 28, 2026cs.LG

MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework for predicting whole-brain fMRI responses to naturalistic audiovisual stimuli. MIRAGE achieves state-of-the-art performance via a native multimodal backbone and adaptive feature gating across layers. These representations are then combined with a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. Controlled comparisons show that natively multimodal features consistently outperform post-hoc aggregation of independent unimodal features, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable to interpret the modality-specific gating profile over the backbone, and each modality traces a distinct anatomical pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a generalizable, interpretable, and accurate approach for whole-brain encoding.
May 5, 2026q-bio.NC

A foundation model of vision, audition, and language for in-silico neuroscience

Cognitive neuroscience is fragmented into specialized models, each tailored to specific experimental paradigms, hence preventing a unified model of cognition in the human brain. Here, we introduce TRIBE v2, a tri-modal (video, audio and language) foundation model capable of predicting human brain activity in a variety of naturalistic and experimental conditions. Leveraging a unified dataset of over 1,000 hours of fMRI across 720 subjects, we demonstrate that our model accurately predicts high-resolution brain responses for novel stimuli, tasks and subjects, superseding traditional linear encoding models, delivering several-fold improvements in accuracy. Critically, TRIBE v2 enables in silico experimentation: tested on seminal visual and neuro-linguistic paradigms, it recovers a variety of results established by decades of empirical research. Finally, by extracting interpretable latent features, TRIBE v2 reveals the fine-grained topography of multisensory integration. These results establish artificial intelligence as a unifying framework for exploring the functional organization of the human brain.
May 15, 2025cs.CL

Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing

Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal processing. Inspired by the brain's associative mechanisms, where viewing an image can evoke related sounds and linguistic representations, we propose a unified framework that leverages Multimodal Large Language Models (MLLMs) to align brain signals with a shared semantic space encompassing text, images, and audio. A router module dynamically selects and fuses modality-specific brain features according to the characteristics of each stimulus. Experiments on various fMRI datasets with textual, visual, and auditory stimuli demonstrate state-of-the-art performance, achieving an 8.48% improvement on the most commonly used benchmark. We further extend our framework to EEG and MEG data, demonstrating flexibility and robustness across varying temporal and spatial resolutions. To our knowledge, this is the first unified BCI architecture capable of robustly decoding multimodal brain activity across diverse brain signals and stimulus types, offering a flexible solution for real-world applications.