Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.
Figures & tables
Figure 1: Cross-attention encoding model. (A) A video backbone (V-JEPA 2) encodes each clip into a grid of spatiotemporal tokens. A learned query for each cortical parcel attends over the tokens, and the attended feature predicts the responses of the parcel’s voxels through a voxel-specific linear readout. (B) Constrained routing models, shown as example attention patterns over the token grid with the form of their attention weights. (C) Attention of the joint model for a parcel in right MT, overlaid on selected frames of a test clip of a weightlifter.
Figure 2: Comparing accuracy of joint spatiotemporal attention versus other mappings (V-JEPA 2 ViT-g, layer 32; 102 test clips). (A) Joint model minus the factorized, spatial and mean-pool (“uniform”) models and ridge regression on mean-pooled features in voxelwise signed r2 , averaged over subjects (fsaverage flatmap, unthresholded; each subject in Appendix 6.7 ). (B) Region-mean signed r2 of every routing configuration (T = temporal factor, S = spatial factor; each uniform, fixed or routed), the joint model and ridge regression. Large dots: mean ± SEM over subjects; small dots: individual subjects; grey band: noise ceiling (mean ± SEM). Region means include only voxels with a noise ceiling ≥ 5%. (C) Zoomed inset from (B); shows region means of the mean-pool, spatial, factorized and joint models in each subject (colored lines) and their mean ± SEM (black).
Region
Ridge
Mean pool
Temporal
Spatial
Factorized
Joint
Noise ceiling
All voxels
.174
.180
.180
.187
.186
.191
.34
Early visual
.178
.177
.178
.189
.189
.189
.36
Lateral/face/body
.242
.249
.248
.255
.255
.263
.40
Scene
.118
.115
.114
.116
.115
.119
.24
Parietal
.109
.122
.121
.125
.125
.129
.29
Outside ROIs
.165
.172
.171
.178
.177
.181
.33
Table 1: Region-mean signed r2 on the 102 test clips (V-JEPA 2 ViT-g, layer 32; mean over 10 subjects; voxels with noise ceiling ≥ 5%). The best model in each row is in bold. All routing configurations and the other backbones are given in Table 7 .
Figure 3: Joint attention follows moving content, whereas factorized attention cannot. Attention maps of one parcel in each of two test clips, for the joint model (upper two rows of each block) and the factorized model (lower two rows), over all 16 frames of the clip. Left of each block: the subject, the parcel and its location on the flatmap within its functional ROI. Attention is shown in units of the uniform level 1/N on the same scale for both models: tokens at or below the uniform level are transparent, and color and opacity increase up to 5 times the uniform level.
Figure 4: The attention of the joint model differentiates category-selective regions. (A) Localizing category-selective regions from attention alone in one subject (S8). Left: t of each localizer contrast of z-scored IoU across test clips, per parcel. Middle: the K highest- t parcels ( K = number of ROI parcels) and their overlap with the ROIs. Right: the localizer-defined ROIs (parcels with at least 20% of their voxels inside). (B) Raw IoU of the 92 most-attended tokens (top 1%) with the masks of five categories, for eight category-selective ROIs and for all parcels (mean ± SEM over 10 subjects). (C) The same IoU z-scored across all parcels, so that all parcels are 0 by definition. Outlined bars mark the category that defines each ROI (for LOC, all four foreground categories).
Backbone
Best layer
R2
V-JEPA 2 ViT-g ( Assran et al., 2025 )
32
.134
V-JEPA 2 ViT-H
24
.133
V-JEPA 2 ViT-L
18
.132
V-JEPA 2 ViT-L, fine-tuned on SSv2
18
.125
VideoMAE v2 ViT-g ( Wang et al., 2023b )
20
.109
V-JEPA 2.1 ViT-g [cite]
44
.108
Table 2: Best layer of each backbone in the ridge-regression layer sweep (mean-pooled features; test R2 , mean over all voxels in 10 subjects).
Region
Comparison
Δ signed r2
t(9)
p
pFDR
Joint better
dz
All voxels
factorized
+0.0043±0.0008
5.39
4.4e-04
8.7e-04
10/10
1.71
spatial
+0.0040±0.0007
5.85
2.4e-04
5.8e-04
10/10
1.85
mean pool
+0.0103±0.0017
6.16
1.7e-04
4.5e-04
10/10
1.95
ridge
+0.0162±0.0019
8.52
1.3e-05
1.6e-04
10/10
2.69
Early visual
factorized
−0.0000±0.0010
-0.03
0.976
0.976
5/10
-0.01
spatial
+0.0000±0.0009
0.04
0.966
0.976
4/10
0.01
Table 3: Joint model minus each comparison model in region-mean signed r2 on the 102 test clips (mean ± SEM over 10 subjects), with two-sided paired t -tests across subjects ( df=9 ). p is uncorrected; pFDR is Benjamini–Hochberg corrected over all 24 tests; values below 0.05 are in bold. Positive Δ indicates that the joint model performed better. “Joint better” column indicates how many individual subjects have a numerical advantage for the joint model in each comparison. dz : mean divided by the standard deviation of the per-subject differences.
Figure 5: Voxelwise accuracy of the joint model (V-JEPA 2 ViT-g, layer 32; signed r2 on the 102 test clips; all modeled voxels), averaged over the 10 subjects (top left) and in each subject, on the fsaverage flatmap. Values outside 0–0.6 are shown at the ends of the color scale.
Figure 6: Joint model minus the factorized, spatial and mean-pool (“uniform”) models and ridge regression on mean-pooled features in voxelwise signed r2 , in each subject (rows), on the flatmap of Fig. 2 A. All panels share one color scale; differences beyond ±0.08 are shown at its ends.
Spatial − mean pool
Spatial − ridge
Region
Δ
t(9)
pFDR
n+
Δ
t(9)
pFDR
n+
V-JEPA 2 ViT-g, layer 32 of 40
All voxels
+6.3
5.07
8.9e-4
9/10
+12.2
8.26
3.3e-5
10/10
Early visual
+12.2
5.58
4.7e-4
10/10
+10.8
4.24
0.003
9/10
Lateral/face/body
+5.9
3.97
0.004
9/10
+13.0
8.73
2.5e-5
10/10
Scene
+0.5
0.37
0.734
6/10
−2.2
-1.11
0.313
5/10
Table 4: Spatial model minus the mean-pool model and minus ridge regression on mean-pooled features, in region-mean signed r2 on the 102 test clips ( Δ in units of 10−3 , mean over 10 subjects; voxels with noise ceiling ≥ 5%). t(9) : two-sided paired t -test across subjects. pFDR : Benjamini–Hochberg corrected over all 96 tests in the table; values below 0.05 are in bold. n+ : subjects in which the spatial model is better.
Figure 7: Contributions of the readout components (V-JEPA 2 ViT-g, layer 32; region-mean signed r2 on the 102 test clips; voxels with noise ceiling ≥ 5%). From left to right: ridge regression on mean-pooled features, a linear readout of the mean-pooled features trained by gradient descent, the mean-pool model without and with the FFN (“uniform”), and the joint model without and with the FFN. Open markers: models without an FFN. Colored lines: individual subjects; markers: mean ± SEM over subjects.
Contrast
Region
Δ signed r2
t(9)
p
pFDR
n+
Linear readout (SGD) − ridge
All voxels
−0.0045±0.0006
-7.11
5.6e-05
4.2e-04
0/10
Early visual
−0.0053±0.0010
-5.37
4.5e-04
0.001
0/10
Lateral/face/body
−0.0069±0.0012
-5.88
2.3e-04
0.001
0/10
Scene
−0.0055±0.0010
-5.26
5.2e-04
0.001
0/10
Parietal
−0.0021±0.0008
-2.53
0.032
0.039
1/10
Outside ROIs
−0.0037±0.0006
-6.71
8.8e-05
4.4e-04
0/10
Table 5: Each model of Fig. 7 minus the model it extends, in region-mean signed r2 (mean ± SEM over 10 subjects), with two-sided paired t -tests across subjects ( df=9 ). p is uncorrected; pFDR is Benjamini–Hochberg corrected over all 30 tests; values below 0.05 are in bold. n+ : subjects in which the first model is better.
Figure 8: Region-mean signed r2 of every routing configuration, the joint model and ridge regression in each functional ROI (V-JEPA 2 ViT-g, layer 32; 102 test clips; voxels with noise ceiling ≥ 5%), grouped by region as in Fig. 2 . Large dots: mean ± SEM over subjects; small dots: individual subjects; grey band: noise ceiling (mean ± SEM).
Joint − factorized
Joint − spatial
Joint − mean pool
Joint − ridge
ROI
Δ
t(9)
pFDR
n+
Δ
t(9)
pFDR
n+
Δ
t(9)
pFDR
n+
Δ
t(9)
pFDR
n+
Early visual
V1v
−1.7
-1.26
0.281
5
−1.7
-1.53
0.203
3
+8.3
3.38
0.017
8
+3.8
1.40
0.240
6
V1d
−0.8
-0.59
0.590
4
−1.9
-2.26
0.072
3
+13.9
4.30
0.006
9
+12.0
3.63
0.013
8
V2v
−0.5
-0.40
0.717
4
−1.5
-1.84
0.130
4
+5.5
2.18
0.080
6
+2.5
0.95
0.402
7
V2d
−1.8
-1.23
0.290
4
−0.7
-0.68
0.546
4
+18.3
5.92
0.001
9
+18.0
5.38
0.002
9
Table 6: Joint model minus each comparison model in ROI-mean signed r2 on the 102 test clips ( Δ in units of 10−3 , mean over 10 subjects; voxels with noise ceiling ≥ 5%). t(9) : two-sided paired t -test across subjects. pFDR : Benjamini–Hochberg corrected over all 92 tests in the table; values below 0.05 are in bold. n+ : subjects (of 10) in which the joint model is better.
Temporal uniform
Temporal fixed
Temporal routed
Region
Ridge
U
F
R
U
F
R
U
F
R
Joint
V-JEPA 2 ViT-g ( Assran et al., 2025 ) , layer 32 of 40
All voxels
.174
.180
.181
.187
.181
.181
.187
.180
.180
.186
.191
Early visual
.178
.177
.178
.189
.177
.177
.190
.178
.179
.189
.189
Lateral/face/body
.242
.249
.250
.255
.250
.250
.255
.248
.249
.255
.263
Scene
.118
.115
.117
.116
.117
.117
.116
.114
.114
.115
.119
Table 7: Region-mean signed r2 on the 102 test clips (mean over 10 subjects; voxels with noise ceiling ≥ 5%) for every routing configuration and backbone. The factorized models are labelled by their temporal factor (column groups) and spatial factor (U = uniform, F = fixed, R = routed), where a fixed factor is a learned, stimulus-independent distribution for each parcel. Temporal uniform with spatial uniform is the mean-pool model, temporal uniform with spatial routed is the spatial model, temporal routed with spatial uniform is the temporal model, and temporal routed with spatial routed is the factorized model. Ridge is ridge regression on the mean-pooled features of the same layer. The best model in each row is in bold.
Figure 9: Joint model minus the factorized, spatial and mean-pool (“uniform”) models and ridge regression on mean-pooled features of the same layer, in region-mean signed r2 on the 102 test clips, for every backbone and layer of the routing sweep (mean ± SEM over 10 subjects; voxels with noise ceiling ≥ 5%).
Figure 10: Fig. 2 B,C for V-JEPA 2 ViT-g, layer 4 of 40: region-mean signed r2 of every routing configuration, the joint model and ridge regression (top), and of the mean-pool, spatial, factorized and joint models in each subject (bottom).
Figure 11: As Fig. 10 , for VideoMAE v2 ViT-g, layer 20 of 40.
Figure 12: As Fig. 10 , for VideoMAE v2 ViT-g, layer 4 of 40.
Figure 13: As Fig. 10 , for DINOv2 ViT-B/14, block 12 of 12.
Figure 14: As Fig. 10 , for DINOv2 ViT-B/14, block 2 of 12.
Figure 15: As Fig. 10 , for CLIP RN50, stage 4 of 4.
Figure 16: As Fig. 10 , for CLIP RN50, stage 2 of 4.
Figure 17: Further examples of attention maps in the format of Fig. 3 : joint model (upper two rows of each block) and factorized model (lower two rows) over all 16 frames of each clip, in units of the uniform level 1/N , from transparent at or below the uniform level to saturated at 5 times the uniform level.
Figure 18: Further examples of attention maps, as in Fig. 17 .
Figure 19: Further examples of attention maps, as in Fig. 17 .
Correct masks
Mismatched masks
Correct − mismatched
Localizer
Chance
Overlap
t(9)
pFDR
AUC
Subjects
Overlap
t(9)
pFDR
Subjects
t(9)
p
pFDR
Face
0.09
0.19 (0.11)
2.81
0.021
0.61
5/10
0.10 (0.05)
0.25
0.81
0/10
3.13
0.012
0.036
Body
0.02
0.14 (0.09)
4.02
0.005
0.76
3/10
0.08 (0.09)
2.00
0.11
3/10
1.70
0.12
0.18
Scene
0.10
0.28 (0.12)
4.85
0.003
0.77
8/10
0.26 (0.14)
3.74
0.014
8/10
1.11
0.30
0.30
Table 8: Localization of category-selective ROIs from z-scored IoU contrasts with the correct masks and with the masks of a different clip (joint model, 10 subjects, 102 test clips; face: FFA, OFA and STS; body: EBA; scene: PPA, RSC and TOS). Overlap: fraction of the top- K parcels that belong to the ROIs, mean (SD) over subjects; chance is the fraction of scored parcels in the ROIs. t(9) and pFDR : two-sided one-sample t -test of overlap minus chance across subjects, Benjamini–Hochberg corrected over the three localizers within each condition. Subjects: number with a one-sided hypergeometric p<0.05 (uncorrected). Correct − mismatched: two-sided paired t -test of the overlaps across subjects, FDR-corrected over the three localizers.