ReGraph: A Computational Account of Emergent Generalization in the "what" and "where" Dual Visual Streams
Authors: Hyewon Kang, Jungmin Lee, Ilgyu Lee, Seok-Jun Hong
Organizations: Department of Intelligent Precision Healthcare Convergence, Sungkyunkwan University, Suwon, Republic of Korea · Center for Neuroscience Imaging Research, Institute for Basic Science (IBS), Suwon, Republic of Korea · Department of Biomedical Engineering, Sungkyunkwan University, Suwon, Republic of Korea
Where generalization capacity--the ability to extract context-invariant relational structures--first emerges remains a central question in AI and neuroscience. The foundation for this capacity lies upstream of the hippocampus, within the entorhinal cortex, where parallel pathways dissociate relational structure in the medial entorhinal cortex (MEC) from sensory content in the lateral entorhinal cortex. However, as Eichenbaum argued, such factorization likely originates earlier, driven by the segregation of the dorsal ('where') and ventral ('what') visual streams. Supporting this, grid-like firing patterns--a signature of MEC (context-invariant codes)--also appear in preceding neocortical regions along the dorsal pathway. Yet, how such representations are computationally formed along upstream pathways remains unknown. To investigate this in silico, we developed ReGraph, a recurrent dual-stream graph model with biological inductive biases, including retina-driven stream-specialized encoding, dorsal-to-ventral modulation, and dynamic lateral connectivity. Trained on the action benchmark Something-Something V2, ReGraph revealed a pathway-specific emergence of relational mapping: context-invariant codes and grid-like spatial bases uniquely co-emerged along the extended dorsal stream. In contrast, their absence in single-stream, unmodulated variants, and standard baselines implies that these inductive biases are prerequisites for relational structures. Crucially, our post-hoc analyses demonstrated that these grid-like bases serve as reusable routing templates for information processing via lateral connectivity. Together, our findings provide a computational account that generalization may not be a faculty that emerges abruptly within a dedicated region, but a property that already takes shape as sensory information is parsed into factorized streams of hierarchical visual processing.
Figures & tables
Figure 1: From retina to dual visual streams. Left : The retinal M/P dichotomy drives the dorsal and ventral pathways, each comprising primary and extended stages. Grid-like patterns in the extended dorsal pathway motivate a hypothesis that the precursor of relational structures may emerge along this pathway. The primary dorsal/ventral pathways preferentially encode spatiotemporal dynamics and object details, respectively. Right: Schematics of three connectivity principles defined in § 2.2 .
Figure 2: The ReGraph architecture. A. Input is tokenized into M/P-driven asymmetric views (Dorsal: high temporal/low spatial; Ventral: low temporal/high spatial) and processed via specialized pretrained networks for spatiotemporal process (depth, optical flow, saliency) and hierarchical object-oriented processing. B. Stacked layers model cortical hierarchies, with top-down modulation routing the dorsal signals to the ventral layers. C. Self-attention acts as a dynamic graph operator, followed by a GRU-based recurrent update to regulate temporal flow.
Ablation Type
Variant
TD/TV
ND/NV
V-RGB
D→V
Tokens
Top-1(%; ID)
Single-Stream
Dorsal-Only
12 / –
196 / –
–
–
2,352
71.69
References
Ventral-Only
– / 3
– / 784
Yes
–
2,352
60.38
Dorsal–Ventral
(a) Trait-Symmetric
5 / 5
484 / 484
No
×
4,840
69.78
Biological Trait
(b) Dorsal-trait Only
8 / 2
484 / 484
No
✓
4,840
72.42
Asymmetry
(c) Ventral-trait Only
5 / 5
196 / 784
Yes
×
4,900
70.43
(d) Fully Asymmetric (ReGraph)
12 / 3
196 / 784
Yes
✓
4,704
74.57
Table 1: Ablation over design choices. All conditions share a fixed token budget ( TDND+TVNV≈4,800 ). Temporal ( TD>TV ) and spatial ( NV>ND ) imbalances realize the proposed dorsal and ventral traits, respectively. ID: In-distribution evaluation.
Model
Stream
L1 (OOD [Drop])
L2 (OOD [Drop])
L3 (OOD [Drop])
L4 (OOD [Drop])
Avg (OOD [Drop])
Dorsal-Only
Dorsal
72.8 ± 3.6 [ − 25.6]
69.0 ± 3.7 [ − 30.7]
76.7 ± 1.9 [ − 23.2]
93.6 ± 1.6 [ − 6.4]
78.0 [ − 21.5]
Unmodulated
Dorsal
86.8 ± 2.7 [ − 12.7]
82.9 ± 2.1 [ − 17.1]
81.5 ± 2.1 [ − 18.5]
96.9 ± 0.4 [ − 3.1]
87.0 [ − 12.9]
Ventral
52.0 ± 3.8 [ − 44.2]
56.9 ± 2.7 [ − 42.1]
59.2 ± 4.5 [ − 40.2]
65.4 ± 1.7 [ − 33.4]
58.4 [ − 40.0]
ReGraph
Dorsal
82.0 ± 2.7 [ − 16.9]
81.5 ± 3.2 [ − 18.5]
78.5 ± 3.4 [ − 21.3]
97.4 ± 0.6 [ − 2.6]
84.9 [ − 14.8]
Ventral
55.0 ± 5.7 [ − 43.6]
59.7 ± 1.7 [ − 40.0]
95.9 ± 1.1 [ − 4.0]
95.0 ± 2.5 [ − 4.9]
76.4 [ − 23.1]
Table 2: OOD linear probing for context-invariance. OOD [Drop] : OOD accuracy (%) and its train–test gap (pp), averaged over 8 heads (30 samples/class). Bold: superior OOD accuracy in the dorsal streams (avg. > 80%) and the L3–L4 surge in the ventral stream of ReGraph ( D→V ).
Model
L1
L2
L3
L4
Avg
Dorsal-Only
0.0 / –
5.6 / –
4.0 / –
2.6 / –
3.0 / –
Unmodulated
0.1 / 0.1
8.6 / 6.8
4.0 / 5.9
6.3 / 5.9
4.7 / 4.7
ReGraph
0.0 / 0.1
5.1 / 6.5
8.8 / 6.0
15.3 / 7.2
7.3 / 4.9
Table 3: Ratio of grid-like bases across layers. We report the fraction of top 5% singular-value bases exceeding a shuffled null threshold (%), formatted as Dorsal / Ventral (App. J ). Bold indicates a strictly increasing trend unique to ReGraph ’s extended dorsal stream.
Figure 3: Spatial weighting maps in ReGraph . 2D autocorrelograms of pre-softmax attention bases; GS: gridness score. Dorsal (left three) : hexagonal patterns with varied orientation and spacing. Ventral (right three) : irregular or cross-like patterns lacking hexagonal structure, with near-zero GS.
Task
L1 ( Ghigh/Glow[Δ↑] )
L2 ( Ghigh/Glow[Δ↑] )
L3 ( Ghigh/Glow[Δ↑] )
L4 ( Ghigh/Glow[Δ↑] )
Bases Ablation
σ=0
82.0 / 81.7 [ −0.3 ]
82.6 / 82.9 [ +0.3 ]
70.3 / 77.3 [ +7.0 ]
89.1 / 95.1 [ +6.0 ]
σ=0.5
70.3 / 70.1 [ −0.2 ]
75.3 / 75.4 [ +0.2 ]
61.7 / 70.2 [ +8.5 ]
85.8 / 93.6 [ +7.7 ]
σ=1.0
54.3 / 54.4 [ +0.1 ]
60.9 / 60.2 [ −0.7 ]
47.1 / 55.8 [ +8.7 ]
75.4 / 86.9 [ +11.5 ]
σ=1.5
42.0 / 42.0 [ 0.0 ]
47.5 / 46.1 [ −1.4 ]
35.9 / 43.4 [ +7.5 ]
61.2 / 75.2 [ +14.0 ]
Reconstruction ( R2 )
.000 / .014 [ −.014 ]
.008 / .001 [ +.007 ]
.082 / .000 [ +.082 ]
.131 / .037 [ +.094 ]
Table 4: Functional relevance of grid-like bases. Each cell reports results obtained under Ghigh / Glow manipulations [gap ΔAcc./R2↑ ]. A larger Δ indicates greater reliance on Ghigh for both tasks (OOD Acc. under ablation; reconstruction score). Bold ( Δ>row mean ) highlights the L3-4 surge.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
GNN role
Biological analog
ψ(⋅)
Message function
Axonal/synaptic transmission gated by eji
⨁
Permutation-invariant aggregation
Dendritic integration of postsynaptic potentials
ϕ(⋅)
Update function
Non-linear somatic activation producing the next firing state
Appendix
Table 5: Correspondence between message-passing operations and cortical lateral processing.
General MP component
General form
GCN realization
Message ψ
ψ(hiℓ,hjℓ,eji)
d~id~j1Whjℓ
Aggregation ⨁
permutation-invariant op
j∈Ni∪{i}∑
Update ϕ
ϕ(hiℓ,agg)
σ(agg)
Appendix
Table 6: GCN viewed as a specific instance of the general message-passing framework in Eq. equation 4 .
Functional role
GCN
Self-Attention
Adjacency (who connects to whom)
A^ij (scalar, fixed)
dkqi⊤kj (input-dependent)
Edge-weight normalization
1/d~id~j (degree-based)
softmax over j (similarity-based)
Connectivity scope
Local: j∈Ni∪{i}
Global: every pair (i,j) , 1≤i,j≤N
Source of connectivity
Predefined graph topology
Learned via WQ,WK
Temporal behavior
Fixed across all inputs
Recomputed for every Hℓ
Transmitted signal
Whjℓ
Vj=WVhjℓ
Appendix
Table 7: Component-level equivalence between GCN and self-attention. Self-attention preserves the graph-operator structure of GCN but lifts the adjacency from a static, sparse object to a dynamic, dense one.
Stream
Cue
Model
Extraction Block
Output Dim.
Post-processing
Dorsal
Depth
Depth Anything V2
refinenet2
64×16×16
Flatten to ND=196 , project to d=512 , fuse via cross-attention →ED
Flow
GMFlow
Conv1 (upsampler)
256×28×28
Saliency
UniSal
upsample_2
128×28×28
Ventral
Object
ViT-S/8 (DINO)
Block 9
384×28×28
Discard [CLS] , project to d=512→EV ( NV=784 )
Appendix
Table 8: Pretrained network specifications for the primary visual stage. “Extraction Block” denotes the specific layer from which representations are drawn. “Output Dim.” specifies the raw channel and spatial shape ( C′×H′×W′ ) per frame, prior to spatial resampling and linear projection that produce the tokenized form RNS×d .
ID
Action Class
6
Covering something with something
12
Dropping something onto something
15
Hitting something with something
16
Holding something
19
Holding something next to something
36
Moving something and something away from each other
Appendix
Table 9: The 33 action classes selected from SSV2, sorted by class ID. IDs correspond to the original 174-class indexing of the SSV2 dataset.
Figure 4: Example videos and corresponding templated descriptions ( Goyal et al., 2017 ) .
Hyperparameter
Single-Stream
Dual-Stream
Learning rate (ReGraph layers)
5e-4
3e-4
Learning rate (pretrained networks)
5e-5
3e-5
Weight decay
0.03
0.05
Batch size
5
5
Gradient accumulation
4
6
Effective batch size
20
30
Appendix
Table 10: Training hyperparameters. The ”Single-Stream” column applies to both single-stream references (Dorsal-Only and Ventral-Only), and the ”Dual-Stream” column applies to all dual-stream variants.
Layer
H1
H2
H3
H4
H5
H6
H7
H8
Mean ( ± std)
V1
0.292
0.283
0.249
0.294
0.326
0.262
0.253
0.269
0.279 ( ± 0.024)
V2
0.277
0.287
0.313
0.298
0.277
0.283
0.290
0.283
0.289 ( ± 0.011)
V3
0.444
0.479
0.478
0.491
0.481
0.444
0.482
0.472
0.471 ( ± 0.017)
Appendix
Table 11: Learned D→V gate values σ(γ) per head at each ventral layer of ReGraph . Vk ( k∈{1,2,3} ) denotes the k -th ventral layer; V4 is omitted as it receives no D→V modulation. Bold highlights the sharp opening at V3 across all heads.
Model
Stream
L1
L2
L3
L4
Avg
Dorsal-Only
Dorsal
0 (0.0%)
105 (5.6%)
75 (4.0%)
48 (2.6%)
228 (3.0%)
Unmodulated
Dorsal
1 (0.1%)
161 (8.6%)
74 (4.0%)
118 (6.3%)
354 (4.7%)
Ventral
9 (0.1%)
553 (6.8%)
479 (5.9%)
478 (5.9%)
1519 (4.7%)
ReGraph
Dorsal
0 (0.0%)
95 (5.1%)
164 (8.8%)
286 (15.3%)
545 (7.3%)
Ventral
9 (0.1%)
525 (6.5%)
483 (6.0%)
588 (7.2%)
1605 (4.9%)
Appendix
Table 12: Counts and ratios of grid-like spatial bases (Top-5% retention rate). Each cell shows ngrid ( Ratio% ), aggregated across classes and heads. Bold highlights ReGraph’s dorsal stream, exhibiting monotonic emergence of grid-like bases.
Model
Stream
L1
L2
L3
L4
Avg
Dorsal-Only
Dorsal
.244 / –
.240 / .272
.244 / .324
.242 / .341
.243 / .312
Unmodulated
Dorsal
.242 / .269
.240 / .409
.242 / .390
.241 / .382
.241 / .363
Ventral
.216 / .274
.216 / .340
.198 / .301
.186 / .294
.204 / .302
ReGraph
Dorsal
.243 / –
.242 / .366
.245 / .390
.244 / .351
.244 / .369
Ventral
.216 / .276
.216 / .340
.193 / .313
.178 / .283
.201 / .303
Appendix
Table 13: Mean significance thresholds and gridness scores (Top-5% retention rate). Each cell shows θˉnull/gˉsig ; “–” indicates no grid-like bases were identified. Bold highlights ReGraph’s dorsal stream.
Layer
3%
5%
7%
L1
0.0
0.0
0.2
L2
1.4
5.1
5.6
L3
12.1
8.8
8.3
L4
20.2
15.3
12.1
Appendix
Table 14: Ratio of grid-like bases (%) in the ReGraph dorsal stream across retention rates. The monotonic increase across layers is consistently reproduced at all rates.
Figure 5: RPA pipeline. A. a) SSV2 attention matrices ( M ) decomposed via HOSVD to isolate spatial bases ( U2,U3 ). b-c) Top 3, 5, 7% of bases (ranked by singular value) reshaped into 2D spatial weighting maps. d) Gridness scores computed per map. B. a) Logistic-regression classifier trained on spatiotemporally pooled representations. b) Evaluation on the test set (unseen objects) assesses context-invariance; smaller accuracy drop indicates higher invariance.
Model
Initialization
Fine-tuning
Top-1
VideoMAE
K400 (self-supervised)
SSV2-33 subset
80.4
ReGraph (ours)
depth/flow/saliency/object
74.6
SlowFast
K400 (supervised)
66.1
TimeSformer
ImageNet-21K → K400 (supervised)
62.3
Appendix
Table 15: Action-recognition comparison on the 33-class SSV2 subset. Top-1 accuracy (%) under the shared 6-view evaluation protocol.
Model
L1
L2
L3
L4
ReGraph — dorsal
0.0
5.1
8.8
15.3
ReGraph — ventral
0.1
6.5
6.0
7.2
VideoMAE
2.4
1.6
1.0
0.9
TimeSformer
1.1
0.9
1.0
1.4
SlowFast — fast
1.6
1.6
0.0
0.0
SlowFast — slow
1.6
0.0
0.0
0.0
Appendix
Table 16: Ratio of grid-like bases (%) in baseline models.
Figure 6: Protocols for testing the functional relevance of grid-like bases. A. Bases-ablation Protocol. a) Top 5% spatial bases sorted by gridness into Ghigh and Glow . b) Attention matrices ablated by projecting out Ghigh or Glow subspaces. c) OOD accuracy (Acc.) evaluated under Gaussian noise perturbations. d) OOD Acc. gap, ΔAcc.=Acc.(Glow)−Acc.(Ghigh) , quantifies reliance on Ghigh . B. Connectivity Reconstruction. a) Disjoint top- k ( Ghigh ) and bottom- k ( Glow ) subsets selected per mode. b) Unseen attention maps reconstructed via bilateral projection. c) Layer-wise attention maps and R2 scores evaluated for each group. d) Reconstruction gap, ΔR2=Rhigh2−Rlow2 , assesses preservation of unseen connectivity.
Layer
Number of Ablated Bases ( k )
k=1 : Ghigh / Glow[ΔAcc.↑]
k=5 : Ghigh / Glow[ΔAcc.↑]
k=9 : Ghigh / Glow[ΔAcc.↑]
L1
82.0 / 81.9 [ − 0.1]
81.9 / 81.6 [ − 0.3]
82.0 / 81.7 [ − 0.3]
L2
81.3 / 81.5 [ + 0.2]
80.7 / 81.9 [ + 1.1]
82.6 / 82.9 [ + 0.3]
L3
76.0 / 78.3 [ + 2.4]
69.7 / 77.5 [ + 7.8]
70.3 / 77.3 [ + 7.0]
L4
97.2 / 97.4 [ + 0.2]
93.6 / 96.6 [ + 3.0]
89.1 / 95.1 [ + 6.0]
Appendix
Table 17: OOD performance to the number of ablated bases ( k ). Cells report OOD accuracy under Ghigh / Glow ablation [gap ΔAcc.↑ ]. A larger positive ΔAcc. indicates greater reliance on Ghigh . Bold highlights the prominent reliance surge at higher layers ( L3 – L4 ), which is most pronounced at k=9 .
Pool
k
L1
L2
L3
L4
30%
5
.005 / .000 [ + 0.005]
.001 / .003 [ − .002]
.055 / .000 [ + .055 ]
.105 / .022 [ + .084 ]
30%
9
.025 / .000 [ + 0.025]
.003 / .066 [ − .062]
.120 / .001 [ + .118 ]
.153 / .052 [ + .101 ]
40%
5
.004 / .000 [ + 0.004]
.001 / .000 [ + 0.001]
.048 / .000 [ + .048 ]
.084 / .016 [ + .068 ]
40%
9
.014 / .000 [ + 0.014]
.001 / .008 [ − .007]
.082 / .000 [ + 0.082]
.131 / .037 [ + 0.094]
50%
5
.004 / .000 [ + 0.004]
.000 / .000 [ + 0.000]
.044 / .000 [ + .044 ]
.060 / .009 [ + .051 ]
50%
9
.012 / .000 [ + 0.012]
.001 / .000 [ + 0.001]
.058 / .000 [ + .058 ]
.123 / .022 [ + .101 ]
Appendix
Table 18: Reconstruction R2 across pool sizes and subset sizes. Each cell reports reconstruction scores under Ghigh / Glow [gap ΔR2=RGhigh2−RGlow2↑ ], averaged over 13×8 (class, head) pairs ( k=1 omitted as R2<0.01 ). The bottom row reports the layer-wise mean of the gaps ( ΔR2 ). Bold highlights the prominent surge at L3 and peak at L4 on average across configurations.
Variant
Top-1 (%)
Epochs
h/epoch
Wall-clock
GPU-hours
Single-stream references
Dorsal-Only
71.69
24
1.7
∼ 43 h
∼ 86
Ventral-Only
60.38
15
0.2
∼ 4 h
∼ 8
Dorsal–Ventral biological trait asymmetry
(a) Trait-Symmetric
69.78
22
0.8
∼ 19 h
∼ 38
(b) Dorsal-trait Only
72.42
29
1.3
∼ 38 h
∼ 76
Appendix
Table 19: Compute resources and wall-clock training time for each reported ablation. All models trained on 2 × NVIDIA RTX 5090 GPUs (32 GB each), with maximum 100 epochs and early stopping at patience = 8. Wall-clock includes both training and the final 6-view ensemble evaluation; GPU-hours = wall-clock × 2.
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.
Hardik Rajpal, Dan Goodman
1I-X Centre for AI in Science, Imperial College London, W12 0BZ, UK · Department of Electrical and Electronic Engineering, Imperial College London, SW7 2AZ, UK
The spatial and functional organization of the primate visual cortex is a fundamental problem in neuroscience. While recent computational frameworks like the Topographic Deep Artificial Neural Network (TDANN) have successfully modeled spatial organization in the ventral stream, the computational origins of the dorsal stream's distinct topographies, such as direction-selective maps in the middle temporal (MT) area, remain largely unresolved. In this work, we present a spatiotemporal TDANN to investigate whether MT topography is governed by the same universal principles. By training a 3D ResNet on naturalistic videos via a Momentum Contrast (MoCo) self-supervised paradigm alongside a biologically inspired spatial loss, we demonstrate the spontaneous emergence of brain-like direction maps and topological pinwheel structures. Crucially, we reveal that MT tuning properties, characterized by strong direction selectivity paired with a residual axial component, arise from a strict optimization trade-off between task-driven discriminative pressure and spatial regularization. The model's representations quantitatively match in vivo macaque MT physiological baselines, including direction selectivity index, circular variance, and pinwheel density. These findings unify the computational origins of the ventral and dorsal streams, establishing a general mechanism for cortical self-organization.
Zhaotian Gu, Molan Li, Jie Su +3
School of System Science, Beijing Normal University, Beijing 100875, China · Qiyuan Laboratory, Beijing 100095, China · State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal2026 University, Beijing 100875, China
Humans abstract experiences into structured representations to facilitate pattern inference and knowledge transfer. While the hippocampal-entorhinal (HPC-MEC) circuit is known to represent both spatial and conceptual spaces, the mechanisms for concurrently extracting abstract structures from continuous, high-dimensional dynamics remain poorly understood. We propose a brain-inspired hierarchical model that simultaneously infers latent transitions and constructs a predictive visual world model. Our architecture employs an inverse model for structural extraction alongside an HPC-MEC coupling model that dissociates relational structures (MEC) from integrated episodic scenes (HPC). Using primitive transformation dynamics as a benchmark, we demonstrate the model's capacity for structural abstraction. By leveraging velocity-driven path integration, the framework enables robust prediction and structural reuse across diverse contexts, thereby achieving structural generalization. This work provides a novel computational framework for understanding how brain-inspired, self-supervised learning of world models facilitates the acquisition of reusable abstract knowledge.
Tianqiu Zhang, Muyang Lyu, Xiao Liu +1
Peking-Tsinghua Center for Life Sciences, Academy for Advanced Interdisciplinary Studies, IDG/McGovern Institute for Brain Research, Center of Quantitative Biology, School of Psychological and Cognitive Sciences, Key Laboratory of Machine Perception (Ministry of Education), Peking University. · HHMI Janelia.