Procedural Core: A Compact Recurrent Initialization for Vision Transformers
Authors: Zachary Shinnick, Christian Internò, Hemanth Saratchandran, Anton van den Hengel, Damien Teney
Organizations: Australian Institute for Machine Learning (AIML), Adelaide University · Bielefeld University · Metacognition AI · Idiap Research Institute
Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.
Figures & tables
Figure 1: We propose a generic initialization for transformers as an alternative to random weights. (Left) We first train a small auxiliary model to capture generic computations. We generate abstract training data with simple algorithms such as formal languages with nested structures (pictured as balanced parentheses). The model is recurrent in depth such that it captures regularities of the data in a minimal set of weights: the Procedural Core . (Middle) We unroll and expand the Procedural Core to initialize a target model of arbitrary width and depth. (Right) This bootstraps the subsequent standard training, improving convergence and generalization compared to a random initialization.
Figure 2: Procedural Core improves ViT-Base training on ImageNet-1k . Left: Final top-1 accuracy (mean ± s.d. over three seeds). Right: Mean training trajectories over the same seeds, showing a persistent advantage over default initialization.
Figure 3: Procedural Core improves representation learning with DINO on ImageNet-1k . Left: k -NN accuracy at selected epochs. Right: k -NN accuracy throughout training.
Figure 4: Procedural Core improves language modeling compared to a default random initialization (lower is better; mean ± std. dev. over 3 seeds).
Figure 5: Architectural choices and ablations of our method. (a) Transfer peaks with three unique auxiliary layers. (b) Removing recurrence substantially reduces transfer despite matching the parameter count. (c) The expansion in depth is effective with various number of layers in the target model. (d) The expansion in width is slightly more effective than a baseline with zero-padding.
Figure 6: Cumulative energy of singular values of different weight matrices in models trained with the Procedural Core recurrent architecture ( ) vs. a direct warm-up ( ) as in Shinnick et al. (2026) .
Figure 7: Rank truncation hinders transfer more sharply with Proc. Core weights.
Figure 8: Token dynamics across depth. Per-token hidden-state norms for default and Procedural Core initializations. High-norm outliers are suppressed from the middle blocks onwards with Procedural Core . Heatmaps are independently normalized; extended examples are in Appendix B.6 .
Figure 9: Procedural Core suppresses high-norm tokens and reduces attention to them. Statistics on 200 ImageNet-1k images. (a) Bimodality across depth. (b-c) Token-norm distributions at peak-bimodality blocks. (d) Attention mass received by high-norm tokens.
Figure 10: Transfer is concentrated in the value/output pathway. (a,b) Component-shuffling accuracy. (c,d) Shuffling V/proj. reverts to the default behavior with high-norm tokens, while shuffling Q/K preserves the Procedural Core effect.
Figure 11: Procedural Core improves representations, particularly for dense prediction tasks (segmentation, localization, and depth estimation; frozen-backbone performance). This is consistent with our finding of preserving local visual information by suppressing high-norm tokens.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Batch size
256
Training steps
15,000
Mask ratio
0.5 (close-only)
Optimizer
AdamW
Learning rate
2×10−3
Weight decay
0.05
Appendix
Table 1: Optimization settings used to train the auxiliary model for learning the Procedural Core.
Setting
Value
Architecture
GPT-2 Small (12L, 768D, 12H)
Context length
1024
Training budget
2B tokens
Effective batch size
32 sequences (32,768 tokens)
Optimizer
AdamW
Learning rate
6×10−4 (cosine to 6×10−5 )
Appendix
Table 2: GPT-2 Small training configuration for FineWeb-Edu and CodeParrot experiments.
Top-1 Acc. (%)
Procedural warm-up
73.9±0.1
Procedural Core (ours)
74.6±0.5
Appendix
Table 3: Comparison with procedural warm-up on CIFAR-100 . ViT-Tiny top-1 accuracy (mean ± std).
Initialization
Top-1 (%)
Default init.
88.4±0.8
Mimetic init.
89.1±0.5
Procedural warm-up
89.4±0.3
Procedural Core
89.7±0.1
Appendix
Table 4: Downstream transfer on CIFAR-100 . Fine-tuning accuracy after ImageNet-1K pretraining (mean ± std over three seeds).
Initialization
Top-1 (%)
Default initialization
75.36
Procedural Core (ours)
75.42
Appendix
Table 5: Linear evaluation of DINO representations on ImageNet-1K . Frozen ViT-S features evaluated with a linear classifier.
Figure 12: Layer-wise singular value spectra under recurrent vs. independent parameterization. Singular value decay for attention and MLP weight matrices across layers 1–10. In the recurrent model, the middle layers share parameters, so the same spectrum appears across layers. Each subplot contrasts this with independently parameterized layers. Recurrent layers exhibit slower spectral decay, indicating a more distributed use of singular directions.
Figure 13: Additional token-norm dynamics across depth. Per-token hidden-state norms across transformer blocks for default and Procedural Core initializations. The models begin with similar token-norm structure but increasingly diverge across depth, with the Procedural Core suppressing high-norm token outliers in later blocks.
Figure 14: Full downstream evaluation across dense and localization tasks. Frozen ViT-B backbones evaluated on segmentation ( ADE20K , ImageNet-S ), object localization ( VOC07 ), and depth estimation ( NYUv2 ). The Procedural Core improves over default initialization across the primary metrics; shuffling V/proj largely removes these gains, while transplanting V/proj+LN recovers much of them. Error bars denote 95% bootstrap intervals ( NYUv2 : s.d. over three readout seeds).
Figure 15: Under the stronger DeiT-III recipe, gains concentrate on attention-based tasks. Full downstream evaluation under DeiT-III ( Touvron et al., 2022 ) . Procedural Core improves zero-shot segmentation on ImageNet-S and unsupervised object localization on VOC07 , while performance is broadly unchanged on the remaining tasks.
Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temporal modeling, and VLM integration often call for customized SOTA-level ViTs. Training such encoders remains beyond the reach of much of the community, as it requires massive image-text data, while standard softmax attention makes high-resolution or dynamic-resolution pretraining prohibitively costly and often forces low-resolution pretraining followed by post-hoc adaptation. TuringViT addresses these challenges with three key designs: Turing Linear Attention (TLA) for efficient sequence modeling, VISTA-Curation to construct supervision-rich image-video training data, and native dynamic-resolution pretraining that supports flexible inputs from the start and transfers seamlessly to downstream VLMs. As a result, TuringViT outperforms leading open-source ViT baselines with only 10% of the data, achieves stronger downstream VLM performance, and delivers substantially better latency scaling on high-resolution inputs. Our scaling-law analysis further shows that TuringViT continues to improve predictably with curated data scale, far from saturation. Its fast adaptation, hardware-friendly design, and efficient deployment have made it a unified visual foundation across XPeng's AI systems. More broadly, TuringViT provides a reproducible pipeline that dramatically lowers the cost for the community to train, customize, and deploy SOTA-level ViTs, moving toward making such Vision Transformers accessible to all.
Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.
Dhananjay Tomar, Marius Aasan, Andreas Kleppe +1
Institute for Cancer Genetics and Informatics Oslo University Hospital Oslo, Norway · Department of Informatics University of Oslo Oslo, Norway · SFI Visual Intelligence UiT The Arctic University of Norway Tromsø, Norway
Vision Transformers (ViTs) implement depth by stacking independently parameterized blocks, but it remains unclear how much of this parameterization is necessary and how much can be replaced by recurrent reuse. We study this question with bViT, a single-block recurrent ViT that repeatedly applies the same transformer block while preserving the iterative computation of a deep model. On ImageNet-1K, bViT-B reaches 0.779 validation accuracy compared with 0.789 for ViT-B under the same training recipe and computational budget, while using 8.6M rather than 86.6M parameters. This correspondence becomes stronger with model width, while narrow recurrent models exhibit a substantial performance gap. Beyond classification, the single-block formulation provides a controlled testbed for studying how transformer computation evolves with depth, since the same heads, neurons, and weight matrices can be tracked across recurrent steps. Analyses of attention, activation patterns, and step-conditioned spectral pruning reveal temporally organized behavior and step-dependent utilization of the shared parameters. bViT also transfers competitively to downstream tasks while enabling highly parameter-efficient adaptation. Our work shows that much of the performance associated with independently parameterized ViT depth can be recovered through recurrent reuse of a single sufficiently wide transformer block.
Michal Byra, Pawel Olszowiec, Grzegorz Stefanski +2
Samsung AI Center, Warsaw, Poland · Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland