VANDAM: Viewing a nucleotide sequence with DNA molecular priors
Authors: Jeremy Levy, Ariel Larey, Yury Nahshan, Raizy Kellerman, Elay Dahan, Amit Bleiweiss, Guy Leib, Omri Nayshool, +13 more
Organizations: Applied AI Architecture, NVIDIA, Israel. · Worldwide Field Ops, NVIDIA, Israel. · Developer Programs, NVIDIA, Israel. · Cancer Research Center and Wohl Institute of Translational Medicine, Sheba Medical Center, Tel Hashomer, Israel. · Dina Recanati School of Medicine Reichman University, Israel. · Windreich Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, USA.
Contemporary Genomic Foundation Models (GFMs) rely on a DNA-as-a-string paradigm that employs masked token prediction objectives for pretraining. However, this abstraction does not explicitly model the biochemical, structural, and physical properties essential to biological function. Many molecular properties can be estimated from sequence using established biophysical models, so their utility lies not in providing an independent modality, but in introducing priors that training objectives can explicitly exploit. We introduce VANDAM, a framework that extends the training of GFMs with DNA molecular priors. In self-supervised training, VANDAM predicts regional molecular properties from pooled representations. When functional labels are available and can reward retaining molecular priors, local features are additionally injected at the input. VANDAM consistently improves downstream performance across four architecture families and nine held-out genomic tasks by complementing token-based objectives. Probing experiments further demonstrate that the use of molecular priors generalizes to other unseen molecular properties.
Figures & tables
Figure 1: VANDAM architecture and molecular feature hierarchy. a , High-level training flow. Local Feature Injection (LFI; dashed) augments input token representations with local sequence-derived descriptors and is used only in VANDAM-MT, where functional-task supervision Ltask can reward their use. Regional Property Prediction (RPP; blue) pools representations for genomic regions and predicts tile-level properties in both VANDAM-PT and VANDAM-MT, alongside the native token objective Ltoken . b , RPP mechanism. Backbone representations are partitioned into tiles, attention-pooled within each tile, passed through a shared prediction head, and compared with transformed and calibrated sequence-derived regional targets using LRPP . c , Local features comprise per-nucleotide geometric and electrostatic descriptors; regional targets comprise tile-level measures of curvature, composition, duplex stability, and non-B-DNA structural propensity.
Model
Params
Architecture
Tokenizer
Pre-training objective
NTv3-8M
8M
U-Net conv-transformer
Single-nucleotide
MLM
Caduceus-PS
7.73M
Bi-directional Mamba
Single-nucleotide
MLM
DNABERT-2
117M
Flat transformer
BPE
MLM
HyenaDNA
3.28M
Hyena operator
Single-nucleotide
NTP
Table 1: Backbone models used in all experiments. Parameter counts are trainable backbone parameters in our implementation, excluding VANDAM and downstream heads.
NTv3-8M
Caduceus
DNABERT-2
HyenaDNA
Task
Baseline- PT
VANDAM- PT
Baseline- PT
VANDAM- PT
Baseline- PT
VANDAM- PT
Baseline- PT
VANDAM- PT
Clinical pathogenicity (ClinVar)
58.0
61.7 +
60.5
61.9 +
59.3
65.9 +
66.9
75.3 +
Disease variants (OMIM)
49.8
59.0 +
48.6
55.0 +
55.7 –
51.2
52.2
52.1
Mendelian variants (TraitGym)
49.7
55.8 +
47.2
55.9 +
54.2
53.7
53.9
54.2
BRCA1 variant effect
50.9
51.6
49.6
54.0 +
56.9
57.6
51.9
55.3 +
Transcription-factor binding
79.7
80.8 +
68.9
75.2 +
79.0
80.3 +
73.7
80.0 +
Table 2: Pre-training results per task. Bold denotes the higher central score; superscript + / − denotes a BH-adjusted paired difference favoring VANDAM/Baseline.
NTv3-8M
Caduceus
DNABERT-2
HyenaDNA
Task
Baseline- MT
VANDAM- MT
Baseline- MT
VANDAM- MT
Baseline- MT
VANDAM- MT
Baseline- MT
VANDAM- MT
Clinical pathogenicity (ClinVar)
68.1
70.2 +
54.4 –
51.4
59.5
69.7 +
55.7
71.2 +
Disease variants (OMIM)
54.9
57.4
47.4
53.1 +
52.0
66.4 +
47.9
51.0
Mendelian variants (TraitGym)
51.5
56.8 +
47.8
55.6 +
55.0
57.5
51.4
54.8
BRCA1 variant effect
51.9
54.6
51.7
52.5
52.5
56.2 +
51.6
53.4
Transcription-factor binding
78.7
79.8 +
53.7
66.4 +
78.1
82.6 +
75.2
81.0 +
Table 3: Multi-task training results per task. Bold denotes the higher central score; superscript + / − denotes a BH-adjusted paired difference favoring VANDAM/Baseline.
Configuration
Pre-training
Multi-task
Original checkpoint
59.5
59.5
Baseline-PT / Baseline-MT
60.5
63.0
LFI only
60.5
64.2
RPP only
63.8
63.2
LFI + RPP
63.4
65.1
RPP without MLM loss, masking retained
60.2
60.2
Table 4: Component ablation across pre-training and multi-task training on NTv3-8M. Values are mean AUROC over the eight AUROC tasks.
Figure 2: Biological-content controls on NTv3-8M (mean AUROC over eight tasks). a , Pre-training RPP-target controls. b , Multi-task LFI controls with biological RPP. c , Multi-task RPP-target controls with biological LFI.
Descriptor
Original
Residual
Descriptor
Original
Residual
Hole central population
0.060
0.054
Electron central population
0.083
0.057
Hole maximum transfer
−0.041
−0.032
Electron maximum transfer
−0.055
−0.046
Hole delocalization (IPR)
0.071
0.091
Electron delocalization (IPR)
0.067
0.051
Exciton lifetime
0.073
0.084
Electron–hole separation
0.069
0.076
Table 5: qDNA transfer gain ( ΔR2 , VANDAM-PT minus Baseline-PT). Bold denotes statistical significant improvement.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Feature
Full name
Units
Structural category
Biophysical interpretation
MGW
Minor Groove Width
Å
Groove geometry
Width of the minor groove at the center base; narrow grooves deepen the local negative electrostatic potential.
HelT
Helical Twist
degrees
Inter-bp rotation
Rotation about the helix axis between successive base-pair steps.
ProT
Propeller Twist
degrees
Intra-bp rotation
Out-of-plane twist of the two bases within a pair.
Roll
Roll
degrees
Inter-bp rotation
Rotation about the long axis of a base-pair step; a principal component of intrinsic bending.
EP
Electrostatic Potential
kT/e
Groove/electrostatic
Electrostatic potential in the minor groove at the center base.
Stretch
Stretch
Å
Intra-bp translation
Separation of the two bases along the base-pair long axis.
Appendix
Table 6: The 14 DNAshapeR channels used by VANDAM. Rotational quantities are measured in degrees and translational quantities in Å.
Table 11: The eight-task multi-task training pool. The training tasks are task-disjoint from the evaluation suite; record-level filtering further removes training examples that overlap an evaluation example under the policy below.
Training task
Visible
Excluded
Retained
GUE promoter
47,356
40,183 (84.85%)
7,173
GUE splice site
36,496
18,006 (49.34%)
18,490
LRB enhancer
99,995
28,987 (28.99%)
71,008
Common-vs-rare
193,908
50,620 (26.11%)
143,288
meQTL
73,601
25,108 (34.11%)
48,493
sQTL
1,000,741
385,758 (38.55%)
614,983
Appendix
Table 12: Record-level filtering of the eight-task multi-task training pool. Counts are examples before filtering; the retained pool was used for all reported multi-task results.
Task
Access
Prediction
Metric
Test n
Len (bp)
BRCA1 variant effect
zero-shot
cosine dissimilarity
AUROC
3,893
8,192
Clinical pathogenicity (ClinVar)
zero-shot
cosine dissimilarity
AUROC
258,944
8,192
Disease variants (OMIM)
zero-shot
cosine dissimilarity
AUROC
200,406
8,192
Mendelian variants (TraitGym)
zero-shot
cosine dissimilarity
AUROC
3,380
8,192
Transcription-factor binding
linear probe
probability
AUROC
5,000
100
Regulatory variants (eQTL)
linear probe
probability
AUROC
8,862
8,192
Appendix
Table 13: The nine-task evaluation suite. Len denotes the task-level sequence window. NTv3, Caduceus, and HyenaDNA use the listed context; DNABERT-2 uses 2,500-bp crops.
Backbone
Task
Δ [95% CI]
q
Supported
NTv3-8M
ClinVar
3.6 [ 3.3,4.0 ]
<0.001
VANDAM
NTv3-8M
OMIM
9.3 [ 5.4,13.1 ]
<0.001
VANDAM
NTv3-8M
TraitGym
6.1 [ 1.4,10.7 ]
0.008
VANDAM
NTv3-8M
BRCA1
0.6 [ −2.1,3.3 ]
0.425
–
NTv3-8M
TF binding
1.1 [ 0.4,1.7 ]
0.001
VANDAM
NTv3-8M
eQTL
1.9 [ 1.4,2.3 ]
<0.001
VANDAM
Appendix
Table 14: Paired inference for pre-training comparisons. Δ is VANDAM minus Baseline in points; brackets give the paired 95% confidence interval. The q column reports the BH-adjusted one-sided q -value in the direction of the observed Δ . “Supported” denotes the direction meeting the paired-interval and BH-adjusted q<0.05 criteria.
Backbone
Task
Δ [95% CI]
q
Supported
NTv3-8M
ClinVar
2.1 [ 1.9,2.4 ]
<0.001
VANDAM
NTv3-8M
OMIM
2.5 [ −1.0,6.0 ]
0.102
–
NTv3-8M
TraitGym
5.3 [ 0.9,9.8 ]
0.015
VANDAM
NTv3-8M
BRCA1
2.7 [ −0.2,5.7 ]
0.051
–
NTv3-8M
TF binding
1.1 [ 0.2,2.0 ]
0.010
VANDAM
NTv3-8M
eQTL
1.2 [ 0.7,1.6 ]
<0.001
VANDAM
Appendix
Table 15: Paired inference for multi-task comparisons. Δ is VANDAM minus Baseline in points; brackets give the paired 95% confidence interval. The q column reports the BH-adjusted one-sided q -value in the direction of the observed Δ . “Supported” denotes the direction meeting the paired-interval and BH-adjusted q<0.05 criteria.
Task
MFM (bio)
MFM (rand)
RPP (alg)
RPP (shuf)
RPP (bio)
Clinical pathogenicity (ClinVar)
67.2
60.7
54.0
53.8
61.7
Disease variants (OMIM)
52.6
53.5
51.9
47.0
59.0
Mendelian variants (TraitGym)
49.0
48.0
48.9
47.2
55.8
BRCA1 variant effect
51.5
49.8
50.2
48.9
51.6
Transcription-factor binding
81.2
81.1
78.9
78.7
80.8
Regulatory variants (eQTL)
69.1
68.4
69.0
67.2
70.2
Appendix
Table 16: Local feature reconstruction versus regional property prediction on NTv3-8M. MFM (bio) uses the composite biological feature provider and MFM (rand) a random 5-mer provider. RPP (alg) predicts generic regional sequence statistics, whereas RPP (shuf) predicts properties recomputed from a shuffled pentamer shape table. Values are AUROC except for CAGE expression, which uses macro Pearson r ; mean AUROC averages the eight AUROC tasks. Bold marks the best value in each row.
Figure 3: Learned utilization of the DNAshape local-feature-injection pathway. A: Distribution across valid positions of the centered DNAshape-branch contribution relative to the centered sequence-embedding contribution after the learned mixing projection. B: Mixed-embedding change induced by a one-standard-deviation perturbation of each feature, computed from the composed feature-encoder and mixer map. These analyses describe learned activation and parameter scales, not causal downstream-task importance.
Task
Full RPP
− folding
− origin
− curvature
Transcription-factor binding
80.91
80.44
80.50
80.72
Regulatory variants (eQTL)
70.20
68.54
68.56
68.72
Gene-expression variants
63.44
63.77
63.09
63.33
BRCA1 variant effect
51.55
52.49
53.02
48.87
Disease variants (OMIM)
59.00
49.72
53.28
49.51
Mendelian variants (TraitGym)
55.78
56.83
56.67
52.16
Appendix
Table 17: NTv3-8M VANDAM-PT ablation of RPP targets. “ − folding” removes G-quadruplex, cruciform, and Z-DNA propensities; “ − origin” removes AT content, melting propensity, and intrinsic curvature; “ − curvature” removes intrinsic curvature. For supervised tasks, all columns use the same fixed probe seed, so differences isolate the effect of removing each RPP target. Zero-shot scores are seed-independent.
Backbone
Objective
Loss ↓
Accuracy ↑
NTv3-8M
MLM
1.098→1.110
50.6%→50.2%
Caduceus
MLM
1.095→1.109
51.1%→50.2%
DNABERT-2
MLM (BPE)
5.060→5.058
16.4%→16.4%
HyenaDNA
Causal NTP
1.265→1.250
40.2%→41.3%
Appendix
Table 18: Intrinsic test metrics after continual pre-training. Each cell gives Baseline-PT → VANDAM-PT.
Backbone
Objective
Loss ↓
Accuracy ↑
NTv3-8M
MLM
1.136→1.118
48.6%→49.7%
Caduceus
MLM
1.207→1.195
44.0%→44.8%
DNABERT-2
MLM (BPE)
5.007→5.673
16.8%→14.3%
HyenaDNA
Causal NTP
1.267→1.252
40.1%→41.1%
Appendix
Table 19: Intrinsic test metrics after multi-task training. Each cell gives Baseline-MT → VANDAM-MT.
Descriptor
Baseline R2
Baseline-PT R2
VANDAM-PT R2
Original ΔR2 [95% CI]
Residual ΔR2 [95% CI]
Hole central population
0.717
0.709
0.769
0.060 [ 0.023,0.104 ]
0.054 [ 0.027,0.083 ]
Hole maximum transfer
0.300
0.310
0.268
−0.041 [ −0.113,0.006 ]
−0.032 [ −0.078,0.007 ]
Hole delocalization (IPR)
0.766
0.742
0.812
0.071 [ 0.040,0.106 ]
0.091 [ 0.063,0.122 ]
Electron central population
0.671
0.663
0.746
0.083 [ 0.050,0.122 ]
0.057 [ 0.027,0.089 ]
Electron maximum transfer
0.343
0.353
0.298
−0.055 [ −0.186,0.025 ]
−0.046 [ −0.194,0.041 ]
Electron delocalization (IPR)
0.772
0.771
0.838
0.067 [ 0.047,0.090 ]
0.051 [ 0.021,0.082 ]
Appendix
Table 20: Frozen linear probing of qDNA descriptors at 300 bp on the independent confirmation split. Checkpoint R2 values use the original targets. Both ΔR2 columns compare VANDAM-PT with Baseline-PT; intervals are paired source-sequence-bootstrap 95% confidence intervals.
Handling
ClinVar
OMIM
TraitGym
BRCA1
TF
eQTL
Gene expression
Histone
CAGE
Zeroing (VANDAM)
70.2
57.4
56.8
54.6
79.8
69.3
63.1
69.3
35.9
Uniform
68.6
58.1
56.1
59.4
78.9
69.8
70.7
71.0
33.3
Conditional
70.8
55.9
54.3
61.2
78.6
70.1
70.0
71.4
33.5
Appendix
Table 21: Sensitivity of NTv3-8M VANDAM-MT to masked-base handling. Scores are AUROC except CAGE expression, which uses macro Pearson r .
Apr 17, 2026·Zhijiang Tang, Jiaxin Qi, Yan Cui +3Pretraining
Computer Network Information Center, Chinese Academy of Sciences, Beijing, China · Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Zhejiang, China
College of Computer Science, Inner Mongolia University, Hohhot, Inner Mongolia, 010021, China · National and Local Joint Engineering Research Center of Intelligent Information Processing Technology for Mongolian, Hohhot, Inner Mongolia, 010021, China · Inner Mongolia Key Laboratory of Multilingual Artificial Intelligence Technology, Hohhot, Inner Mongolia, 010021, China +1