Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
Figures & tables
Figure 1: Folding performance and transfer to general reasoning. (a) Full-target folding scores on all 334 FoldBench proteins ( Xu et al., 2025 ) , for Fold2Reason and for general-purpose language models prompted to emit C α coordinates directly. (b) Gains of Fold2Reason over its base model on the ten General-10 benchmarks, in percentage points, grouped into the four reasoning categories used throughout the paper.
Figure 2: The Fold2Reason pipeline from protein structures to general reasoning transfer. (a) Known protein structures generate deterministic FoldingCorpus labels and continuous geometry targets; ground-truth coordinates are used only for supervision. (b) The adapted model transfers to general reasoning benchmarks without protein inputs at inference. (c) A frozen language-model backbone with trainable LoRA and a shared workspace feeds FoldingCorpus and geometry readouts during post-training.
Category
Benchmark
Base
Post-training
Hidden Geom.
Format Copy
Fixed Shuffle
Fold2Reason
Spatial reasoning
FTB-Core
36.18
37.25 (+1.07)
36.04 (-0.14)
35.41 (-0.77)
40.26 (+4.08)
SpatialViz
26.61
23.16 (-3.45)
28.56 (+1.95)
25.76 (-0.85)
32.71 (+6.10)
VSI Bench
58.53
59.22 (+0.69)
56.32 (-2.21)
58.42 (-0.11)
60.16 (+1.63)
Graph reasoning
GraphQA Easy
65.14
66.90 (+1.76)
63.95 (-1.19)
64.91 (-0.23)
69.70 (+4.56)
GraphQA Hard
32.65
34.83 (+2.18)
33.95 (+1.30)
35.88 (+3.23)
39.48 (+6.83)
Table 1: Dataset-level General-10 results for matched FoldingCorpus-format controls and the full Fold2Reason recipe. Values are mean accuracies in percent over three seeds; parentheses report absolute change from Base in percentage points. Categories follow Figure 1 (b). Green/red cells indicate gains/declines; stronger shading indicates larger absolute changes on a shared scale. The three controls share the same LoRA recipe, supervised-token budget, seeds, and full General-10 evaluation.
Figure 3: FoldingCorpus supervision across model scales and families. (a) General-10 change relative to each model’s own base checkpoint. Bars show the three-seed mean and open circles show the individual seeds. (b) The same change resolved per benchmark, with one marker shape and colour per model and the categories of Figure 1 (b). The horizontal axis is broken to accommodate the large SpatialViz gains of the two smaller Qwen3.5 models.
Figure 4: Data scaling of structural readout and general transfer. The three panels report FoldBench lDDT-C α , FoldBench Contact F1, and General-10 change for an independent three-seed Qwen3.5-9B Full-RG scaling run over seven nested subsets spanning 50–4,000 proteins. Every subset is trained for three epochs. Thin curves show individual seeds, center curves show means, and transparent bands with boundary lines show three-seed 95% t -intervals. The upper axis gives the corresponding number of FoldingCorpus records at 12 targets per protein.
Arm
G10
3D macro
FTB
SpatialViz
VSI
Text G7
Base
45.09
40.44
36.18
26.61
58.53
47.09
w/o FoldingCorpus
45.77 (+0.68)
41.47 (+1.03)
38.44 (+2.26)
27.40 (+0.79)
58.58 (+0.05)
47.61 (+0.52)
w/o Geometry
48.02 (+2.93)
43.30 (+2.86)
39.00 (+2.81)
31.81 (+5.20)
59.08 (+0.54)
50.05 (+2.96)
Fold2Reason-full
48.33 (+3.23)
44.38 (+3.94)
40.26 (+4.08)
32.71 (+6.10)
60.16 (+1.63)
50.02 (+2.93)
Table 2: Component results on external benchmarks over three seeds. Values are mean accuracies in percent; parentheses report absolute change from Base in percentage points. The 3D macro averages FTB-Core, SpatialViz, and VSI; Text G7 averages the remaining seven General-10 datasets.
Figure 5: Contact and distance maps for three example proteins. Rows show three FoldBench334 proteins of increasing length (8wt3_A, L=134 ; 8qjp_A, L=250 ; 7xg9_A, L=286 ), comparing ground truth, FoldingCorpus-only, and Fold2Reason. Contacts use C α distance <8 Å and sequence separation ≥6 ; predicted maps show frequency across all three seeds, with gray diagonal bands marking excluded near-sequence pairs. Distance maps show the three-seed mean, clipped at 20 Å. Panels are unsmoothed and cover the complete protein.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Base model
Qwen3.5-9B
Protein views
336 sequence; 332 +MSA; 332 +MSA+template
Length range / mean
29–199 / 144.42 residues
FoldingCorpus records
12 per protein; 12,000 total
Packed supervision
12 labels + EOS per protein
Workspace
width 256; 16 pooled tokens; at most 2,048 pairs
Appendix
Table 3: Canonical Fold2Reason-full training contract.
Partition
Proteins
Questions/protein
FoldingCorpus records
Packed samples
Train
1,000
12
12,000
1,000
Development
100
12
1,200
100
Frozen test
100
12
1,200
100
Total
1,200
12
14,400
1,200
Appendix
Table 4: Protein data partitions and supervision counts. A FoldingCorpus record is one operator applied to one protein; a packed sample contains all 12 records.
Operator / train labels
Definition and threshold
Sampling rule
CONTACT_SHORT A470/B530
A iff d(i,j)<8 Å
6≤sep≤11
CONTACT_MEDIUM A458/B542
A iff d(i,j)<8 Å
12≤sep≤23
CONTACT_LONG A447/B553
A iff d(i,j)<8 Å
sep≥24
DISTANCE_ORDER_1 A481/B519
A iff d(i,j)<d(k,l)
low/high distance deciles; local/medium pairs
DISTANCE_ORDER_2 A512/B488
A iff d(i,j)<d(k,l)
low/high distance deciles; long-range pairs
SEGMENT_ORIENTATION A331/B297/C372
For endpoint vectors u,v : A if cos(u,v)≥0.5 ; B if cos(u,v)≤−0.5 ; C otherwise
segment starts separated by at least eight residues
Appendix
Table 5: The 12 FoldingCorpus operators. Counts are observed labels in the 1,000-protein training partition. Indices refer to valid residues; sep is sequence separation.
Task
Answer/parser
Representative generated query
nearest_neighbor
point ID / exact
choose the nearest labeled 3D point
bond_angle
degrees / ±0.5
compute an angle from three coordinates
signed_dihedral
degrees / ±0.5
compute a signed four-point dihedral
tetrahedral_chirality
positive/negative
determine the handedness of four points
proper_rigid_equivalence
yes/no
distinguish a proper rigid transform from reflection
sparse_constraints_candidate_selection
A–D
choose the structure satisfying sparse distances
Appendix
Table 6: FTB-Core v1 task composition. The final column gives one prompt schema per task; each task contributes 500 test and 500 test-OOD examples to the headline score.
Benchmark (revision)
n
Modality
New tokens
Score / extraction
FTB-Core v1.0.0
12,000
text
32
numeric tolerance or categorical EM
SpatialViz ( f38482a8 )
1,180
image+text
128
option-letter accuracy
VSI ( d7cb1a39 )
5,130
video+text
16
MC accuracy/MCA; aggregation below
GraphQA Easy ( 18f55c29 )
21,600
text
128
canonical answer-tail EM
GraphQA Hard ( b7e910dd )
21,600
text
128
canonical answer-tail EM
BBH ( 982bb89f )
6,511
text
128
normalized EM
Appendix
Table 7: Per-benchmark evaluation contract. All rows use greedy decoding. EM denotes the benchmark-specific normalized exact matcher; MCA is mean relative accuracy.
Item
Protocol
Initialization
Pretrained Phase-0 head-only decoder
Trainable in Phase 0
Coordinate and distogram heads only; base LM and LoRA frozen
Phase-0 data
1,000 training proteins from the OpenFold high-confidence split
Fold2Reason training
Decoder loaded with requires_grad=False and excluded from optimization
Held-out exclusions
No FoldBench334, General-10, FTB-Core, SpatialViz, VSI, GraphQA, BBH, ChemBench, ChemBench4K, Lab-Bench, or SciBench examples
Extra structural data
No additional PDB-derived proteins beyond the 1K training split
Appendix
Table 8: Frozen geometry decoder provenance.
Input: training proteins (s,v,Y,m,w) , tokenizer T , fixed data seed sd ; v is the assigned evidence view.
Output: frozen protein inputs x , marker indices I , FoldingCorpus tokens z , masked labels ℓ , and geometry targets.
1
Compute valid C α distances and structural summaries from Y,m .
2
Build the training-pool hard-negative index used by the 32-way summary question.
3
For each protein p in the training partition:
4
Serialize (sp,vp) and a residue skeleton with exactly one marker per residue; tokenize to obtain (xp,Ip) .
5
For each of the 12 operators in Table 5 :
Appendix
Algorithm 1 Construct a FoldingCorpus training cache
Input: frozen cache, base θ0 , pretrained decoder gψ , training seed, three-epoch schedule.
Output: per-example predictions, per-seed metrics, and seed-aggregated results.
1
For each training seed, load its declared checkpoint in evaluation mode.
2
Held-out FoldingCorpus questions: for each protein, compute its matched workspace memory for Full; omit memory for Pure-LoRA.
3
Ask each question independently using the sequence and that question; do not append earlier gold answers.
4
Take the first next-token argmax over the complete vocabulary and compare with the canonical label token.
5
Average correctness within each operator, then average the 12 operator accuracies. Record candidate-restricted accuracy separately as a diagnostic.
Appendix
Algorithm 3 Evaluate fixed checkpoints on FoldingCorpus, FoldBench, and General-10
Model
Epochs
Steps
Peak LR
LoRA (M)
GPUs/seed
Qwen3.5-2B
3
375
1e-04
16.82
4
Qwen3.5-4B
3
375
1e-04
32.46
4
Qwen3.5-9B
3
375
1e-04
43.28
4
InternVL3.5-8B
3
375
1e-04
43.65
4
Gemma-4-12B-IT
3
375
1e-04
65.57
4
Appendix
Table 9: Model-family training contracts from the saved run configurations. All arms use rank-16, alpha-32 LoRA with dropout 0.05; the parameter count includes only trainable adapters.
Model
S1
S2
S3
Mean ± SD
95% t interval
Qwen3.5-2B
+6.448
+5.048
+4.747
5.414±0.907
[+3.160, +7.668]
Qwen3.5-4B
+4.636
+5.155
+4.742
4.844±0.274
[+4.164, +5.525]
Qwen3.5-9B
+3.216
+3.204
+3.249
3.223±0.024
[+3.164, +3.281]
InternVL3.5-8B
+1.754
+1.225
+1.607
1.529±0.273
[+0.851, +2.206]
Gemma-4-12B-IT
-0.690
+0.509
+0.398
0.072±0.663
[-1.574, +1.718]
Appendix
Table 10: General-10 gains by model and training seed. Each row is paired with its own base checkpoint and archived evaluation contract.
Benchmark
Qwen 2B
Qwen 4B
Qwen 9B
InternVL 8B
Gemma 12B
FTB-Core
+10.45
+7.46
+4.10
+2.64
-1.23
SpatialViz
+21.21
+22.20
+6.69
-0.45
-0.79
VSI
-0.28
+3.99
+1.50
+0.20
+2.95
GraphQA Easy
+2.80
-2.82
+1.96
+0.22
-0.43
GraphQA Hard
+6.51
+3.69
+9.95
+3.42
-1.93
BBH
+1.35
+2.91
+2.38
+1.35
+1.53
Appendix
Table 11: Complete Pure-LoRA mean changes from each model’s own base (percentage points). All benchmarks and model settings are retained.
FTB-Core task
Base
Full
Δ
SpatialViz task
Base
Full
Δ
Bond angle
0.60
0.23
-0.37
2DRotation
0.00
5.00
+5.00
Distance audit
9.40
7.50
-1.90
3DRotation
30.00
35.42
+5.42
Fragment assembly
28.30
30.10
+1.80
3ViewProjection
31.00
38.67
+7.67
Nearest neighbor
41.50
50.47
+8.97
ArrowMoving
30.00
33.75
+3.75
Noisy template
25.80
26.97
+1.17
BlockMoving
31.25
31.25
+0.00
Rigid equivalence
47.20
48.20
+1.00
CrossSection
14.17
11.39
-2.78
Appendix
Table 12: Complete task-level mean scores for Full. Scores are percentages; changes are percentage points. Abbreviated FTB task names follow the construction table.
Task
n
Base PF
Full PF
Base VO
Full VO
Common Δ
all
12000
0.00
0.00
36.18
40.26
+4.08
Bond angle
1000
0.00
0.00
0.60
0.23
-0.37
Constraint audit
1000
0.00
0.00
9.40
7.50
-1.90
Fragment assembly
1000
0.00
0.00
28.30
30.10
+1.80
Nearest neighbor
1000
0.00
0.00
41.50
50.47
+8.97
Template selection
1000
0.00
0.00
25.80
26.97
+1.17
Appendix
Table 13: FTB-Core parser audit. PF: parse-failure rate (percent of all rows); VO: accuracy conditional on a valid output (percent). Common-valid Δ uses the same IDs for Base and Full in each seed. Full values are means of three seed-wise rates, not pooled votes. Overall rates use question weighting.
Task
n
Base PF
Full PF
Base VO
Full VO
Common Δ
all
1180
1.69
0.00
27.07
32.71
+5.46
2DRotation
80
0.00
0.00
0.00
5.00
+5.00
3DRotation
80
0.00
0.00
30.00
35.42
+5.42
3ViewProjection
100
0.00
0.00
31.00
38.67
+7.67
ArrowMoving
80
0.00
0.00
30.00
33.75
+3.75
BlockMoving
80
0.00
0.00
31.25
31.25
+0.00
Appendix
Table 14: SpatialViz parser audit. PF: parse-failure rate (percent of all rows); VO: accuracy conditional on a valid output (percent). Common-valid Δ uses the same IDs for Base and Full in each seed. Full values are means of three seed-wise rates, not pooled votes. Overall rates use question weighting.
Benchmark
Base
w/o FoldingCorpus
w/o Geometry
Fold2Reason-full
FTB-Core
36.18
38.44 (+2.26)
39.00 (+2.81)
40.26 (+4.08)
SpatialViz
26.61
27.40 (+0.79)
31.81 (+5.20)
32.71 (+6.10)
VSI
58.53
58.58 (+0.05)
59.08 (+0.54)
60.16 (+1.63)
GraphQA Easy
65.14
67.34 (+2.20)
68.72 (+3.58)
69.70 (+4.56)
GraphQA Hard
32.65
32.63 (-0.02)
44.44 (+11.79)
39.48 (+6.83)
BBH
54.45
54.64 (+0.19)
55.17 (+0.72)
56.81 (+2.36)
Appendix
Table 15: Dataset-level component ablation results. Values are mean accuracies in percent over three seeds; parentheses report absolute change from Base in percentage points. The w/o FoldingCorpus arm removes FoldingCorpus CE, the w/o Geometry arm removes the Geometry objective, and Fold2Reason-full combines FoldingCorpus CE and Geometry. The historical w/o FoldingCorpus arm is geometry-dominant and includes the legacy retrieval auxiliary.
Comparison
S1
S2
S3
Mean ± SD
95% t interval
w/o FoldingCorpus
+0.688
+0.783
+0.556
0.676±0.114
[+0.392, +0.959]
w/o Geometry
+2.611
+2.629
+3.551
2.931±0.538
[+1.595, +4.266]
Full RG
+3.186
+3.115
+3.400
3.234±0.148
[+2.866, +3.601]
Full − w/o Geometry
+0.575
+0.486
-0.152
0.303±0.396
[-0.681, +1.287]
Appendix
Table 16: Component gains and the paired Geometry increment on General-10 (percentage points).
Readout condition
TM-score ↑
lDDT-C α↑
Contact F1 ↑
C α MAE (Å) ↓
Frozen-base readout
0.1742
0.2400
0.0664
10.753
w/o Geometry
0.1710
0.2435
0.0627
12.096
w/o FoldingCorpus
0.1685
0.2534
0.0696
12.262
Fold2Reason-full
0.1688
0.2532
0.0690
12.244
Appendix
Table 17: Complete FoldBench334 structural readout comparison over three seeds. The frozen-base row trains a decoder of the same architecture on frozen Qwen features; the three adapted rows use the Phase-0 decoder as a fixed readout. The historical w/o FoldingCorpus arm is geometry-dominant and includes a legacy retrieval auxiliary. Lower C α MAE is better.
Figure 6: Geometry changes local readouts more consistently than global topology. Each point is one FoldBench334 protein after averaging that protein over three training seeds; distributions show Fold2Reason-full minus FoldingCorpus-only. Thick bars span the interquartile range, white circles mark medians, and diamonds mark means. Positive values favor Full except for C α MAE, where negative is better. Full improves per-protein lDDT-C α and Contact F1 for 75.4% and 78.7% of proteins, respectively, while TM-score and MAE contain influential tails.
Selection
Protein
L
TM-score FC-only → Full
lDDT-C α FC-only → Full
Contact F1 FC-only → Full
Distance MAE (Å) FC-only → Full
Panel A: proteins displayed in Figure 5
Rank 1
8wt3_A
134
0.2438 → 0.2742 (+0.0304)
0.2767 → 0.3034 (+0.0267)
0.1146 → 0.1169 (+0.0023)
6.0422 → 5.3820 (-0.6602)
Rank 2
8qjp_A
250
0.2255 → 0.2586 (+0.0331)
0.2462 → 0.2629 (+0.0167)
0.0377 → 0.0552 (+0.0175)
10.8590 → 9.1586 (-1.7004)
Rank 3
7xg9_A
286
0.2310 → 0.2568 (+0.0258)
0.2607 → 0.2834 (+0.0227)
0.0435 → 0.0493 (+0.0058)
12.2372 → 10.8577 (-1.3794)
Panel B: Contact-F1-change quantiles displayed in Figure 7
Lower (10th)
7urp_A
159
0.1917 → 0.2035 (+0.0118)
0.2125 → 0.2275 (+0.0150)
0.0707 → 0.0595 (-0.0112)
9.3934 → 8.0564 (-1.3371)
Appendix
Table 18: Preselected FoldBench334 cases for qualitative structural audit. Panel A matches Figure 5 : the three proteins have the highest Full mean TM-score among all 334 targets. Panel B matches Figure 7 : proteins are selected at the 10th, 50th, and 90th percentiles of the per-protein Contact F1 change. Selection uses the three-seed mean, with protein ID breaking ties, and precedes visual inspection. Each metric cell reports FC-only → Full, followed by Full minus FC-only in parentheses; lower distance MAE is better.
Figure 7: Local distance-map audit for the preselected proteins of Table 18 . For each protein, the anchor is the residue with the most ground-truth contacts; the displayed neighborhood contains that anchor and its 15 nearest ground-truth C α neighbors. Ground truth alone determines the selection. Full changes the local distance MAE from 6.311 to 6.178 Å in the lower case, from 2.620 to 2.821 Å in the median case, and from 5.982 to 5.533 Å in the higher case. The middle case worsens while the lower and higher cases improve, matching the heterogeneous Geometry effect in the aggregate analysis. Green squares mark the anchor residue.
Representation
N
Min
Q1
Median
Q3
Max
Mean
Manifest sequence
334
28.0
140.2
225.5
344.8
1414.0
261.43
Evaluated residues
334
26.0
140.2
225.5
344.8
1414.0
261.10
Appendix
Table 19: FoldBench334 cohort size and residue-count distribution. Manifest sequence length and the residue count in the structural evaluation cache are reported separately.
Proteins
FoldingCorpus targets
Steps
FoldingCorpus macro
lDDT-C α
Contact F1
General-10 Δ
50
600
21
0.478
0.244
0.0666
0.49
100
1,200
39
0.486
0.244
0.0671
0.61
250
3,000
96
0.485
0.246
0.0666
1.37
500
6,000
189
0.492
0.247
0.0675
2.92
1,000
12,000
375
0.501
0.252
0.0694
3.31
2,000 †
24,000
750
0.546
0.252
0.0739
3.70
Appendix
Table 20: Fixed-epochs data-scaling endpoints. Values are three-seed means from the independent scaling run. General-10 values are changes from the base model in percentage points.
Figure 8: Training and transfer dynamics at pre-registered checkpoints. Columns show training losses, held-out FoldingCorpus answer accuracy, FoldBench structural readouts, and General-10 changes. Legend entries carry the run identifiers full_rg for Fold2Reason-full (blue) and pure_lora for FoldingCorpus-only Pure-LoRA (orange). The upper axis reports cumulative protein and answer-label exposures. All benchmark evaluations were run offline after training and were not used for checkpoint selection.
Dataset
N
Canonical base
Scaling base
FTB-Core
12000
36.18
36.23
SpatialViz
1180
26.61
26.61
VSI
5130
58.53
56.93
GraphQA Easy
21600
65.14
67.71
GraphQA Hard
21600
32.65
33.39
BBH
6511
54.45
54.00
Appendix
Table 21: The two frozen base-evaluation contracts used by the canonical endpoint study and the independent scaling study. Each reported gain uses the base from its own column.
Proteins
Seed
Steps
FC acc.
lDDT
Contact F1
General-10 Δ
50
20260729
21
0.4608
0.2384
0.0650
+0.406
50
20260803
21
0.4875
0.2458
0.0667
+0.686
50
20260804
21
0.4858
0.2474
0.0680
+0.392
100
20260729
39
0.4758
0.2432
0.0660
+0.232
100
20260803
39
0.4975
0.2432
0.0670
+0.593
100
20260804
39
0.4850
0.2469
0.0682
+1.013
Appendix
Table 22: All fixed-three-epoch data-scaling endpoints by seed. FoldingCorpus and structure metrics are fractions; General-10 changes are percentage points.
Dataset
50
100
250
500
1K
2K †
4K †
FTB-Core
-0.13
-0.44
+0.50
+4.26
+4.26
+5.00
+4.56
SpatialViz
+0.76
+0.65
+1.53
+5.17
+5.93
+7.12
+6.92
VSI
+0.12
-0.08
+0.33
+1.17
+1.13
+1.82
+1.58
GraphQA Easy
+1.36
+2.13
+1.77
+2.58
+2.49
+2.76
+1.47
GraphQA Hard
+0.78
+0.56
+1.94
+3.90
+6.35
+8.67
+8.46
BBH
+0.57
+0.74
+2.37
+3.30
+3.77
+3.68
+2.90
Appendix
Table 23: Dataset-level changes along the complete fixed-three-epoch curve (percentage points). The final row reports the macro mean and training-seed SD.
Figure 9: Dataset-level transfer along the fixed-three-epoch curve. Cells show the three-seed mean change from the scaling study’s frozen base in percentage points. A common diverging color scale is centered at zero. Asterisks mark the fresh 2K and 4K extension runs; the 4K pool also changes the confidence-quality support described above.
Labels/protein
S1
S2
S3
Δ± SD
FC acc.
lDDT
Contact F1
3
+2.971
+3.731
+3.472
3.391±0.386
0.4942±0.0175
0.2544±0.0020
0.0690±0.0007
6
+3.145
+2.441
+3.268
2.951±0.446
0.4994±0.0047
0.2542±0.0028
0.0693±0.0003
12
+3.145
+2.941
+3.848
3.311±0.476
0.5014±0.0155
0.2523±0.0012
0.0694±0.0005
Appendix
Table 24: FoldingCorpus-label density at fixed 1,000 proteins and 375 optimizer steps. S1–S3 and the first mean report General-10 changes (pp); the remaining metrics are fractions.
Figure 10: FoldingCorpus-label density with 1,000 proteins and 375 steps. Solid lines show means over three training seeds. Shaded regions extend from mean minus one sample SD to mean plus one sample SD; the thin boundary lines mark those limits. FoldingCorpus accuracy uses the frozen test split. FoldBench uses all 334 targets, and external transfer uses the scaling evaluation contract.
File
Contents
endpoint_seed_scores.json
All ten dataset scores by seed for Fold2Reason-full, the component arms, the model families, and the source-control aggregates.
model_training_contracts.json
Saved LoRA modules, parameter counts, steps, seed, and hardware world size for the fifteen model-family runs.
task_breakdown.json
Complete FTB split/family/task, SpatialViz category/task, and VSI question-group scores.
foldbench_audit.json
The 334 target IDs, sequence and evaluation length distributions, and structural metrics by arm and seed.
scaling_evidence.json
Fixed-three-epoch endpoints, label-density results, and numerical checkpoint trajectories.
scaling_evaluation_manifest.json
Archived checkpoint and split assignments and expected evaluation counts for the included original scaling runs.
Appendix
Table 25: Machine-readable evidence included with the appendix source package.
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order. However, applying artificial intelligence to this process presents a joint challenge of representation and reasoning: models must preserve domain-native structural information while showing how specific evidence supports predictions under these constraints. Here we introduce SciReasoner, a multimodal scientific foundation model for native structural reasoning across proteins, small molecules and inorganic crystals. SciReasoner discretizes coordinates, topologies and periodic connectivities into a unified structure-aware vocabulary, treating structural tokens as addressable evidence units during reasoning. In homology-controlled Gene Ontology prediction, SciReasoner improves Cellular Component annotation for low-homology and orphan-like proteins, increasing Fmax from 0.42 to 0.55. In chemistry, it raises single-step retrosynthesis accuracy from 0.63 to 0.72 while generating fragment-level disconnection and precursor-verification traces. In materials science, its representations separate elemental and compound phases and resolve high- and low-band-gap regimes. Across 86 benchmarks, SciReasoner achieves state-of-the-art performance on 67 tasks. Double-blind expert evaluation rates its reasoning traces as preferred or at least comparable to those of a frontier large language model in 98% of cases. By making structure an inspectable substrate for reasoning under scientific constraints, SciReasoner connects accurate prediction with interpretable scientific inference.
Chen Tang, Yizhou Wang, Jianyu Wu +26
Shanghai Artificial Intelligence Laboratory, China. · The Chinese University of Hong Kong, Hong Kong. · Shanghai Jiao Tong University, China. +8
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through post-training, yet how each stage shapes reasoning and generalization remains poorly understood. We study when post-training improves performance and when it induces over-specialization. Across genomics, transcriptomics, and proteins, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that each post-training stage reshapes generalization in a distinct way rather than contributing uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that biological reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed post-training budgets, the strongest ID-OOD trade-off comes from brief SFT, larger RL allocations, and asymmetric adaptation capacity across stages.
Lukas Fesser, Hanlin Zhang, Michelle M. Li +5
Harvard University · Google DeepMind · Google Research
When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic protocol using a minimal, target-label-free additive correction. Fitting just two parameters on as few as 25 unlabeled examples recovers 9--34 accuracy points for Qwen3.5 models, transferring successfully to OLMo-2-1B and Llama-3.1-8B. Crucially, these recovered decisions persist on hard instances unresolved by simple lexical overlap and significantly exceed count-preserving permutation baselines. Our results show that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.
Qiyao Yan, Chenpeng Wang, Liangming Pan
State Key Laboratory of Multimedia Information Processing, Peking University · School of Computer Science, Peking University · YiXin-AILab, YIXIN, Beijing, China +1