Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
Figures & tables
Figure 1: Overview of Latent-JEPA . (a) Anticipating breakage or a poem’s theme and rhyme illustrates prediction at the level of useful abstractions. The verse gives one possible continuation of the first line. (b) A continuous reasoner updates latent thoughts before generating the remaining reasoning and molecular outcome. (c) The final latent thought also predicts textual and molecular representations of the reference solution during training. The molecular optimization illustration replaces an aldehyde with a carboxyl group; R denotes the unchanged scaffold.
Molecule editing (%)
Property improvement
Method
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
General-purpose language models
GPT-4o
80.0
80.0
65.0
-0.09
0.92
0.13
0.07
-0.02
0.00
o1-mini
55.0
80.0
58.3
-0.42
1.78
0.07
-0.03
-0.10
-0.08
o3-mini
65.0
55.0
80.0
0.26
0.81
0.21
0.19
-0.03
0.01
Gemini 2.5 Pro (thinking)
100.0
85.0
81.7
-0.22
1.06
0.28
0.36
-0.02
0.06
Table 1: Molecule editing and optimization on ChemCoTBench. Editing reports correctness (%); optimization reports mean property improvement. All metrics are higher-is-better. Dark-red bold marks Latent-JEPA scores equal to or better than Coconut-Chem. Optimization success rates appear in Table 11 .
Molecular understanding
Major product
Byproduct
Method
FG ↓
Ring ↓
Murcko ↑
Ring-sys. (%) ↑
FTS ↑
Valid (%) ↑
FTS ↑
Valid (%) ↑
General-purpose language models
GPT-4o
0.17
1.35
0.21
80.0
0.58
–
0.20
–
o1-mini
0.21
1.25
0.25
61.7
0.31
–
0.17
–
o3-mini
0.13
0.60
0.39
75.0
0.71
–
0.27
–
Gemini 2.5 Pro (thinking)
0.11
0.60
0.51
87.5
0.89
–
0.51
–
Table 2: Molecular understanding and forward-reaction prediction. FG and Ring report count MAE; Murcko and FTS report similarity; ring-system accuracy and reaction validity are percentages. Dark-red bold marks Latent-JEPA scores equal to or better than Coconut-Chem. Ring-system scoring and the treatment of published reference scores are specified in Appendix L.2 .
Figure 2: Molecular-outcome prediction on examples excluded from reasoner training and probe fitting. A: Ridge-probe cosine similarity across recurrent updates. B: Within-task Hit@5; the dashed line denotes uniform retrieval. C: Corresponding-target similarity versus within-task shuffled targets. Bands summarize evaluation-example uncertainty; Appendix K specifies the protocol.
Figure 3: Representational similarity between recurrent feedback embeddings and molecular outcomes. A: Agreement with calibrated SMI-TED target distances. B: Agreement with Morgan fingerprint Tanimoto distances. Curves are task-averaged within-task Spearman RSA and single-run point estimates. Appendix K specifies representations, configurations, and distance measures.
Target calibration
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Both views centered
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Without molecular centering
0.8701
0.9703
0.2151
0.2645
0.1163
0.2329
Without text centering
0.7854
0.8311
0.2036
0.2644
0.1174
0.2232
Table 3: Effect of target centering on molecular optimization. All entries are mean property improvements ( ↑ ). Full task results are reported in Table 6 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
LatentChem
LatentChem + Latent-JEPA
Molecule editing (%) ↑
Addition
65.0
63.3
Deletion
78.3
85.0
Substitution
56.7
60.0
Molecular optimization: property improvement ↑
LogP
0.7051
0.8040
Appendix
Table 4: Answer-representation prediction with LatentChem. Both models receive three epochs of supervised adaptation with the molecular updater frozen and are evaluated at T=0.7,p=0.9,n=3 . Latent-JEPA predicts an independently encoded answer representation from the final latent thought. In the Latent-JEPA column, dark-red bold indicates improvement over the adapted LatentChem baseline. Editing accuracy, ring-system accuracy, and reaction validity are percentages; optimization entries are mean property improvements, count entries are MAE, and Murcko and FTS are similarities.
Molecule editing (%) ↑
Property improvement ↑
Text target
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Online, dropout
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Online, no dropout
80.0
80.0
56.1
0.8613
0.9398
0.2168
0.2705
0.1104
0.2222
EMA, no dropout
81.7
80.0
64.4
0.8410
0.9549
0.2308
0.2957
0.1295
0.2278
Appendix
Table 5: Text-target encoder ablation. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE; Murcko and FTS are similarities. Dark-red bold marks the best task score in each column, including ties at the displayed precision. Shading identifies the default configuration. Lower average rank is better.
Molecule editing (%) ↑
Property improvement ↑
Target calibration
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Both views centered
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Without molecular centering
75.0
73.3
62.2
0.8701
0.9703
0.2151
0.2645
0.1163
0.2329
Without text centering
78.3
75.0
58.9
0.7854
0.8311
0.2036
0.2644
0.1174
0.2232
Appendix
Table 6: Target-centering ablation with both predictive views and ℓ2 normalization retained. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE; Murcko and FTS are similarities. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies the default configuration. Lower average rank is better.
Objective
Center removed
Metrics
Task-score changes
Original / shuffled distance
Text only
Text
17
10 / 3 / 4
0.102 / 0.330
Molecular only
Molecular
17
4 / 3 / 10
0.003 / 0.009
Dual
Molecular
17
4 / 3 / 10
0.003 / 0.009
Dual
Text
17
3 / 1 / 13
0.095 / 0.217
Appendix
Table 7: Target centering and source-specific prediction. Task-score changes count improved, unchanged, and worse metrics relative to the corresponding centered model after removing the indicated center. The final column gives cosine distances using the original and shuffled latent sources. Settings and aggregation are described in the surrounding text.
Prediction views
Add ↑
Delete ↑
Sub. ↑
LogP ↑
Solubility ↑
QED ↑
DRD2 ↑
JNK3 ↑
GSK3- β↑
Textual
81.7
80.0
58.9
0.7839
0.9178
0.2229
0.2684
0.1100
0.2415
Molecular (SMI-TED)
83.3
90.0
62.8
0.8624
0.9359
0.2333
0.2579
0.1141
0.2442
Textual + molecular
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Appendix
Table 8: Chemical performance with different prediction views. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies dual prediction. Loss coefficients and rank aggregation are specified in the text.
Molecule editing (%) ↑
Property improvement ↑
Prediction views
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Vanilla Coconut
71.7
88.3
56.1
0.7726
0.7250
0.1968
0.2300
0.0861
0.1861
Textual
76.7
86.7
56.1
0.8656
0.9001
0.2048
0.2268
0.0992
0.1875
Molecular (SMI-TED)
68.3
85.0
57.2
0.7960
0.8107
0.2105
0.2420
0.0769
0.2153
Textual + molecular
66.7
90.0
55.6
0.8429
0.8799
0.2307
0.2376
0.0796
0.1972
Appendix
Table 9: Prediction-view comparison in Vanilla Coconut without source ChemTokens. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE. Training and evaluation settings are given in the text. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies textual prediction. Lower average rank is better.
Figure 10: Example-wise molecular-outcome prediction from the final latent thought. Each point compares probe cosine for Coconut-Chem (horizontal) and dual-view Latent-JEPA (vertical); colors denote task families. Points above the diagonal favor Latent-JEPA . The probe split is shared with Figure 2 .
Figure 12: Molecular-outcome retrieval by task and recurrent depth. A: Coconut-Chem Hit@5. B: Dual-view Latent-JEPA Hit@5. C: Their difference; dots mark 95% bootstrap intervals excluding zero. The heatmaps use the same latent hidden states, probe holdout, and within-task candidates as Figure 2 . Averaging each column over tasks recovers the corresponding aggregate Hit@5 curve; Appendix K gives the protocol.
Analysis
Cohort
Main linear probing and Hit@5
1,224 training examples across 13 tasks for probe fitting; 321 test examples across six tasks for evaluation
SMI-TED RSA
321 molecular-target-eligible test examples across six tasks
Morgan RSA
294 test examples with valid molecular outcomes across the same six tasks
Latent PCA and task organization
16,943 training examples; the separate retrieval analysis holds out 3,390 from probe fitting
Target calibration and task affinity
16,949 molecular training targets
Appendix
Table 10: Cohorts used for representation analysis, with counts in examples.
Method
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
LatentChem
77.7
73.0
78.3
63.7
45.7
69.3
Coconut-Chem
82.3
85.0
89.0
77.3
58.3
81.3
Coconut-Chem + Latent-JEPA
82.3
86.0
86.3
79.3
64.0
84.0
Appendix
Table 11: Molecular optimization success rates (%). Dark-red bold values in the Latent-JEPA row match or improve upon Coconut-Chem at the reported precision; regressions remain black.