Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
Figures & tables
Figure 1: Overview of Latent-JEPA . (a) Anticipating breakage or a poem’s theme and rhyme illustrates prediction at the level of useful abstractions. The verse gives one possible continuation of the first line. (b) A continuous reasoner updates latent thoughts before generating the remaining reasoning and molecular outcome. (c) The final latent thought also predicts textual and molecular representations of the reference solution during training. The molecular optimization illustration replaces an aldehyde with a carboxyl group; R denotes the unchanged scaffold.
Molecule editing (%)
Property improvement
Method
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
General-purpose language models
GPT-4o
80.0
80.0
65.0
-0.09
0.92
0.13
0.07
-0.02
0.00
o1-mini
55.0
80.0
58.3
-0.42
1.78
0.07
-0.03
-0.10
-0.08
o3-mini
65.0
55.0
80.0
0.26
0.81
0.21
0.19
-0.03
0.01
Gemini 2.5 Pro (thinking)
100.0
85.0
81.7
-0.22
1.06
0.28
0.36
-0.02
0.06
Table 1: Molecule editing and optimization on ChemCoTBench. Editing reports correctness (%); optimization reports mean property improvement. All metrics are higher-is-better. Dark-red bold marks Latent-JEPA scores equal to or better than Coconut-Chem. Optimization success rates appear in Table 11 .
Molecular understanding
Major product
Byproduct
Method
FG ↓
Ring ↓
Murcko ↑
Ring-sys. (%) ↑
FTS ↑
Valid (%) ↑
FTS ↑
Valid (%) ↑
General-purpose language models
GPT-4o
0.17
1.35
0.21
80.0
0.58
–
0.20
–
o1-mini
0.21
1.25
0.25
61.7
0.31
–
0.17
–
o3-mini
0.13
0.60
0.39
75.0
0.71
–
0.27
–
Gemini 2.5 Pro (thinking)
0.11
0.60
0.51
87.5
0.89
–
0.51
–
Table 2: Molecular understanding and forward-reaction prediction. FG and Ring report count MAE; Murcko and FTS report similarity; ring-system accuracy and reaction validity are percentages. Dark-red bold marks Latent-JEPA scores equal to or better than Coconut-Chem. Ring-system scoring and the treatment of published reference scores are specified in Appendix L.2 .
Figure 2: Molecular-outcome prediction on examples excluded from reasoner training and probe fitting. A: Ridge-probe cosine similarity across recurrent updates. B: Within-task Hit@5; the dashed line denotes uniform retrieval. C: Corresponding-target similarity versus within-task shuffled targets. Bands summarize evaluation-example uncertainty; Appendix K specifies the protocol.
Figure 3: Representational similarity between recurrent feedback embeddings and molecular outcomes. A: Agreement with calibrated SMI-TED target distances. B: Agreement with Morgan fingerprint Tanimoto distances. Curves are task-averaged within-task Spearman RSA and single-run point estimates. Appendix K specifies representations, configurations, and distance measures.
Target calibration
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Both views centered
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Without molecular centering
0.8701
0.9703
0.2151
0.2645
0.1163
0.2329
Without text centering
0.7854
0.8311
0.2036
0.2644
0.1174
0.2232
Table 3: Effect of target centering on molecular optimization. All entries are mean property improvements ( ↑ ). Full task results are reported in Table 6 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
LatentChem
LatentChem + Latent-JEPA
Molecule editing (%) ↑
Addition
65.0
63.3
Deletion
78.3
85.0
Substitution
56.7
60.0
Molecular optimization: property improvement ↑
LogP
0.7051
0.8040
Appendix
Table 4: Answer-representation prediction with LatentChem. Both models receive three epochs of supervised adaptation with the molecular updater frozen and are evaluated at T=0.7,p=0.9,n=3 . Latent-JEPA predicts an independently encoded answer representation from the final latent thought. In the Latent-JEPA column, dark-red bold indicates improvement over the adapted LatentChem baseline. Editing accuracy, ring-system accuracy, and reaction validity are percentages; optimization entries are mean property improvements, count entries are MAE, and Murcko and FTS are similarities.
Molecule editing (%) ↑
Property improvement ↑
Text target
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Online, dropout
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Online, no dropout
80.0
80.0
56.1
0.8613
0.9398
0.2168
0.2705
0.1104
0.2222
EMA, no dropout
81.7
80.0
64.4
0.8410
0.9549
0.2308
0.2957
0.1295
0.2278
Appendix
Table 5: Text-target encoder ablation. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE; Murcko and FTS are similarities. Dark-red bold marks the best task score in each column, including ties at the displayed precision. Shading identifies the default configuration. Lower average rank is better.
Molecule editing (%) ↑
Property improvement ↑
Target calibration
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Both views centered
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Without molecular centering
75.0
73.3
62.2
0.8701
0.9703
0.2151
0.2645
0.1163
0.2329
Without text centering
78.3
75.0
58.9
0.7854
0.8311
0.2036
0.2644
0.1174
0.2232
Appendix
Table 6: Target-centering ablation with both predictive views and ℓ2 normalization retained. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE; Murcko and FTS are similarities. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies the default configuration. Lower average rank is better.
Objective
Center removed
Metrics
Task-score changes
Original / shuffled distance
Text only
Text
17
10 / 3 / 4
0.102 / 0.330
Molecular only
Molecular
17
4 / 3 / 10
0.003 / 0.009
Dual
Molecular
17
4 / 3 / 10
0.003 / 0.009
Dual
Text
17
3 / 1 / 13
0.095 / 0.217
Appendix
Table 7: Target centering and source-specific prediction. Task-score changes count improved, unchanged, and worse metrics relative to the corresponding centered model after removing the indicated center. The final column gives cosine distances using the original and shuffled latent sources. Settings and aggregation are described in the surrounding text.
Prediction views
Add ↑
Delete ↑
Sub. ↑
LogP ↑
Solubility ↑
QED ↑
DRD2 ↑
JNK3 ↑
GSK3- β↑
Textual
81.7
80.0
58.9
0.7839
0.9178
0.2229
0.2684
0.1100
0.2415
Molecular (SMI-TED)
83.3
90.0
62.8
0.8624
0.9359
0.2333
0.2579
0.1141
0.2442
Textual + molecular
78.3
86.7
61.1
0.8318
0.9659
0.2296
0.2848
0.1296
0.2515
Appendix
Table 8: Chemical performance with different prediction views. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies dual prediction. Loss coefficients and rank aggregation are specified in the text.
Molecule editing (%) ↑
Property improvement ↑
Prediction views
Add
Delete
Sub.
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
Vanilla Coconut
71.7
88.3
56.1
0.7726
0.7250
0.1968
0.2300
0.0861
0.1861
Textual
76.7
86.7
56.1
0.8656
0.9001
0.2048
0.2268
0.0992
0.1875
Molecular (SMI-TED)
68.3
85.0
57.2
0.7960
0.8107
0.2105
0.2420
0.0769
0.2153
Textual + molecular
66.7
90.0
55.6
0.8429
0.8799
0.2307
0.2376
0.0796
0.1972
Appendix
Table 9: Prediction-view comparison in Vanilla Coconut without source ChemTokens. Editing, ring-system accuracy, and validity are percentages; property scores are mean improvements; counts are MAE. Training and evaluation settings are given in the text. Dark-red bold marks column-best task scores, including ties at the displayed precision; shading identifies textual prediction. Lower average rank is better.
Figure 10: Example-wise molecular-outcome prediction from the final latent thought. Each point compares probe cosine for Coconut-Chem (horizontal) and dual-view Latent-JEPA (vertical); colors denote task families. Points above the diagonal favor Latent-JEPA . The probe split is shared with Figure 2 .
Figure 12: Molecular-outcome retrieval by task and recurrent depth. A: Coconut-Chem Hit@5. B: Dual-view Latent-JEPA Hit@5. C: Their difference; dots mark 95% bootstrap intervals excluding zero. The heatmaps use the same latent hidden states, probe holdout, and within-task candidates as Figure 2 . Averaging each column over tasks recovers the corresponding aggregate Hit@5 curve; Appendix K gives the protocol.
Analysis
Cohort
Main linear probing and Hit@5
1,224 training examples across 13 tasks for probe fitting; 321 test examples across six tasks for evaluation
SMI-TED RSA
321 molecular-target-eligible test examples across six tasks
Morgan RSA
294 test examples with valid molecular outcomes across the same six tasks
Latent PCA and task organization
16,943 training examples; the separate retrieval analysis holds out 3,390 from probe fitting
Target calibration and task affinity
16,949 molecular training targets
Appendix
Table 10: Cohorts used for representation analysis, with counts in examples.
Method
LogP
Solubility
QED
DRD2
JNK3
GSK3- β
LatentChem
77.7
73.0
78.3
63.7
45.7
69.3
Coconut-Chem
82.3
85.0
89.0
77.3
58.3
81.3
Coconut-Chem + Latent-JEPA
82.3
86.0
86.3
79.3
64.0
84.0
Appendix
Table 11: Molecular optimization success rates (%). Dark-red bold values in the Latent-JEPA row match or improve upon Coconut-Chem at the reported precision; regressions remain black.
Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete text tokens, but this textual interface also introduces redundancy and inference overhead. Latent reasoning offers a promising alternative by carrying part of the computation in continuous representations. However, existing methods typically predefine when latent computation is invoked and how it is allocated during decoding, leaving a key problem unresolved: when to invoke latent computation, what type of computation to perform, and how much budget to allocate. We propose \textbf{Ty}ped \textbf{L}at\textbf{e}nt \textbf{R}easoning (Tyler), a typed and budget-aware framework for latent reasoning during autoregressive decoding. Tyler learns a policy that, at each decoding step, chooses between emitting a text token and switching to a latent computation module specialized for a particular reasoning function. Once invoked, an operator maps the current reasoning state into latent tokens that support global planning, local state updates, or reusable procedural abstraction. Across extensive experiments on three backbone LLMs, Tyler improves accuracy by up to 14.49 points over CoT and by up to 4.30 points over the strongest competing baseline. It further generalizes across diverse reasoning domains and achieves the best final-stage performance with the lowest forgetting.
Deployed large reasoning models (LRMs) often behave unexpectedly. Test-time steering controls LRM outputs by intervening on their hidden representations, but it can degrade output quality. We argue that prior steering work implicitly relies on internal features that detect behavior in already generated text. We show that these detection features are poor predictors of future behavioral outcomes, and thus not the natural intervention target. Instead, we train activation probes to predict future behavior likelihoods from intermediate reasoning steps. These probes predict the most likely behavior with 64%-91% accuracy, revealing a separate type of internal prediction features. Building on these prediction features, we introduce a text-level steering method, Future Probe Controlled Generation. FPCG samples multiple candidate sentences and chooses the best one according to a probe predicting the future behavior likelihood. This enables steering with almost no output quality degradation. FPCG also enables steering in several evaluations where activation steering fails. These results show that distinguishing detection and prediction features enables a more nuanced approach to controlling LRM behaviors.
Evgenii Kortukov, Piotr Komorowski, Florian Klein +5
1Fraunhofer HHI · 2Northeastern University · 3KAIST
Reasoning problems often admit multiple valid ways to proceed. Continuous reasoning promises to move computation beyond language tokens into a more compact latent space, but representing several plausible ways to think next remains difficult. We introduce Autoregressive Thought Flow (ATF), which models the next continuous thought as a multimodal distribution. A causal autoregressive model performs the reasoning computation, while a lightweight diffusion head generates a plausible next thought from the resulting condition. The sampled thought is fed back into the model, allowing continuous reasoning to unfold for a variable number of steps while preserving the pretrained backbone. Across mathematical reasoning tasks, ATF improves accuracy with compact latent traces and benefits from reinforcement learning and additional test-time thinking. Multi-sample evaluation shows broader solution coverage, indicating that its multimodal predictions capture useful diversity among reasoning paths. Our results suggest that continuous reasoning is more effective when multiple possible next thoughts remain available rather than being collapsed into a single prediction.
Yang Li, Yi Wang, Shiyuan Huang +3
Rutgers University · Amazon · University of Illinois at Urbana-Champaign