While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.
Figures & tables
Model
Molecule Mapping
Reaction Mapping
PDDL Grounding
Molecule ID
SMILES
Coverage
Domain
Problem
GPT-5.2
1.0000
0.9932
1.0000
0.9303
1.0000
DeepSeek V3.2
0.9992
0.7870
1.0000
0.9300
1.0000
Gemini 3.1
0.6000
0.5964
1.0000
0.8402
1.0000
Gemini 2.5 Flash
0.6000
0.5933
0.9996
0.8351
1.0000
Qwen3-30B-Thinking
0.2976
0.0285
0.7000
0.4670
0.9128
Table 1: Performance comparison across staged subtasks in retrosynthesis planning (averaged). Molecule Mapping includes Molecule ID (identifier consistency) and SMILES (exact-match string accuracy). Reaction Mapping (Coverage) measures correctness of reactant–product structure. PDDL Grounding includes Domain (action generation) and Problem (initial and goal specification).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
100
200
300
400
500
600
700
800
900
1000
Qwen3-30B-A3B-Thinking-2507
0.9913 (0.1652)
0.9954 (0.0868)
–
–
0.9896 (0.0329)
–
–
–
–
–
Qwen2.5-14B-Instruct-1M
0.9913 (0.9304)
–
–
–
–
–
–
–
–
–
DeepSeek V3.2
1.0 (1.0)
0.9954 (0.6256)
1.0 (0.9907)
1.0 (0.9928)
1.0 (0.9931)
0.9987 (0.1797)
1.0 (0.9800)
1.0 (0.9839)
0.9982 (0.1238)
1.0 (1.0)
GPT-5.2
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9938)
1.0 (0.9904)
1.0 (0.9913)
1.0 (0.9935)
1.0 (0.9944)
1.0 (0.9930)
1.0 (0.9928)
1.0 (0.9921)
Gemini 2.5 Flash
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9907)
1.0 (0.9876)
1.0 (0.9844)
–
1.0 (0.9789)
–
–
–
Gemini 3.1
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9907)
1.0 (0.9928)
–
1.0 (0.9961)
1.0 (0.9933)
–
–
–
Appendix
Table 2: Task 1 (Molecule Mapping) performance across varying dataset sizes. Each cell reports identifier exact-match rate and SMILES exact-match rate (in parentheses). “–” indicates invalid outputs.
Model
100
200
300
400
500
600
700
800
900
1000
Qwen3-30B-A3B-Thinking-2507
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
–
–
–
Qwen2.5-14B-Instruct-1M
1.0000
1.0000
–
–
–
–
–
–
–
–
DeepSeek V3.2
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
GPT-5.2
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
Gemini 2.5 Flash
1.0000
1.0000
1.0000
0.9975
1.0000
1.0000
1.0000
0.9988
1.0000
1.0000
Gemini 3.1
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
Appendix
Table 3: Task 2 (Reaction ID Mapping) integrity across varying dataset sizes. Each cell reports the exact mapping rate (EMR), defined as the proportion of reactions for which RXN_ID , Reactants_IDs , and Products_IDs exactly match the input. “–” indicates missing or invalid outputs (e.g., empty or unparsable model responses).
Model
100
200
300
400
GPT-5.2
1.0
1.0
1.0
1.0
DeepSeek V3.2
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
0.0
0.1176
0.2045
0.1531
Qwen2.5-14B-Instruct-1M
1.0
1.0
1.0
1.0
Gemini 2.5 Flash
1.0
1.0
1.0
1.0
Gemini 3.1
1.0
1.0
1.0
1.0
Appendix
Table 4: Syntactic validity rate of generated PDDL problems across different dataset sizes. A generation is considered valid if it satisfies basic structural constraints, including the correct problem definition, domain declaration, start molecule fact, goal predicates, and balanced parentheses.
Model
Metric
100
200
300
400
500
600
700
800
900
1000
GPT-5.2
Coverage
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Validity
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
DeepSeek V3.2
Coverage
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Validity
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
Coverage
1.0
1.0
0.9800
0.7960
0.5940
0
0
0
0.3380
0
Validity
1.0
1.0
0
0
0
0
0
0
0
0
Appendix
Table 5: Syntactic validity and action completion rate of generated PDDL domains across different dataset sizes (100–1000 reactions). Validity indicates whether the generated domain satisfies basic syntactic constraints such as correct domain structure, valid action blocks, and balanced parentheses. Completion measures the fraction of expected actions successfully generated (e.g., number of generated actions divided by the expected number of reactions).
Model
Metric
100
200
300
400
500
600
700
800
900
1000
GPT-5.2
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0.8522
0.8350
DeepSeek V3.2
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0.8489
0.8350
Qwen3-30B-A3B-Thinking-2507
Grounding
0.9900
0.9950
0.9733
0.7884
0.5880
0
0
0.3356
0
0
Qwen2.5-14B-Instruct-1M
Grounding
0.9900
0.7900
0.5367
0.3501
0.3240
0.2683
0.2300
0.2012
0.1622
0.1580
Gemini 2.5 Flash
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0
0.8350
Gemini 3.1
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.0767
0.8857
0.8612
0.8522
0.8350
Appendix
Table 6: Action grounding performance of generated PDDL domains across different dataset sizes (100–1000 reactions). Grounding measures whether each generated action correctly includes the corresponding reactant and product identifiers for the given reaction. This metric evaluates semantic correctness beyond syntactic validity and action coverage.
Model
Metric
100
200
300
400
GPT-5.2
Problem Coverage
1.0
1.0
1.0
1.0
DeepSeek V3.2
Problem Coverage
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
Problem Coverage
1.0
0.8897
0.8636
0.8980
Qwen2.5-14B-Instruct-1M
Problem Coverage
0.1304
0.3309
0.5909
0.5442
Gemini 2.5 Flash
Problem Coverage
1.0
1.0
1.0
1.0
Gemini 3.1
Problem Coverage
1.0
1.0
1.0
1.0
Appendix
Table 7: Problem coverage of generated PDDL problem files across dataset sizes (100–400 reactions). Problem coverage measures the fraction of ground-truth problem instances that are correctly generated, where each problem file is aligned with the input molecule (Mol_Name), and includes both the corresponding final product (Final_Product_ID) and all required goal candidate molecules.
Model
Solve Rate
Path Accuracy
GPT-5.2
1.0
0.9864
DeepSeek V3.2
1.0
0.9864
Qwen3-30B-A3B-Thinking-2507
0.0
0.0
Qwen2.5-14B-Instruct-1M
0.0
0.0
Gemini 2.5 Flash
0
0
Gemini 3.1
1.0
0.9864
Appendix
Table 8: Retrosynthesis planning performance on the 400-problem test set. Solve Rate indicates the fraction of problems for which the planner successfully generated a valid plan. Path Accuracy measures whether the predicted reaction sequence matches the ground-truth pathway, allowing both forward and reverse order matches to account for the retrosynthesis direction.
Model
Solve Rate
Path Accuracy
GPT-5.2
0
0
DeepSeek V3.2
0
0
Qwen3-30B-A3B-Thinking-2507
0
0
Qwen2.5-14B-Instruct-1M
0
0
Gemini 2.5 Flash
0
0
Gemini 3.1
0
0
Appendix
Table 9: Retrosynthesis planning performance using direct SMILES-based PDDL generation. In this setting, models are provided with the reactant and product SMILES strings and must directly generate both the domain.pddl and problem.pddl required for planning. This experiment evaluates a more direct generation pipeline compared to the reaction-template-based setting, serving as a comparison of end-to-end reasoning capability.
Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of feasible reactions. The vast combinatorial search space makes this task challenging even for expert chemists. Traditional methods combine tree search with offline-trained value networks that score candidates in isolation, without reasoning about complete multi-step routes. Recent work leverages Large Language Models (LLMs) for this task, but relies on simple interfaces that limit exploration of the full search space. We introduce RetroAgent, an LLM agent that bridges symbolic search and neural reasoning through a harness with structured memory. Through memory and chemistry tools, the agent observes the full search state, including explored routes, available alternatives, and properties of intermediates, enabling informed decisions grounded in both global progress and domain knowledge. Experiments on in-distribution and out-of-distribution benchmarks demonstrate that RetroAgent delivers strong performance and generalization.
Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule. Although molecular large language models (LLMs) have recently shown promising results, most existing methods either generate reactants directly or provide only generic product-level analysis, without explicitly reasoning about bond-disconnection strategies that justify specific reactant choices. This paper proposes RetroReasoner, a retrosynthetic reasoning model that captures chemists' strategic disconnection-based thinking. RetroReasoner is trained with supervised fine-tuning and reinforcement learning. For supervised fine-tuning, SyntheticRetro generates structured disconnection rationales paired with reactant predictions. For reinforcement learning, a round-trip reward evaluates predicted reactants by passing them through a forward synthesis model and rewarding predictions that reconstruct the original product. RetroReasoner can also be applied to multi-step retrosynthetic planning by incorporating it into a parallelized Monte Carlo tree search framework, reducing search time while increasing the number and diversity of valid synthetic pathways. Experimental results show that RetroReasoner outperforms prior baselines, including not only molecular LLMs but also retrosynthesis-specific expert models, and generates a broader range of feasible reactant proposals, especially for challenging reaction instances. The code is available at https://github.com/KU-AGI/RetroReasoner.
Hanbum Ko, Chanhui Lee, Ye Rin Kim +4
Department of Artificial Intelligence, Korea University · Materials Intelligence Lab, LG AI Research · Department of Statistics, Korea University
Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Recent progress has produced both specialized deep-learning retrosynthesis systems and general-purpose large language models, but objective comparison remains difficult due to the lack of flexible, chemically interpretable benchmarking protocols. In the current study, we are introducing the URSA (Utilitarian RetroSynthesis Assessment) evaluation framework that provides the opportunity to benchmark the synthetic routes not only from a formal perspective, such as convergence to commercially available starting materials, but also from a chemical plausibility perspective, mimicking the way expert chemists evaluate the reactions and routes. The study covers a comprehensive evaluation of both conventional end-to-end retrosynthesis solutions and LLMs for the synthesis planning task on a set of novel, diverse target molecules with undisclosed synthetic routes, which represent realistic tasks in the daily drug design routine. We find that while LLMs can support high-level strategic planning, they currently underperform specialized retrosynthesis models in reliably solving synthesis planning tasks.
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev +7
Insilico Medicine AI Limited, Masdar City, Abu Dhabi, UAE · Independent researcher · Insilico Medicine Canada Inc., Montreal, Quebec, Canada +1