While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL generation, we achieve high success rates where end-to-end approaches fail. This provides evidence that a primary bottleneck lies in representation alignment rather than raw model capacity. Our structural analysis demonstrates that intermediate representations are essential in retrosynthesis planning, highlighting the importance of representation-centric design in future systems.
Figures & tables
Model
Molecule Mapping
Reaction Mapping
PDDL Grounding
Molecule ID
SMILES
Coverage
Domain
Problem
GPT-5.2
1.0000
0.9932
1.0000
0.9303
1.0000
DeepSeek V3.2
0.9992
0.7870
1.0000
0.9300
1.0000
Gemini 3.1
0.6000
0.5964
1.0000
0.8402
1.0000
Gemini 2.5 Flash
0.6000
0.5933
0.9996
0.8351
1.0000
Qwen3-30B-Thinking
0.2976
0.0285
0.7000
0.4670
0.9128
Table 1: Performance comparison across staged subtasks in retrosynthesis planning (averaged). Molecule Mapping includes Molecule ID (identifier consistency) and SMILES (exact-match string accuracy). Reaction Mapping (Coverage) measures correctness of reactant–product structure. PDDL Grounding includes Domain (action generation) and Problem (initial and goal specification).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
100
200
300
400
500
600
700
800
900
1000
Qwen3-30B-A3B-Thinking-2507
0.9913 (0.1652)
0.9954 (0.0868)
–
–
0.9896 (0.0329)
–
–
–
–
–
Qwen2.5-14B-Instruct-1M
0.9913 (0.9304)
–
–
–
–
–
–
–
–
–
DeepSeek V3.2
1.0 (1.0)
0.9954 (0.6256)
1.0 (0.9907)
1.0 (0.9928)
1.0 (0.9931)
0.9987 (0.1797)
1.0 (0.9800)
1.0 (0.9839)
0.9982 (0.1238)
1.0 (1.0)
GPT-5.2
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9938)
1.0 (0.9904)
1.0 (0.9913)
1.0 (0.9935)
1.0 (0.9944)
1.0 (0.9930)
1.0 (0.9928)
1.0 (0.9921)
Gemini 2.5 Flash
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9907)
1.0 (0.9876)
1.0 (0.9844)
–
1.0 (0.9789)
–
–
–
Gemini 3.1
1.0 (1.0)
1.0 (0.9909)
1.0 (0.9907)
1.0 (0.9928)
–
1.0 (0.9961)
1.0 (0.9933)
–
–
–
Appendix
Table 2: Task 1 (Molecule Mapping) performance across varying dataset sizes. Each cell reports identifier exact-match rate and SMILES exact-match rate (in parentheses). “–” indicates invalid outputs.
Model
100
200
300
400
500
600
700
800
900
1000
Qwen3-30B-A3B-Thinking-2507
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
–
–
–
Qwen2.5-14B-Instruct-1M
1.0000
1.0000
–
–
–
–
–
–
–
–
DeepSeek V3.2
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
GPT-5.2
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
Gemini 2.5 Flash
1.0000
1.0000
1.0000
0.9975
1.0000
1.0000
1.0000
0.9988
1.0000
1.0000
Gemini 3.1
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
Appendix
Table 3: Task 2 (Reaction ID Mapping) integrity across varying dataset sizes. Each cell reports the exact mapping rate (EMR), defined as the proportion of reactions for which RXN_ID , Reactants_IDs , and Products_IDs exactly match the input. “–” indicates missing or invalid outputs (e.g., empty or unparsable model responses).
Model
100
200
300
400
GPT-5.2
1.0
1.0
1.0
1.0
DeepSeek V3.2
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
0.0
0.1176
0.2045
0.1531
Qwen2.5-14B-Instruct-1M
1.0
1.0
1.0
1.0
Gemini 2.5 Flash
1.0
1.0
1.0
1.0
Gemini 3.1
1.0
1.0
1.0
1.0
Appendix
Table 4: Syntactic validity rate of generated PDDL problems across different dataset sizes. A generation is considered valid if it satisfies basic structural constraints, including the correct problem definition, domain declaration, start molecule fact, goal predicates, and balanced parentheses.
Model
Metric
100
200
300
400
500
600
700
800
900
1000
GPT-5.2
Coverage
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Validity
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
DeepSeek V3.2
Coverage
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Validity
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
Coverage
1.0
1.0
0.9800
0.7960
0.5940
0
0
0
0.3380
0
Validity
1.0
1.0
0
0
0
0
0
0
0
0
Appendix
Table 5: Syntactic validity and action completion rate of generated PDDL domains across different dataset sizes (100–1000 reactions). Validity indicates whether the generated domain satisfies basic syntactic constraints such as correct domain structure, valid action blocks, and balanced parentheses. Completion measures the fraction of expected actions successfully generated (e.g., number of generated actions divided by the expected number of reactions).
Model
Metric
100
200
300
400
500
600
700
800
900
1000
GPT-5.2
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0.8522
0.8350
DeepSeek V3.2
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0.8489
0.8350
Qwen3-30B-A3B-Thinking-2507
Grounding
0.9900
0.9950
0.9733
0.7884
0.5880
0
0
0.3356
0
0
Qwen2.5-14B-Instruct-1M
Grounding
0.9900
0.7900
0.5367
0.3501
0.3240
0.2683
0.2300
0.2012
0.1622
0.1580
Gemini 2.5 Flash
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.8883
0.8857
0.8612
0
0.8350
Gemini 3.1
Grounding
0.9900
0.9950
0.9933
0.9950
0.9420
0.0767
0.8857
0.8612
0.8522
0.8350
Appendix
Table 6: Action grounding performance of generated PDDL domains across different dataset sizes (100–1000 reactions). Grounding measures whether each generated action correctly includes the corresponding reactant and product identifiers for the given reaction. This metric evaluates semantic correctness beyond syntactic validity and action coverage.
Model
Metric
100
200
300
400
GPT-5.2
Problem Coverage
1.0
1.0
1.0
1.0
DeepSeek V3.2
Problem Coverage
1.0
1.0
1.0
1.0
Qwen3-30B-A3B-Thinking-2507
Problem Coverage
1.0
0.8897
0.8636
0.8980
Qwen2.5-14B-Instruct-1M
Problem Coverage
0.1304
0.3309
0.5909
0.5442
Gemini 2.5 Flash
Problem Coverage
1.0
1.0
1.0
1.0
Gemini 3.1
Problem Coverage
1.0
1.0
1.0
1.0
Appendix
Table 7: Problem coverage of generated PDDL problem files across dataset sizes (100–400 reactions). Problem coverage measures the fraction of ground-truth problem instances that are correctly generated, where each problem file is aligned with the input molecule (Mol_Name), and includes both the corresponding final product (Final_Product_ID) and all required goal candidate molecules.
Model
Solve Rate
Path Accuracy
GPT-5.2
1.0
0.9864
DeepSeek V3.2
1.0
0.9864
Qwen3-30B-A3B-Thinking-2507
0.0
0.0
Qwen2.5-14B-Instruct-1M
0.0
0.0
Gemini 2.5 Flash
0
0
Gemini 3.1
1.0
0.9864
Appendix
Table 8: Retrosynthesis planning performance on the 400-problem test set. Solve Rate indicates the fraction of problems for which the planner successfully generated a valid plan. Path Accuracy measures whether the predicted reaction sequence matches the ground-truth pathway, allowing both forward and reverse order matches to account for the retrosynthesis direction.
Model
Solve Rate
Path Accuracy
GPT-5.2
0
0
DeepSeek V3.2
0
0
Qwen3-30B-A3B-Thinking-2507
0
0
Qwen2.5-14B-Instruct-1M
0
0
Gemini 2.5 Flash
0
0
Gemini 3.1
0
0
Appendix
Table 9: Retrosynthesis planning performance using direct SMILES-based PDDL generation. In this setting, models are provided with the reactant and product SMILES strings and must directly generate both the domain.pddl and problem.pddl required for planning. This experiment evaluates a more direct generation pipeline compared to the reaction-template-based setting, serving as a comparison of end-to-end reasoning capability.