Organizations: Westlake University · The University of Hong Kong · Shanghai Innovation Institute · Zhejiang University · Sichuan University · Shanghai Artificial Intelligence Laboratory · The Hong Kong University of Science and Technology (Guangzhou) · University of Illinois Urbana-Champaign
Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90% accuracy on Easy VQA, performance drops to56--66% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7--46.2% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20% for most models and below8% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
Figures & tables
Figure 1: Overview of R-GroundBench . A : R-group grounding requires aligning disconnected molecular, textual, and symbolic information. B : R-GroundBench evaluates this capability through complementary VQA and Generation tracks with controlled difficulty and modality settings. C : Results reveal a gap between candidate recognition and autonomous molecular instantiation, exposing shortcut reliance in current models.
Figure 2: Overview of R-GroundBench construction and task formulation. A : Dataset construction from real Markush structures with filtering, annotation, and instruction generation. B : Dual-track evaluation design, including VQA with controlled difficulty/modality settings and Generation without candidate options.
Mode
Input
Candidates
Tests
i2i
Markush image
mol. images
Visual grounding
i2s
Markush image
SMILES options
Image → symbol
s2i
E-SMILES
mol. images
Symbol → visual
s2s
E-SMILES
SMILES options
Symbolic editing
Table 2: The four input-output modality settings in R-GroundBench . Notation: i2i = image in, image candidates; i2s = image in, SMILES candidates; s2i = E-SMILES in, image candidates; s2s = E-SMILES in, SMILES candidates.
Model Type
Model
Modality
Basic
Advanced
ΔEH
Avg.
Easy
Med
Hard
Avg.
Easy
Med
Hard
Avg.
Basic
Adv
General VLMs
Qwen3-VL-8B ( Qwen Team 2025b )
i2i
25.5
24.9
25.1
25.2
21.3
20.4
21.1
20.9
0.4
0.2
23.0
i2s
93.6
74.8
55.7
74.7
55.2
54.7
55.7
55.2
37.9
-0.5
64.9
s2i
78.4
68.7
49.6
65.6
37.7
38.2
37.2
37.7
28.8
0.5
51.6
s2s
93.4
72.7
55.9
74.0
49.4
48.7
49.2
49.1
37.5
0.2
61.6
Qwen3-VL-32B ( Qwen Team 2025b )
i2i
98.1
90.2
55.5
81.3
44.6
44.2
47.0
45.3
42.6
-2.4
63.3
Table 3: Full VQA results on R-GroundBench (% accuracy). Columns: Basic/Advanced accuracy by difficulty (Easy/Med/Hard), their means, ΔEH (Easy minus Hard), and overall mean across six splits. Rows: each model × modality (i2i, i2s, s2i, s2s); LLMs on s2s only. bold : best; underline : second best among models within each modality and model-type block. Chemical-domain models are not marked because model group contains fewer than three models.
Easy
Hard
Model
EM
Scaff.
Tani.
EM
Scaff.
Tani.
VLMs — s2s
Qwen3-VL-8B
4.2
49.2
44.0
3.0
52.2
45.6
Qwen3-VL-32B
12.8
38.0
57.7
10.0
41.4
52.9
Claude Sonnet 4.6
24.0
31.2
88.6
16.2
23.0
87.5
GPT-4o
19.4
53.0
59.9
10.4
50.8
54.0
Table 4: Generation track results (%). EM = Exact Match, Scaff. = Scaffold Match, Tani. = Tanimoto Similarity. Chemical-domain models are not marked.
Figure 3: Four representative failure cases of GPT-4.1 on R-GroundBench .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Filter step
Remaining
Removed
Raw MolParser-7M sft_real
91,166
—
+ Has <a> tag (pre-filter)
46,727
44,439
+ Standard R-group token only
19,282
27,445
+ No non-standard tokens
14,476
4,806
+ R-group count ≤10
14,394
82
+ RDKit-parseable backbone
14,394
0
Appendix
Table 5: Sequential filtering statistics. The pre-filter (row 2) is applied during dataset download and retains only samples containing at least one <a> annotation tag; the remaining four steps correspond to the criteria listed above.
Table 6: Seven similarity clusters of the 22-substituent pool. Within-group members share a common scaffold but differ in ring size, substitution pattern, or heteroatom position, providing challenging fine-grained alternatives for Hard splits.
Style
Split
Example
Direct
Easy/Med
Replace R 1 with 4-methylphenyl.
Conversational
Easy/Med
Can you change R 1 to 4-methylphenyl?
Result-oriented
Easy/Med
I need the R 1 position to be 4-methylphenyl.
Relational
Hard
Among the four candidates, select the one whose R 1 substituent contains the greatest number of halogen atoms.
Chemically descriptive
Hard
Replace R 1 with a para-halogenated phenyl ring, specifically the one whose halogen is fluorine rather than chlorine.
Appendix
Table 7: Five instruction templates used in R-GroundBench . The upper three are direct-name styles deployed in Easy and Medium splits. The lower two are Hard templates that require additional grounding beyond explicit substituent names: Relational identifies the target through relations among candidates, while Chemically descriptive specifies the target through structural descriptions rather than direct naming.
Figure 4: Physicochemical property distributions of edited molecules ( n=55,982 ). Green histograms with log-normal fits (red).
Statistic
Value
Source subset
MolParser-7M sft_real
Raw patent/lit. mols
91,166
Markush-containing
46,727
Clean samples used
14,394
Unique Markush scaffolds
12,962
Unique editing operations
55,982
Appendix
Table 8: R-GroundBench dataset statistics.
Figure 5: Scaffold diversity. Left : cumulative coverage curve (23,822 unique scaffolds). Right : frequency group distribution (82.1% singletons).
Figure 6: ChEMBL SAR data quality. Left : target class distribution (308 clusters, 13 types). Right : pIC 50 distribution ( μ=6.92 , σ=1.19 , n=2,387 ).
Figure 7: Cross-track degradation from VQA to Generation under the s2s modality. For each model, blue bars show Basic VQA accuracy, green bars show Advanced VQA accuracy, and purple bars show Generation Exact Match (EM). The black curve connects all scores to highlight the diagnostic performance cliff: models that perform well on VQA, especially Basic Easy/Medium splits, collapse when required to generate the edited molecule directly. This pattern indicates that multiple-choice success often reflects shortcut exploitation rather than robust atomic-level R-group grounding.
NOTA is correct
Model
Struct. err
Non-exist R
NOTA as distr.
VLMs (img-smi)
Qwen3-VL-8B
61.0
75.8
73.3
Qwen3-VL-32B
88.6
97.0
72.1
Claude Sonnet 4.6
92.9
98.0
80.2
GPT-4o
90.0
94.9
72.1
Appendix
Table 9: NOTA robustness (%, img-smi for VLMs; smi-smi for LLMs). Red bold : best; underline : second best.
Model
Physchem
Bioactivity
Indication
VLMs (img-smi, Hard)
Qwen3-VL-8B
85.5
26.3
83.7
Qwen3-VL-32B
89.3
26.7
88.4
Claude Sonnet 4.6
66.7
23.9
82.6
GPT-4o
86.2
29.5
82.1
GPT-4.1
84.9
26.0
88.9
Appendix
Table 10: Advanced property reasoning accuracy (%, Hard split) by property type. Red bold : best; underline : second best.
Gap Δsim
Easy
Medium
Hard
<−0.05
—
91.9
60.3
[−0.05,0.00)
94.3
93.3
63.4
[−0.00,0.05)
95.7
93.6
61.2
[−0.05,0.10)
98.4
98.4
62.1
[−0.10,0.20)
98.0
—
—
[−0.20,0.50)
98.6
—
—
Appendix
Table 11: Accuracy (%) by similarity gap Δsim (Eq. 1 ), averaged over 8 VLMs under the img-smi setting. NOTA questions are excluded. The lower rows correspond to Easy-Basic samples with larger scaffold-similarity advantages, which are absent in Medium- and Hard-Basic due to controlled distractor construction. “—”: n<10 .
Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing machine learning model training sets, they must be transformed into line notations. The two common forms of this task are translating an image of a single molecule (optical chemical structure recognition - OCSR) and translating a Markush structure that represents a family of molecules. While prior work in the former case is quite mature, Markush structure parsing remains a challenging task. In this work, we treat both tasks as an image-to-text translation problem. We then propose OCSRGlyph, a state-of-the-art OCSR model, improving performance over prior methods by carefully considering stereochemistry. For the Markush task, we introduce MarkushGlyph, a vision-language model that reads the entire Markush structure as an image. This contrasts with prior systems, which often use multiple stages to separately process visual and text input content. Finally, we introduce a new metric for determining the accuracy of Markush structure translations, handling failure modes present in prior metrics.
Alex Andonian, Samuel G Rodriques, Andrew D White +1
Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.
Hongyu Guo, Hao Li, He Cao +2
Peking University, Shenzhen Graduate School · International Digital Economy Academy (IDEA)
Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.
Jiatong Li, Yuxuan Ren, Weida Wang +2
Blue Whale Lab, National University of Singapore · Fudan University · Hong Kong Polytechnic University