Organizations: Westlake University · The University of Hong Kong · Shanghai Innovation Institute · Zhejiang University · Sichuan University · Shanghai Artificial Intelligence Laboratory · The Hong Kong University of Science and Technology (Guangzhou) · University of Illinois Urbana-Champaign
Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90% accuracy on Easy VQA, performance drops to56--66% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7--46.2% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20% for most models and below8% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
Figures & tables
Figure 1: Overview of R-GroundBench . A : R-group grounding requires aligning disconnected molecular, textual, and symbolic information. B : R-GroundBench evaluates this capability through complementary VQA and Generation tracks with controlled difficulty and modality settings. C : Results reveal a gap between candidate recognition and autonomous molecular instantiation, exposing shortcut reliance in current models.
Figure 2: Overview of R-GroundBench construction and task formulation. A : Dataset construction from real Markush structures with filtering, annotation, and instruction generation. B : Dual-track evaluation design, including VQA with controlled difficulty/modality settings and Generation without candidate options.
Mode
Input
Candidates
Tests
i2i
Markush image
mol. images
Visual grounding
i2s
Markush image
SMILES options
Image → symbol
s2i
E-SMILES
mol. images
Symbol → visual
s2s
E-SMILES
SMILES options
Symbolic editing
Table 2: The four input-output modality settings in R-GroundBench . Notation: i2i = image in, image candidates; i2s = image in, SMILES candidates; s2i = E-SMILES in, image candidates; s2s = E-SMILES in, SMILES candidates.
Model Type
Model
Modality
Basic
Advanced
ΔEH
Avg.
Easy
Med
Hard
Avg.
Easy
Med
Hard
Avg.
Basic
Adv
General VLMs
Qwen3-VL-8B ( Qwen Team 2025b )
i2i
25.5
24.9
25.1
25.2
21.3
20.4
21.1
20.9
0.4
0.2
23.0
i2s
93.6
74.8
55.7
74.7
55.2
54.7
55.7
55.2
37.9
-0.5
64.9
s2i
78.4
68.7
49.6
65.6
37.7
38.2
37.2
37.7
28.8
0.5
51.6
s2s
93.4
72.7
55.9
74.0
49.4
48.7
49.2
49.1
37.5
0.2
61.6
Qwen3-VL-32B ( Qwen Team 2025b )
i2i
98.1
90.2
55.5
81.3
44.6
44.2
47.0
45.3
42.6
-2.4
63.3
Table 3: Full VQA results on R-GroundBench (% accuracy). Columns: Basic/Advanced accuracy by difficulty (Easy/Med/Hard), their means, ΔEH (Easy minus Hard), and overall mean across six splits. Rows: each model × modality (i2i, i2s, s2i, s2s); LLMs on s2s only. bold : best; underline : second best among models within each modality and model-type block. Chemical-domain models are not marked because model group contains fewer than three models.
Easy
Hard
Model
EM
Scaff.
Tani.
EM
Scaff.
Tani.
VLMs — s2s
Qwen3-VL-8B
4.2
49.2
44.0
3.0
52.2
45.6
Qwen3-VL-32B
12.8
38.0
57.7
10.0
41.4
52.9
Claude Sonnet 4.6
24.0
31.2
88.6
16.2
23.0
87.5
GPT-4o
19.4
53.0
59.9
10.4
50.8
54.0
Table 4: Generation track results (%). EM = Exact Match, Scaff. = Scaffold Match, Tani. = Tanimoto Similarity. Chemical-domain models are not marked.
Figure 3: Four representative failure cases of GPT-4.1 on R-GroundBench .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Filter step
Remaining
Removed
Raw MolParser-7M sft_real
91,166
—
+ Has <a> tag (pre-filter)
46,727
44,439
+ Standard R-group token only
19,282
27,445
+ No non-standard tokens
14,476
4,806
+ R-group count ≤10
14,394
82
+ RDKit-parseable backbone
14,394
0
Appendix
Table 5: Sequential filtering statistics. The pre-filter (row 2) is applied during dataset download and retains only samples containing at least one <a> annotation tag; the remaining four steps correspond to the criteria listed above.
Table 6: Seven similarity clusters of the 22-substituent pool. Within-group members share a common scaffold but differ in ring size, substitution pattern, or heteroatom position, providing challenging fine-grained alternatives for Hard splits.
Style
Split
Example
Direct
Easy/Med
Replace R 1 with 4-methylphenyl.
Conversational
Easy/Med
Can you change R 1 to 4-methylphenyl?
Result-oriented
Easy/Med
I need the R 1 position to be 4-methylphenyl.
Relational
Hard
Among the four candidates, select the one whose R 1 substituent contains the greatest number of halogen atoms.
Chemically descriptive
Hard
Replace R 1 with a para-halogenated phenyl ring, specifically the one whose halogen is fluorine rather than chlorine.
Appendix
Table 7: Five instruction templates used in R-GroundBench . The upper three are direct-name styles deployed in Easy and Medium splits. The lower two are Hard templates that require additional grounding beyond explicit substituent names: Relational identifies the target through relations among candidates, while Chemically descriptive specifies the target through structural descriptions rather than direct naming.
Figure 4: Physicochemical property distributions of edited molecules ( n=55,982 ). Green histograms with log-normal fits (red).
Statistic
Value
Source subset
MolParser-7M sft_real
Raw patent/lit. mols
91,166
Markush-containing
46,727
Clean samples used
14,394
Unique Markush scaffolds
12,962
Unique editing operations
55,982
Appendix
Table 8: R-GroundBench dataset statistics.
Figure 5: Scaffold diversity. Left : cumulative coverage curve (23,822 unique scaffolds). Right : frequency group distribution (82.1% singletons).
Figure 6: ChEMBL SAR data quality. Left : target class distribution (308 clusters, 13 types). Right : pIC 50 distribution ( μ=6.92 , σ=1.19 , n=2,387 ).
Figure 7: Cross-track degradation from VQA to Generation under the s2s modality. For each model, blue bars show Basic VQA accuracy, green bars show Advanced VQA accuracy, and purple bars show Generation Exact Match (EM). The black curve connects all scores to highlight the diagnostic performance cliff: models that perform well on VQA, especially Basic Easy/Medium splits, collapse when required to generate the edited molecule directly. This pattern indicates that multiple-choice success often reflects shortcut exploitation rather than robust atomic-level R-group grounding.
NOTA is correct
Model
Struct. err
Non-exist R
NOTA as distr.
VLMs (img-smi)
Qwen3-VL-8B
61.0
75.8
73.3
Qwen3-VL-32B
88.6
97.0
72.1
Claude Sonnet 4.6
92.9
98.0
80.2
GPT-4o
90.0
94.9
72.1
Appendix
Table 9: NOTA robustness (%, img-smi for VLMs; smi-smi for LLMs). Red bold : best; underline : second best.
Model
Physchem
Bioactivity
Indication
VLMs (img-smi, Hard)
Qwen3-VL-8B
85.5
26.3
83.7
Qwen3-VL-32B
89.3
26.7
88.4
Claude Sonnet 4.6
66.7
23.9
82.6
GPT-4o
86.2
29.5
82.1
GPT-4.1
84.9
26.0
88.9
Appendix
Table 10: Advanced property reasoning accuracy (%, Hard split) by property type. Red bold : best; underline : second best.
Gap Δsim
Easy
Medium
Hard
<−0.05
—
91.9
60.3
[−0.05,0.00)
94.3
93.3
63.4
[−0.00,0.05)
95.7
93.6
61.2
[−0.05,0.10)
98.4
98.4
62.1
[−0.10,0.20)
98.0
—
—
[−0.20,0.50)
98.6
—
—
Appendix
Table 11: Accuracy (%) by similarity gap Δsim (Eq. 1 ), averaged over 8 VLMs under the img-smi setting. NOTA questions are excluded. The lower rows correspond to Easy-Basic samples with larger scaffold-similarity advantages, which are absent in Medium- and Hard-Basic due to controlled distractor construction. “—”: n<10 .