Organizations: The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · Institute of Computing Technology, Chinese Academy of Sciences.
Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection and correction before use, making large-scale data curation costly and difficult to scale. We therefore study Selective Structure Recognition (SSR), a post-recognition setting that automatically produces reliable structured outputs while rejecting unresolved cases. Selection-only approaches can improve reliability by rejection, but cannot create additional correct outputs beyond those produced by the base recognizer. We propose ViCoR, a repair-before-rejection framework for iterative VerIfiCatiOn and Revision. Its key idea is to make observation-prediction correspondence explicit: coordinate-preserving rendering establishes spatial correspondence between the source image and predicted structure, while index anchoring maps localized visual discrepancies to executable graph edits without full-structure regeneration. A shared VLM is progressively trained from verification to revision. On two real-world OCSR benchmarks, ViCoR improves overall accuracy from 73.53% to 88.26% and from 61.83% to 84.32%, while achieving over 97% accepted accuracy at 85--89% coverage. The resulting molecular data further improve reaction-extraction F1 by 15.5 points and literature-sourced reaction prediction accuracy by 7.7 and 5.8 points, demonstrating the value of automated reliability control for scientific data curation and downstream chemical learning.
Figures & tables
Figure 1: From unreliable OCSR outputs to reliable scientific data with ViCoR. Recognition errors can corrupt chemical data. ViCoR uses spatially aligned verification and executable revision to recover correct structures, improving both reliability–coverage and downstream chemical learning.
Figure 2: Overview of ViCoR. Starting from an OCSR-predicted initial graph, ViCoR iteratively verifies and revises the prediction against the source image. Coordinate-preserving rendering establishes spatial correspondence for verification, while atom indexing maps localized visual discrepancies to executable graph edits for revision. Each revised graph is re-rendered and re-verified until it is accepted or reaches the maximum iteration T.
Category
Method
OCIE
JIE
OA ↑
AA ↑
Cov ↑
OA ↑
AA ↑
Cov ↑
VLM (e2e)
Qwen2.5-VL-3B ( Bai et al., 2025 )
0.54
0.54
100.0
0.90
0.90
100.0
GPT-4o ( OpenAI, 2026 )
13.91
13.91
100.0
23.58
23.58
100.0
GPT-5.6-sol ( OpenAI, 2026 )
49.04
49.04
100.0
51.95
51.95
100.0
Expert
OSRA ( Filippov et al., 2009 )
36.17
36.17
100.0
19.88
19.88
100.0
DECIMER ( Rajan et al., 2020 )
35.07
35.07
100.0
16.42
16.42
100.0
Table 1: Selective Structure Recognition on real-world OCSR benchmarks. OA/AA denote exact-match accuracy over all/accepted predictions, and Cov denotes acceptance coverage. Agreement filter denotes that only consistent results are accepted. VLM SMILES revision used GPT-5.6-sol.
Category
Method
Synthetic
Realistic
Indigo
ChemDraw
CLEF
UOB
JPO
USPTO
Staker
ACS
Rule-based
MolVec ( Peryea et al., 2019 )
95.4
87.9
82.8
80.6
67.8
88.4
0.8
47.4
OSRA ( Filippov et al., 2009 )
95.0
87.3
84.6
78.5
55.3
87.4
0.0
55.3
End-to-end
DECIMER ( Rajan et al., 2020 )
69.6
86.1
62.7
88.2
55.2
41.1
40.8
46.5
MolParser ( Fang et al., 2025 )
—
—
91.0
91.6
75.6
93.0
—
—
MolSight ( Zhang et al., 2026 )
—
—
85.5
87.4
57.6
92.0
—
—
Table 2: Exact-match accuracy (OA) on eight standard OCSR benchmarks. Δ denotes the gain of ViCoR over its base recognizer, MolScribe.
Table 5
Interface
Design
OCIE
JIE
Spat. Align
SBS
Index
OA ↑
AA ↑
Cov ↑
OA ↑
AA ↑
Cov ↑
Base recognizer
–
–
–
73.53
73.53
100.0
61.83
61.83
100.0
Source-prediction interfaces
Image & SMILES
✗
✗
✗
74.26
82.70
68.8
62.91
77.80
59.6
Image & graph
✗
✗
✓
76.48
86.90
73.4
65.74
81.90
65.1
Rendered interfaces
Table 5: Ablation of the spatially aligned visual comparison. Spat. Align denotes coordinate-preserving rendering, SBS side-by-side presentation, and Index explicit atom indexing.
Figure 4: Qualitative results of ViCoR. (a) Representative Revision Examples. (b) ViCoR improves recognition through iterative verification and revision. (c) Data efficiency on corner cases.
Table 8
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Examples of human filtered and retained corner case training samples.
Figure 6: An example of the input and output of the ViCoR for revision.
Task
Prompt
Verification
Do the left and right molecular images represent the same molecule? Ignore the drawing style and other noises. Answer only ‘Yes.’ or ‘No.’ .
Revision
Objective: Compare the left with the right molecular image. Predict all revision actions to reconstruct the exact graph shown in the reference image. Instructions: 1. Check for missing, redundant, or incorrect atoms and bonds. 2. Match indices to specific visual regions strictly using spatial proximity. 3. Formulate the corrections as a sequence of discrete graph edit operations. Output Format: Return ONLY a JSON list of dictionaries containing the required edit instructions. You must use the following templates: • {"Add atom": [<index>, "<symbol>", ["<x>", "<y>"]]} • {"Del atom": [<index>, "<symbol>"]} • {"Rev atom": [<index>, "<target_symbol>"]} • {"Add bond": [[<src>, <tgt>], "<bond_type>"]} • {"Del bond": [[<src>, <tgt>], "<bond_type>"]} • {"Rev bond": [[<src>, <tgt>], "<target_bond_type>"]} Example: [{"Add atom": [3, "C", ["<0.314>", "<0.576>"]]}, {"Add bond": [[3, 4], "solid wedge"]}]
Appendix
Table 8: Prompts used for verification and revision in ViCoR
Dataset
Source domain
Primary focus
Images
Median size
Color (%)
IP5-M
Patents
Markush / R-groups
878
1024×1024
0.2
USPTO-10K-Abb
Patents
Abbreviated superatoms
10,000
648×394
0.0
OCIE
Journal figures (2022–2023)
Mixed OCSR factors
1,141
220×155
72.0
JIE
Journal figures (2022–2024)
Mixed OCSR factors
1,779
233×158
19.6
Appendix
Table 9: Comparison of OCIE and JIE with existing real-world OCSR benchmarks. OCIE and JIE complement patent-domain benchmarks with recent journal figures containing smaller image crops, frequent color usage, and naturally co-occurring recognition challenges.
Dataset
Type
Total Images
Abbreviations
Indigo ( Qian et al., 2023 )
Synthetic
5,719
×
ChemDraw ( Qian et al., 2023 )
Synthetic
5,719
×
CLEF ( Piroi et al., 2010 )
Real
992
✓
UOB ( Sadawi et al., 2012 )
Real
5,740
✓
JPO ( Fujiyoshi et al., 2011 )
Real
450
×
USPTO ( Filippov et al., 2009 )
Real
5,719
✓
Appendix
Table 10: Summary of the test datasets
OCIE
JIE
Method
Throughput ↑
Latency ↓
Rel. latency
OA ↑
Δ
OA ↑
Δ
(img/s)
(s/img)
vs. MolScribe
vs. MolScribe
vs. MolScribe
OSRA
6.64
0.151
0.24 ×
36.17
−37.36
19.88
−41.95
DECIMER
0.28
3.571
5.75 ×
35.07
−38.46
16.42
−45.41
MolGrapher
0.42
2.381
3.83 ×
41.20
−32.33
25.40
−36.43
MolNexTR
1.42
0.704
1.13 ×
73.24
−0.29
61.62
−0.21
Appendix
Table 11: End-to-end accuracy and computational cost on a single NVIDIA RTX 3090. Relative latency is normalized by MolScribe.
OCIE
JIE
Method
Throughput ↑
Latency ↓
Rel. latency
OA ↑
Δ
OA ↑
Δ
(img/s)
(s/img)
vs. base
vs. base
vs. base
MolScribe base
1.61
0.621
1.00 ×
73.53
–
61.83
–
ViCoR ( T=1 )
1.33
0.752
1.21 ×
80.82
+7.29
72.45
+10.62
ViCoR ( T=2 )
1.21
0.826
1.33 ×
86.45
+12.92
81.36
+19.53
ViCoR ( T=3 )
1.18
0.847
1.36 ×
88.26
+14.73
84.32
+22.49
Appendix
Table 12: Accuracy–cost trade-off under different ViCoR iteration budgets. Relative latency is normalized by the MolScribe base recognizer.
Training comparison
OCIE OA ↑
JIE OA ↑
RDKit → RDKit
85.42
80.76
Indigo → RDKit
88.26
84.32
Appendix
Table 13: Effect of rendering diversity during verification training (%).
VLM
OCIE OA ↑
JIE OA ↑
ViCoR(GPT-4o) + MolScribe
76.60
73.97
ViCoR(GPT-5.6-sol) + MolScribe
77.83
75.77
ViCoR(Trained Qwen2.5-VL-3B) + MolScribe
88.26
84.32
ViCoR(Trained Qwen2.5-VL-7B) + MolScribe
88.48
84.53
Appendix
Table 14: Effect of the VLM backbone and model scale on strict recognition accuracy (%).
Figure 7: Examples of comparison between baseline MolScribe and MolScribe + ViCoR.
Figure 8: Example of chemical information extraction (Left: input reaction scope image. Right: expected extracted 7 reactions from the image). This figure comes from ( Chen et al., 2025b ) .
Method
OA ↑
AA ↑
Cov ↑
Faster R-CNN base
41.60
41.60
100.00
YOLO11m base
50.40
50.40
100.00
Faster R-CNN + ViCoR
62.10
95.50
57.80
YOLO11m + ViCoR
68.80
96.30
64.80
Appendix
Table 15: Transfer of ViCoR to hand-drawn BPMN process-graph recognition (%). Results are reported on the official writer-disjoint split of hdBPMN v1.0.0.