ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision
Organizations: The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · Institute of Computing Technology, Chinese Academy of Sciences.
Abstract
Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection and correction before use, making large-scale data curation costly and difficult to scale. We therefore study Selective Structure Recognition (SSR), a post-recognition setting that automatically produces reliable structured outputs while rejecting unresolved cases. Selection-only approaches can improve reliability by rejection, but cannot create additional correct outputs beyond those produced by the base recognizer. We propose ViCoR, a repair-before-rejection framework for iterative VerIfiCatiOn and Revision. Its key idea is to make observation-prediction correspondence explicit: coordinate-preserving rendering establishes spatial correspondence between the source image and predicted structure, while index anchoring maps localized visual discrepancies to executable graph edits without full-structure regeneration. A shared VLM is progressively trained from verification to revision. On two real-world OCSR benchmarks, ViCoR improves overall accuracy from 73.53% to 88.26% and from 61.83% to 84.32%, while achieving over 97% accepted accuracy at 85--89% coverage. The resulting molecular data further improve reaction-extraction F1 by 15.5 points and literature-sourced reaction prediction accuracy by 7.7 and 5.8 points, demonstrating the value of automated reliability control for scientific data curation and downstream chemical learning.
Figures & tables
| Category | Method | OCIE | JIE | ||||
| OA | AA | Cov | OA | AA | Cov | ||
| VLM (e2e) | Qwen2.5-VL-3B ( Bai et al., 2025 ) | ||||||
| GPT-4o ( OpenAI, 2026 ) | |||||||
| GPT-5.6-sol ( OpenAI, 2026 ) | |||||||
| Expert | OSRA ( Filippov et al., 2009 ) | ||||||
| DECIMER ( Rajan et al., 2020 ) | |||||||
| Category | Method | Synthetic | Realistic | ||||||
| Indigo | ChemDraw | CLEF | UOB | JPO | USPTO | Staker | ACS | ||
| Rule-based | MolVec ( Peryea et al., 2019 ) | ||||||||
| OSRA ( Filippov et al., 2009 ) | |||||||||
| End-to-end | DECIMER ( Rajan et al., 2020 ) | ||||||||
| MolParser ( Fang et al., 2025 ) | — | — | — | — | |||||
| MolSight ( Zhang et al., 2026 ) | — | — | — | — | |||||
| Interface | Design | OCIE | JIE | ||||||
| Spat. Align | SBS | Index | OA | AA | Cov | OA | AA | Cov | |
| Base recognizer | – | – | – | 73.53 | 73.53 | 100.0 | 61.83 | 61.83 | 100.0 |
| Source-prediction interfaces | |||||||||
| Image & SMILES | ✗ | ✗ | ✗ | 74.26 | 82.70 | 68.8 | 62.91 | 77.80 | 59.6 |
| Image & graph | ✗ | ✗ | ✓ | 76.48 | 86.90 | 73.4 | 65.74 | 81.90 | 65.1 |
| Rendered interfaces | |||||||||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Prompt |
|---|---|
| Verification | Do the left and right molecular images represent the same molecule? Ignore the drawing style and other noises. Answer only ‘Yes.’ or ‘No.’ . |
| Revision | Objective: Compare the left with the right molecular image. Predict all revision actions to reconstruct the exact graph shown in the reference image. Instructions: 1. Check for missing, redundant, or incorrect atoms and bonds. 2. Match indices to specific visual regions strictly using spatial proximity. 3. Formulate the corrections as a sequence of discrete graph edit operations. Output Format: Return ONLY a JSON list of dictionaries containing the required edit instructions. You must use the following templates: • {"Add atom": [<index>, "<symbol>", ["<x>", "<y>"]]} • {"Del atom": [<index>, "<symbol>"]} • {"Rev atom": [<index>, "<target_symbol>"]} • {"Add bond": [[<src>, <tgt>], "<bond_type>"]} • {"Del bond": [[<src>, <tgt>], "<bond_type>"]} • {"Rev bond": [[<src>, <tgt>], "<target_bond_type>"]} Example: [{"Add atom": [3, "C", ["<0.314>", "<0.576>"]]}, {"Add bond": [[3, 4], "solid wedge"]}] |
| Dataset | Source domain | Primary focus | Images | Median size | Color (%) |
|---|---|---|---|---|---|
| IP5-M | Patents | Markush / R-groups | 878 | 0.2 | |
| USPTO-10K-Abb | Patents | Abbreviated superatoms | 10,000 | 0.0 | |
| OCIE | Journal figures (2022–2023) | Mixed OCSR factors | 1,141 | 72.0 | |
| JIE | Journal figures (2022–2024) | Mixed OCSR factors | 1,779 | 19.6 |
| Dataset | Type | Total Images | Abbreviations |
| Indigo ( Qian et al., 2023 ) | Synthetic | 5,719 | |
| ChemDraw ( Qian et al., 2023 ) | Synthetic | 5,719 | |
| CLEF ( Piroi et al., 2010 ) | Real | 992 | |
| UOB ( Sadawi et al., 2012 ) | Real | 5,740 | |
| JPO ( Fujiyoshi et al., 2011 ) | Real | 450 | |
| USPTO ( Filippov et al., 2009 ) | Real | 5,719 |
| OCIE | JIE | ||||||
| Method | Throughput | Latency | Rel. latency | OA | OA | ||
| (img/s) | (s/img) | vs. MolScribe | vs. MolScribe | vs. MolScribe | |||
| OSRA | 6.64 | 0.151 | 0.24 | 36.17 | 19.88 | ||
| DECIMER | 0.28 | 3.571 | 5.75 | 35.07 | 16.42 | ||
| MolGrapher | 0.42 | 2.381 | 3.83 | 41.20 | 25.40 | ||
| MolNexTR | 1.42 | 0.704 | 1.13 | 73.24 | 61.62 | ||
| OCIE | JIE | ||||||
| Method | Throughput | Latency | Rel. latency | OA | OA | ||
| (img/s) | (s/img) | vs. base | vs. base | vs. base | |||
| MolScribe base | 1.61 | 0.621 | 1.00 | 73.53 | – | 61.83 | – |
| ViCoR ( ) | 1.33 | 0.752 | 1.21 | 80.82 | +7.29 | 72.45 | +10.62 |
| ViCoR ( ) | 1.21 | 0.826 | 1.33 | 86.45 | +12.92 | 81.36 | +19.53 |
| ViCoR ( ) | 1.18 | 0.847 | 1.36 | 88.26 | +14.73 | 84.32 | +22.49 |
| Training comparison | OCIE OA | JIE OA |
|---|---|---|
| RDKit RDKit | 85.42 | 80.76 |
| Indigo RDKit | 88.26 | 84.32 |
| VLM | OCIE OA | JIE OA |
|---|---|---|
| ViCoR(GPT-4o) + MolScribe | 76.60 | 73.97 |
| ViCoR(GPT-5.6-sol) + MolScribe | 77.83 | 75.77 |
| ViCoR(Trained Qwen2.5-VL-3B) + MolScribe | 88.26 | 84.32 |
| ViCoR(Trained Qwen2.5-VL-7B) + MolScribe | 88.48 | 84.53 |
| Method | OA | AA | Cov |
|---|---|---|---|
| Faster R-CNN base | 41.60 | 41.60 | 100.00 |
| YOLO11m base | 50.40 | 50.40 | 100.00 |
| Faster R-CNN + ViCoR | 62.10 | 95.50 | 57.80 |
| YOLO11m + ViCoR | 68.80 | 96.30 | 64.80 |
| Qwen2.5-VL-3B | MolScribe | |
| Image Encoder | ||
| Architecture | Vision Transformer | Swin Transformer (Swin-B) |
| Layers | 32 | 36 |
| Hidden Size | 1280 | – |
| Attention Heads | 16 | – |
| Patch Size | 14 | – |