FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification
Authors: Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, +2 more
Organizations: Department of Computer Science and Engineering, The Ohio State University · Department of Pharmaceutical Sciences, University of Illinois Chicago · College of Pharmacy, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Department of Chemical and Biomolecular Engineering, University of Pennsylvania · Translational Data Analytics Institute, The Ohio State University
As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits by large language models (LLMs), and five negative candidate generation methods. Our evaluation finds that no verifier leads across all sources: LLMs given only the criterion are competitive with dedicated verifiers, while forward models perform best on retrosynthesis proposals but reject most feasible edits of recorded reactions at the evaluated operating points. Looking beyond aggregate scores, both forward models perform below chance when separating infeasible alternative disconnections from feasible generated candidates. To study whether negative supervision addresses these weaknesses, we also release a corpus of over 14 million recorded reactions and generated negative candidates. In matched training comparisons, adding a mixture of generated negatives to forward training raises mean AUROC across sources, but these gains do not extend to retrosynthesis proposals. Varying the generation method further shows that the largest gain on generated candidates coincides with worse proposal screening. These findings motivate evaluating verifiers against experts across sources and designing negatives for transfer to model proposals.
Figures & tables
Source
Produced by
In practice
Controlled
Feas.
Infeas.
Total
Retrosynthesis proposals (RP)
Eight retrosynthesis models
✓
197
85
282
Zero-yield records (ZR)
Experimental records
✓
79
14
93
LLM perturbations (LP)
Two LLMs editing recorded reactions
32
55
87
Generated reactions (GR)
Five generation methods
✓
33
256
289
Total
341
410
751
Table 1: Composition of the FREA benchmark. In practice : reactions produced by retrosynthesis models or attempted in the laboratory. Controlled : candidates obtained from a recorded reaction by a known negative candidate generation method. Table C.1 gives the per-model and per-method breakdown, including the fraction judged infeasible for each GR generation method.
Figure 1: The five methods used to construct GR candidates, illustrated from a common parent reaction (top). RA, RR and NC change the reactants while keeping the product; FT and ET change the product while keeping the reactants. The methods generate intended negatives, whose feasibility is determined by expert annotation.
RP (197/85)
ZR (79/14)
LP (32/55)
GR (33/256)
Avg.
System
AUROC
BAcc
AUROC
BAcc
AUROC
BAcc
AUROC
BAcc
AUROC
BAcc
Large language models, criterion in the prompt
Claude Opus 5
–
69.9
–
64.3
–
76.1
–
61.8
–
68.0
GPT-6 Astra
–
66.8
–
63.7
–
69.6
–
62.7
–
65.7
Forward-model proxies
Chemformer (top-10)
–
70.6
–
67.2
–
63.3
–
50.1
–
62.8
Table 2: Results on FREA by source (feasible/infeasible counts in parentheses). Avg. denotes the four-source mean. Dashes indicate no continuous score. Best values per column are bold. Counts describe the benchmark; one unparseable Claude response on LP is excluded from that model’s metrics.
Reactants changed
Products changed
System
RA(50)
RR(53)
NC(54)
FT(44)
ET(55)
Forward-model proxies
Chemformer (LL)
44.7 [32, 57]
83.5 [75, 93]
26.2 [15, 37]
87.5 [79, 96]
94.9 [90, 100]
BARTSmiles (LL)
46.1 [33, 60]
79.8 [70, 90]
32.5 [20, 45]
83.5 [74, 93]
93.7 [88, 99]
Trained on our generated negative candidates
BARTSmiles (LL + neg.)
72.7 [62, 84]
53.4 [40, 66]
94.2 [90, 99]
67.7 [55, 81]
78.1 [67, 89]
Table 3: AUROC on GR by the method that produced each candidate, with DeLong 95% confidence intervals. Each column ranks the infeasible candidates of one method (count in parentheses) against the same 33 feasible GR candidates, so 50 is chance and a value below 50 ranks infeasible candidates above feasible ones. The measure is threshold-free, so it covers only systems that return a score; rejection rates at each operating point, including those of the language models and the top-10 readouts, are in .
RP
ZR
LP
GR
Avg.
BARTSmiles
78.9
79.5
66.1
66.9
72.8
+ w/o neg.
77.8 −1.1
76.5 −3.0
66.8 +0.7
64.5 −2.4
71.4 −1.4
+ w/ neg.
77.0 −1.8
80.8 +1.4
69.4 +3.4
73.5 +6.7
75.2 +2.4
Δ (w/ − w/o)
−0.7
+4.3
+2.6
+9.0
+3.8
Table 4: One further epoch of training without and with generated negative candidates, with matched positive reactions and training schedules. AUROC by source, log-likelihood readout. w/ neg. is BARTSmiles (LL + neg.) of Tables 2 and 3 .
RP
ZR
LP
GR
Avg.
w/o neg.
77.6
79.0
65.5
64.8
71.7
w/ only RA neg.
78.6 +1.0
76.8 −2.3
67.3 +1.8
68.0 +3.2
72.7 +0.9
w/ only RR neg.
78.9 +1.3
79.0 +0.0
74.4 +8.9
67.4 +2.6
74.9 +3.2
w/ only NC neg.
74.8 −2.8
76.4 −2.6
67.4 +1.9
76.9 +12.1
73.9 +2.2
w/ only FT neg.
70.7 −6.9
76.2 −2.8
62.2 −3.4
52.4 −12.4
65.4 −6.4
w/ only ET neg.
77.3 −0.3
77.3 −1.7
68.5 +3.0
64.5 −0.3
71.9 +0.2
Table 5: The same recipe with the negatives restricted to the candidates of one generation method, and with an equal share of all five. AUROC by source, log-likelihood readout. Grey: difference from w/o neg. , which is a separately trained control using the same positive-only recipe as Table 4 and the shared positive stream of this experiment ( ).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Resource
Scope
Expert labels
Negatives
Source of negatives
Jiang et al. (2021)
One laboratory
—
✓
failed experiments (private)
Zhong et al. (2025)
Acid–amine couplings
—
✓
high-throughput experiments
DORA-XGB ( Chainani et al., 2025 )
Enzymatic
—
✓
unreported reactions
RetroGFN ( Gaiński et al., 2025 )
Small molecule
—
✓
templates, product swap
RetroTrim ( Sadowski et al., 2025 )
Drug-like targets
✓
✓
expert-rated proposals
CREED ( Zagribelnyy et al., 2026b )
Small molecule
—
—
—
Appendix
Table A.1: Selected resources related to reaction feasibility. Expert labels : whether individual reactions carry feasibility judgments by chemists. Negatives : whether the resource includes experimental failures or reactions labeled infeasible.
Role
Scope
Catalysis, ligation, solvation
Species that supply no retained heavy atom.
Acid or base assistance
Proton transfer or pH control.
Reduction, hydrogen donation
Conventional reducing systems and hydrogen sources that transfer no retained heavy atom.
Oxidation
Conventional oxidants that transfer no retained heavy atom.
Water
Water in conventional hydrolysis or workup.
Activation, dehydration, coupling or deprotection assistance
Species that enable these operations without supplying retained heavy atoms, within the permitted single-step or one-pot operation.
Appendix
Table B.1: Permitted auxiliary roles. An omitted species may be assumed only when it acts in one of these roles, within the stated scope. Any other species is substantive and must be listed.
Transformation
Auxiliary (may be assumed)
Substantive (may not)
Cross-coupling
Pd catalyst, ligand, base, solvent
organoboron or halide partner
Esterification
acid catalyst, activating reagent
alcohol
Amide coupling
coupling reagent, base
amine
Reductive amination
NaBHX4 , HX2 /Pd
amine
Wittig olefination
base
ylide or phosphonium salt
Mitsunobu reaction
DEAD, PPhX3
pronucleophile
Appendix
Table B.2: The substantive–auxiliary boundary applied to common transformations. Within a row, the species in the middle column acts in a permitted role and may be assumed when omitted. The species in the right column is a reaction partner and must be listed.
#
Check
Reason recorded if the check fails
1
Can every listed species exist and be present?
unstable_or_nonexistent_reactant
2
Can the listed products exist as drawn?
unstable_or_nonexistent_product
3
Can the listed reactants give the products at all?
missing_substantive_reactant or reactants_cannot_give_product
4
Can the intended pathway yield isolable products despite competing processes?
reactivity_conflict
5
Can the products be obtained with the specified regio- and stereochemistry under some reasonable conditions?
regio_stereo_violation
6
Can the products be obtained in one operation, without isolating an intermediate?
requires_isolated_steps
Appendix
Table B.3: Ordered decision procedure. A failure at checks 1–3 ends the procedure; otherwise, checks 4–6 are assessed and their reasons may be recorded together. Reasons are defined in Table B.4 . The two reasons at check 3 are alternatives ( Section B.3 ).
Reason
Definition
unstable_or_nonexistent_reactant
A listed species, substrate or reagent, is drawn incorrectly, cannot exist, or is too unstable to be present. Lack of bottle stability alone does not qualify a species that can credibly form in situ .
unstable_or_nonexistent_product
A listed product has a structural or stability defect that prevents it from reasonably existing or being obtained as drawn. Unfamiliarity, or failure of software to parse it, does not qualify.
missing_substantive_reactant
The intended transformation is recognizable, but a partner that is not a permitted auxiliary is not listed.
reactants_cannot_give_product
The listed species can exist, but lack a credible mechanistic route to the products: the skeletons do not correspond, a listed reagent cannot perform its required role, or there is a concrete mechanistic obstacle.
reactivity_conflict
A functional group or listed component causes a competing reaction, quenching or decomposition that intercepts the intended pathway, so that the products cannot be obtained in isolable form within the permitted single-step or one-pot operation.
regio_stereo_violation
No reasonable conditions give the products with the specified regio- or stereochemistry in isolable form. Typical cases are an outcome opposite to the one a stereospecific or regiospecific mechanism fixes, and a substitution pattern that no available mechanism produces. An isolable minor isomer does not qualify.
Appendix
Table B.4: Failure reasons recorded for infeasible proposals, in the order of the checks in Table B.3 .
Annotator 2
Annotator 1
Feasible
Infeasible
Uncertain
Total
Feasible
255
77
0
332
Infeasible
90
336
0
426
Uncertain
0
1
0
1
Total
345
414
0
759
Appendix
Table B.5: Independent labels of the two annotators before adjudication, on the 759 reactions of the annotation pool outside the calibration set. Rows give the first annotator’s label and columns the second’s.
Source
Origin
Feasible
Infeasible
Total
Infeas. (%)
RP
LARC (Claude 3.5 Sonnet)
43
12
55
21.8
LARC (Mistral NeMo)
42
12
54
22.2
G2Retro
17
13
30
43.3
EditRetro
18
12
30
40.0
LlaSMol
16
13
29
44.8
UAlign
14
15
29
51.7
Appendix
Table C.1: Breakdown of the benchmark by the model, language model or negative candidate generation method that produced each reaction. Infeas. (%) is the share of a row’s reactions the experts judged infeasible, which for a generation method is the share of its candidates that are genuinely negative.
Rule
Scope
Condition for invalidation
General validity
all_invalid_smiles
All
Canonical reactant or product SMILES is null
product_in_reactant
All
Any product component appears verbatim among the reactant components
unstable_substructure
All
Appendix
Table D.1: Summary of the 16 filtering rules used to construct the generated corpus. “All” denotes filters applied to both recorded positives and generated candidates; these general filters also screen the benchmark before annotation. Methods are described in Appendix D .
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.
Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas +4
1École Polytechnique Fédérale de Lausanne (EPFL), Switzerland · 2Ghent University, Belgium · 3National Centre of Competence in Research (NCCR) Catalysis, Switzerland
Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.
Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.
Hongyu Guo, Hao Li, He Cao +2
Peking University, Shenzhen Graduate School · International Digital Economy Academy (IDEA)