cs.AIOct 5, 2026

FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification

Authors: Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, +2 more

Organizations: Department of Computer Science and Engineering, The Ohio State University · Department of Pharmaceutical Sciences, University of Illinois Chicago · College of Pharmacy, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Department of Chemical and Biomolecular Engineering, University of Pennsylvania · Translational Data Analytics Institute, The Ohio State University

Abstract

As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits by large language models (LLMs), and five negative candidate generation methods. Our evaluation finds that no verifier leads across all sources: LLMs given only the criterion are competitive with dedicated verifiers, while forward models perform best on retrosynthesis proposals but reject most feasible edits of recorded reactions at the evaluated operating points. Looking beyond aggregate scores, both forward models perform below chance when separating infeasible alternative disconnections from feasible generated candidates. To study whether negative supervision addresses these weaknesses, we also release a corpus of over 14 million recorded reactions and generated negative candidates. In matched training comparisons, adding a mixture of generated negatives to forward training raises mean AUROC across sources, but these gains do not extend to retrosynthesis proposals. Varying the generation method further shows that the largest gain on generated candidates coincides with worse proposal screening. These findings motivate evaluating verifiers against experts across sources and designing negatives for transfer to model proposals.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

    Jul 1, 2026Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas +4Reaction

  2. onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

    Aug 3, 2026Brandon Wang, Andrei S. Tyrin, Daniil A. BoikoChemistry

  3. From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

    Jun 2, 2026Hongyu Guo, Hao Li, He Cao +2ChemistryAnswer