cs.AIAug 7, 2026

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Authors: Taolin HanYuchen ZhangJinghang WangYun WuWai Yuet ChiuZhaohai LiYifei ZhangJinxin Wang+17 more

Organizations: 1Alibaba Group · University of Chinese Academy of Sciences · Taolin Han1,3* · 4Tsinghua University · Yuchen Zhang1,4* · Jinghang Wang1* · Yun Wu1 · Wai Yuet Chiu1 · 2Qwen Team, Alibaba Group · Zhaohai Li2 · University of Alberta · Yifei Zhang1,5 · Jinxin Wang1,4 · 6Zhejiang University · Yuhao Zhou1,6 · Chen Zhao1,4 · Jiajia Li1 · Jiaxin Li1 · Qile Jin1 · Kewei Sun1 · Shuang Wu1 · Weiqi Zhai1 · Renquan Lv1,6 · Junchao Li1 · Ruodan Chen1 · Qingteng Chen1 · Zhibo Yang2 · Hu Wei1 · Lin Qu1 · Shuai Bai2,† · Bing Zhao1,†

Abstract

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.

Explore similar work

CardsList