cs.AISep 29, 2026

OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

Authors: Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang

Organizations: Fudan University, Shanghai, China · Beijing Institute of Technology, Beijing, China · Chongqing Ant Consumer Finance Co,. Ltd

Abstract

Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.

Figures & tables

Appendix figures & tables52 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

    Apr 21, 2026Xue Xia, Chengkai Yao, Mingyu Tsoi +10AI Coding AgentsML Reproducibility

  2. Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

    Sep 23, 2026Mengran Li, Bo Li, Chengyang Zhang +3AI Agents for Scientific DiscoveryAI Agent Auditing

  3. CellScientist: From Execution Feedback to Auditable Model-Revision Trajectories for Cellular Perturbation Prediction

    May 8, 2026Mengran Li, Bo Li, Jiaying Wang +7Single-Cell Perturbation ModelingScientific Workflow Automation