OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
Organizations: Fudan University, Shanghai, China · Beijing Institute of Technology, Beijing, China · Chongqing Ant Consumer Finance Co,. Ltd
Abstract
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
Figures & tables
| Benchmark | Scale | What it evaluates / contains | Layer |
|---|---|---|---|
| Virtual Cell Challenge ( Roohani et al., 2025 ) | 300k cells | CRISPRi perturbation responses | Simulation |
| OP3 ( Szałata et al., 2024 ) | 144 compounds | Held-out perturbation expression | Simulation |
| PerturBench ( Wu et al., 2025b ) | 6 datasets | Perturbation-response prediction | Simulation |
| scPerturBench ( Wei et al., 2026b ) | 29 datasets | Unseen perturbations and contexts | Simulation |
| VCBench ( Mao et al., 2026 ) | 7 datasets | In-the-wild perturbation response | Simulation |
| MVCBench ( Li et al., 2026a ) | 1.1M profiles | Transcriptomic and morphology outputs | Simulation |
| MCQ accuracy (%) | AIVC-Judge (1–5) | |||||||
| Model | L1 | L2 | L3 | Avg | L1 | L2 | L3 | Avg |
| Proprietary models | ||||||||
| Claude-Sonnet-4-5 ( Anthropic, 2025 ) | 45.3 | 52.1 | 52.5 | 49.4 | 2.55 | 2.92 | 2.62 | 2.68 |
| Grok-4.6 ( xAI, 2026 ) | 48.6 | 53.9 | 49.2 | 50.3 | 3.26 | 3.39 | 2.56 | 3.10 |
| GPT-5.6-luna ( OpenAI, 2025 ) | 47.2 | 48.3 | 33.9 | 43.7 | 3.25 | 3.47 | 2.72 | 3.16 |
| GPT-5.6-sol ( OpenAI, 2025 ) | 58.6 | 60.6 | 44.2 | 55.1 | 3.37 | 3.68 | 2.76 | 3.28 |
| Method | Resource | L1 | L2 | L3 | Avg | ||
|---|---|---|---|---|---|---|---|
| Base | – | 34.2 | 37.4 | 29.6 | 33.82 | – | – |
| SFT, full corpus | 548k, 2 epochs | 33.8 | 36.5 | 32.7 | 34.28 | +0.46 | 0.42 |
| SFT, rank 32 | 100k, 2 epochs | 34.6 | 38.5 | 34.8 | 35.79 | +1.97 | |
| Multimodal RAG | 548k index, | 35.6 | 39.3 | 31.2 | 35.41 | +1.60 |
Appendix figures & tables52 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Model ID / version | Source and role |
|---|---|---|
| GPT-5.6-sol | gpt-5.6-sol | OpenRouter API; main benchmark |
| GPT-5.6-luna | gpt-5.6-luna | OpenRouter API; main benchmark |
| Grok-4.6 | grok-4.6 | OpenRouter API; main benchmark |
| Claude-Sonnet-4-5 | claude-sonnet-4-5 | OpenRouter API; main benchmark |
| DeepSeek-V4-Flash-Vision-Exp | deepseek-v4-flash-vision-exp | Official DeepSeek API; production AIVC-Judge with image input, temperature 0 |
| GPT-5.6-terra | gpt-5.6-terra | OpenRouter API; auxiliary task-type annotation and distractor generation |
| Level | Required operation | P-E-D alignment | Bloom levels | |
|---|---|---|---|---|
| L1 | Infer a value, direction, relation, or outcome from displayed evidence | Predict analogue | 2, 3, 4 | 2,576 |
| L2 | Explain an observed result through a biologically grounded mechanism | Explain | 4 | 1,763 |
| L3 | Propose a testable hypothesis grounded in the displayed evidence | Discover prerequisite | 5 | 1,738 |
| Task type | Gemini-3.7-flash-high | GPT-5.6-terra |
|---|---|---|
| propose_novel | 1,679 | 1,676 |
| design_experiment | 31 | 35 |
| other | 21 | 20 |
| evaluate_given | 7 | 7 |
| Stage | Retained count |
|---|---|
| OmniScience initial figure–caption pairs | 1,525,179 |
| Stage 0 keyword filtering | 98,632 |
| QA generation + collaborative filtering | 7,923 |
| Unanimous three-annotator audit | 6,077 |
| Theme | Representative terms |
|---|---|
| Virtual cells | virtual cell; AI virtual cell; digital twin cell; whole-cell model |
| Perturbations | perturbation; CRISPR screen; knockout; drug or dose response |
| Single-cell/spatial | single-cell RNA-seq; spatial transcriptomics; UMAP; RNA velocity |
| Multi-omics | multi-omics; ATAC-seq; proteomics; chromatin accessibility |
| Mechanisms | signalling pathway; gene-regulatory network; causal mechanism |
| Disease biology | cancer; drug resistance; therapeutic hypothesis; patient-derived model |
| P-E-D label | Bloom level | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Level | Predict | Explain | Discover | Total | B2 | B3 | B4 | B5 | Total |
| L1 | 2,362 (91.7%) | 214 | 0 | 2,576 | 261 | 773 | 1,542 | 0 | 2,576 |
| L2 | 0 | 1,763 (100%) | 0 | 1,763 | 0 | 0 | 1,763 | 0 | 1,763 |
| L3 | 9 | 9 | 1,720 (99.0%) | 1,738 | 0 | 0 | 0 | 1,738 | 1,738 |
| Bloom’s taxonomy | P-E-D function | ||||||
|---|---|---|---|---|---|---|---|
| Level | Category | Share | Function | Operational meaning | Share | ||
| 2 | Understand | 261 | 4.3% | Predict | Evidence-conditioned inference | 2,371 | 39.0% |
| 3 | Apply | 773 | 12.7% | Explain | Mechanistic explanation | 1,986 | 32.7% |
| 4 | Analyze | 3,305 | 54.4% | Discover | Testable-hypothesis proposal grounded in displayed evidence | 1,720 | 28.3% |
| 5 | Evaluate | 1,738 | 28.6% | ||||
| Unit | OmniVCTrain | OmniVCBench |
|---|---|---|
| Released QA rows | 548,450 | 6,077 |
| Unique source records | 96,007 | 2,625 |
| Unique papers | 33,802 | 1,080 |
| Materialized image paths | 327,601 | 4,428 |
| QA per source record (mean / median) | 5.71 / 6 | 2.32 / 2 |
| QA per paper (mean / median) | 16.23 / 12 | 5.63 / 4 |
| Subject | OmniVCTrain | OmniVCBench |
|---|---|---|
| Biology | 75.5% | 73.3% |
| Medicine | 21.7% | 26.4% |
| Physics | 0.2% | 0.4% |
| Others (train only) | 2.6% | 0.0% |
| Field | Median chars | Mean chars | Median words | Mean words |
|---|---|---|---|---|
| Train questions | 147 | 154.2 | 24 | 25.3 |
| Train answers | 235 | 249.1 | 35 | 37.3 |
| Bench questions | 234 | 238.0 | 37 | 37.3 |
| Bench answers | 358 | 364.1 | 52 | 53.6 |
| Year | Papers | QA items |
|---|---|---|
| 2011 | 30 (2.8%) | 182 (3.0%) |
| 2012 | 121 (11.2%) | 723 (11.9%) |
| 2013 | 233 (21.6%) | 912 (15.0%) |
| 2014 | 25 (2.3%) | 110 (1.8%) |
| 2015 | 160 (14.8%) | 737 (12.1%) |
| 2016 | 152 (14.1%) | 861 (14.2%) |
| Model | 2011 | 2012 | 2013 | 2014 | 2015 | 2016 | 2017 |
|---|---|---|---|---|---|---|---|
| GPT-5.6-sol | 51.1 | 53.3 | 61.0 | 57.3 | 55.8 | 54.1 | 53.8 |
| Grok-4.6 | 49.5 | 55.3 | 42.7 | 59.1 | 57.9 | 50.4 | 49.1 |
| Claude-Sonnet-4-5 | 49.5 | 43.9 | 52.6 | 44.6 | 48.2 | 48.6 | 50.6 |
| GPT-5.6-luna | 47.3 | 38.0 | 48.9 | 38.2 | 41.7 | 42.0 | 44.7 |
| Qwen3-VL-8B-SFT | 29.1 | 32.6 | 36.6 | 38.2 | 30.3 | 35.8 | 34.8 |
| Qwen3-VL-8B | 29.1 | 32.2 | 35.8 | 35.5 | 30.8 | 33.7 | 34.8 |
| Species | Assays | Cell types (top 10) | |||
|---|---|---|---|---|---|
| mouse | 49.9 | microscopy | 61.9 | HeLa | 7.4 |
| human | 38.2 | immunoblot | 41.8 | HEK293T | 4.1 |
| rat | 5.1 | genetic perturbation | 37.4 | HEK293 | 3.4 |
| bacteria | 3.4 | qPCR | 36.9 | U2OS | 2.3 |
| fruit fly | 2.5 | binding interaction | 19.4 | fibroblasts | 2.3 |
| yeast | 2.3 | flow cytometry | 19.4 | MDA-MB-231 | 2.2 |
| Model | L1 | L2 | L3 | Avg |
|---|---|---|---|---|
| GPT-5.6-sol | 37.8 | 39.1 | 25.6 | 34.7 |
| LLaVA-OneVision-7B | 23.6 | 25.3 | 23.3 | 24.0 |
| Qwen3-VL-2B | 22.0 | 23.0 | 19.8 | 21.7 |
| InternVL3.5-2B | 11.0 | 16.1 | 15.1 | 13.7 |
| SmolVLM2-2.2B | 11.8 | 16.1 | 14.0 | 13.7 |
| Criterion | Requirement |
|---|---|
| Strict incorrectness | The screening target is an option contradicted by the evidence or incompatible with the requested answer; an alternative hypothesis requires a substantive scientific distinction. |
| Plausibility | The error should reflect a biologically or visually plausible model failure. |
| Relevance and grounding | The option must answer the question and refer only to entities or patterns present in the item. |
| Diversity | The five distractors should represent distinct error modes rather than paraphrases. |
| Uniqueness | The screening target is one correct option with semantically distinct distractors; subsequent audits examine violations of this target. |
| Style consistency | Length, grammar, specificity, and units should not reveal the correct answer. |
| Distractor generator | Direct | MDHNM | (pp) [95% CI] | Discordant | McNemar |
|---|---|---|---|---|---|
| Grok-4.6 | 87.7 | 57.3 | [ , ] | 111/20 | |
| Gemini-3.1-flash-lite | 95.3 | 57.3 | [ , ] | 120/6 | |
| Kimi-K2.6 | 93.3 | 57.3 | [ , ] | 116/8 |
| Level | Distractor generator | Direct | MDHNM | (pp) |
|---|---|---|---|---|
| L1 | Grok-4.6 | 86.6 | 63.8 | |
| L1 | Gemini-3.1-flash-lite | 91.3 | 63.8 | |
| L1 | Kimi-K2.6 | 92.9 | 63.8 | |
| L2 | Grok-4.6 | 86.2 | 65.5 | |
| L2 | Gemini-3.1-flash-lite | 97.7 | 65.5 | |
| L2 | Kimi-K2.6 | 92.0 | 65.5 |
| Distractor arm | GPT-5.6-sol | Grok-4.6 |
|---|---|---|
| gemini-3.1-flash-lite direct | 93.7 | 94.0 |
| gemini-3.1-flash-lite mined | 90.0 | 91.3 |
| gemini-3.7-flash-high direct | 94.3 | 93.7 |
| gpt-5.6-terra direct | 92.0 | 91.7 |
| original MDHNM | 57.3 | 48.3 |
| Level | Scoring dimensions |
|---|---|
| L1 | correctness; evidence support |
| L2 | faithfulness; causal completeness; mechanistic granularity; evidence consistency |
| L3 | judgment correctness; argumentation quality; evidence weighing |
| Component | Reported configuration |
|---|---|
| Judge | DeepSeek-V4-Flash-Vision-Exp ( deepseek-v4-flash-vision-exp ); single pointwise judge; temperature 0 |
| Response input | question and candidate response |
| Closed evidence | figure caption, surrounding paper context, and complete reference answer |
| Visual input | original figure image set, as shown to the answering model |
| Reference key points | disabled; the complete reference answer provides the scoring anchor |
| Anonymization | evaluated model identity omitted from the judge input |
| Annotator | Correct/assigned | Overall [95% CI] | L1 | L2 | L3 |
|---|---|---|---|---|---|
| Human 1 | 166/300 | 55.3 [50.0, 60.7] | 55.9 | 57.5 | 52.3 |
| Human 2 | 214/300 | 71.3 [65.8, 76.7] | 76.4 | 81.6 | 53.5 |
| Human 3 | 224/300 | 74.7 [69.1, 79.5] | 75.6 | 87.4 | 60.5 |
| Annotator pair | Cohen’s [95% CI] | Exact agreement |
|---|---|---|
| Human 1–Human 2 | 0.456 [0.391, 0.520] | 54.8% |
| Human 1–Human 3 | 0.457 [0.391, 0.523] | 54.8% |
| Human 2–Human 3 | 0.855 [0.805, 0.897] | 88.0% |
| Scorer | GPT-5.6-sol | Qwen3-VL-8B | LLaVA-Med-Mistral-7B |
|---|---|---|---|
| Human 1 | 3.163 | 2.186 | 1.727 |
| Human 2 | 4.127 | 2.657 | 1.917 |
| Human 3 | 3.933 | 2.447 | 1.797 |
| Production judge | 3.354 | 2.257 | 1.790 |
| Scope | Paired cells | Spearman [95% CI] | MAE | Bias |
|---|---|---|---|---|
| All, raw overall | 884 | 0.877 [0.855, 0.897] | 0.464 | +0.188 |
| L1 | 377 | 0.909 [0.887, 0.926] | 0.389 | +0.193 |
| L2 | 257 | 0.875 [0.837, 0.905] | 0.450 | +0.014 |
| L3 | 250 | 0.709 [0.637, 0.774] | 0.592 | +0.360 |
| GPT-5.6-sol | 291 | 0.840 [0.797, 0.874] | 0.608 | +0.379 |
| Qwen3-VL-8B | 296 | 0.865 [0.823, 0.896] | 0.430 | +0.169 |
| Scorer pair | Paired cells | Spearman [95% CI] | QWK [95% CI] |
|---|---|---|---|
| Human 1–Human 2 | 887 | 0.820 [0.788, 0.849] | 0.738 [0.694, 0.776] |
| Human 1–Human 3 | 887 | 0.835 [0.807, 0.859] | 0.800 [0.766, 0.829] |
| Human 2–Human 3 | 900 | 0.953 [0.942, 0.962] | 0.944 [0.930, 0.956] |
| Human 1–judge | 884 | 0.838 [0.810, 0.863] | 0.851 [0.821, 0.876] |
| Human 2–judge | 897 | 0.832 [0.800, 0.860] | 0.779 [0.740, 0.815] |
| Human 3–judge | 897 | 0.854 [0.830, 0.876] | 0.837 [0.806, 0.863] |
| Check | Cross-set candidates | Confirmed duplicates |
|---|---|---|
| Normalized DOI / paper ID (exact) | 0 | 0 |
| Normalized title (exact) | 0 | 0 |
| Title char-TFIDF cosine | 0 | 0 |
| Title char-TFIDF cosine | 0 | 0 |
| Field | Train uniques | Bench uniques | Exact pairs | Near pairs |
|---|---|---|---|---|
| Question | 548,390 | 6,077 | 0 | 0 |
| Caption | 96,007 | 2,625 | 0 | 0 |
| Answer (exact only) | 544,652 | 6,077 | 0 | – |
| Perturbation | Retrieval | Acceptance |
|---|---|---|
| Question: append 10% words | 100% | 100% |
| Question: delete one word | 100% | 98.2% |
| Question: substitute one word | 100% | 87.8% |
| Caption: append 10% words | 100% | 100% |
| Caption: delete one word | 100% | 100% |
| Caption: substitute one word | 100% | 100% |
| Check | Pairs | Bench images involved | Confirmed duplicates |
|---|---|---|---|
| SHA-256 exact | 0 | 0 | 0 |
| Single-hash perceptual candidates | 22,768 | 606 | – |
| Dual-hash (pHash and dHash) Hamming | 45 | 8 | 0 |
| Perturbation | Detection rate |
|---|---|
| JPEG quality 75 | 100% |
| Resize to half | 100% |
| Brightness | 97.7% |
| Crop 2% | 95.0% |
| Add 2% border | 97.3% |
| Model | Item-level [95% CI] | Paper-weighted [95% CI] |
| MCQ accuracy (%) | ||
| GPT-5.6-sol | 55.093 [53.768, 56.444] | 56.033 [54.267, 57.794] |
| Grok-4.6 | 50.304 [48.862, 51.729] | 51.492 [49.738, 53.286] |
| Claude-Sonnet-4-5 | 49.366 [47.937, 50.768] | 50.123 [48.313, 51.889] |
| GPT-5.6-luna | 43.739 [42.415, 45.135] | 44.478 [42.678, 46.314] |
| Qwen3-VL-8B-SFT | 34.277 [33.018, 35.566] | 34.111 [32.444, 35.821] |
| Scope | Pearson | Spearman |
|---|---|---|
| Published full-track, item | 0.938 ( ) | 0.964 ( ) |
| Published full-track, paper | 0.945 | 0.955 |
| Model-specific matched, item | 0.936 | 0.964 |
| Model-specific matched, paper | 0.943 | 0.955 |
| Strict common, item | 0.935 [0.919, 0.946] ( ) | 0.964 [0.927, 0.973] ( ) |
| Strict common, paper | 0.943 [0.923, 0.956] | 0.955 [0.927, 0.973] |
| Excluded family | Pearson | Spearman | |
|---|---|---|---|
| OpenAI GPT | 9 | 0.956 | 0.967 |
| xAI Grok | 10 | 0.932 | 0.964 |
| Anthropic Claude | 10 | 0.971 | 0.964 |
| Qwen3-VL (incl. SFT) | 8 | 0.942 | 0.929 |
| Qwen2.5-VL | 10 | 0.939 | 0.939 |
| InternVL | 10 | 0.943 | 0.952 |
| Setting | Data | Epochs | Rank | Avg | ||
|---|---|---|---|---|---|---|
| Full corpus | 548k | 2 | 16 | 34.28 | +0.46 | 0.42 |
| Rank 32, 100k subset | 100k | 2 | 32 | 35.79 | +1.97 |
| Retrieval signal | L1 | L2 | L3 | Avg | ||
|---|---|---|---|---|---|---|
| Joint image–question | 35.6 | 39.3 | 31.2 | 35.41 | +1.60 | |
| Image only | 34.1 | 37.5 | 28.9 | 33.62 | .64 | |
| Text only | 34.1 | 37.3 | 28.5 | 33.40 | .31 |
| Model | Base | RAG | ||
|---|---|---|---|---|
| Qwen3-VL-8B | 33.82 | 35.41 | +1.60 | |
| Qwen3-VL-2B | 21.69 | 23.19 | +1.50 | |
| InternVL3.5-8B | 30.28 | 31.04 | +0.76 | .10 |
| Pilot | Simulator / data | Prediction target | Evidence choice | Cases / readouts |
|---|---|---|---|---|
| Gene interactions | GEARS ( Roohani et al., 2023 ) ; Norman K562 CRISPRa ( Norman et al., 2019 ) | non-additive double-gene response per program | one cellular program in reserved AB cells | 12 / 48 |
| Drug combinations | CPA trained locally ( Lotfollahi et al., 2023 ) ; ComboSciPlex A549 | program mean change under held-out combinations | one program in 48 reserved cells | 10 / 30 |
| Real microscopy | BBBC021 ( Caie et al., 2010 ; Ljosa et al., 2012 ) ; MCF-7, 5 MoA classes | phenotype-supported mechanism class | F-actin channel or a second field of the same well | 9 / 9 |
| Signaling | discrete Bayesian network; Sachs discretized cells ( Sachs et al., 2005 ) | high-state fraction change under each intervention | one non-target readout in 96 reserved cells | 5 / 30 |
| Pilot | Readouts | Initial | Self-review | After evidence | Corrections |
|---|---|---|---|---|---|
| Gene interactions | 48 | 77.1 | 77.1 | 89.6 | / |
| Drug combinations | 30 | 93.3 | 80.0 | 100.0 | / |
| Real microscopy | 9 | 66.7 | 66.7 | 66.7 | / |
| Signaling | 30 | 66.7 | 60.0 | 70.0 | / |
| L1 | L2 | L3 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | C | ES | F | CC | MG | EC | J | A | EW |
| GPT-5.6-sol | 3.35 | 3.35 | 3.88 | 3.72 | 3.63 | 3.85 | 4.04 | 1.92 | 1.67 |
| GPT-5.6-luna | 3.23 | 3.23 | 3.70 | 3.49 | 3.46 | 3.66 | 3.94 | 1.92 | 1.67 |
| Grok-4.6 | 3.23 | 3.25 | 3.64 | 3.38 | 3.36 | 3.62 | 3.87 | 1.71 | 1.53 |
| Claude-Sonnet-4-5 | 2.52 | 2.55 | 3.07 | 2.99 | 2.97 | 3.10 | 3.64 | 1.99 | 1.73 |
| Qwen3-VL-8B-SFT † | 2.50 | 2.55 | 2.71 | 2.25 | 2.34 | 2.50 | 3.01 | 1.44 | 1.33 |
| Model | Single | Multi | (pp) |
|---|---|---|---|
| GPT-5.6-sol | 54.5 | 55.7 | +1.2 |
| Claude-Sonnet-4-5 | 45.4 | 53.2 | +7.8 |
| Qwen3-VL-8B | 33.6 | 34.0 | +0.4 |
| InternVL3.5-8B | 29.6 | 30.9 | +1.3 |
| LLaVA-OneVision-7B | 25.0 | 26.1 | +1.1 |