Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
Figures & tables
Figure 1: The interpretation component within the AIVC workflow. The simulation layer predicts perturbed cell states; experiments return observations; the interpretation layer compares and explains the evidence, informing the next experiment or model revision (feedback arrow). OmniVCBench evaluates only the interpretation component (highlighted), through three scientific task roles (L1–L3); the remaining stages are application context.
Benchmark
Scale
What it evaluates / contains
Layer
Virtual Cell Challenge ( Roohani et al., 2025 )
∼ 300k cells
CRISPRi perturbation responses
Simulation
OP3 ( Szałata et al., 2024 )
144 compounds
Held-out perturbation expression
Simulation
PerturBench ( Wu et al., 2025b )
6 datasets
Perturbation-response prediction
Simulation
scPerturBench ( Wei et al., 2026b )
29 datasets
Unseen perturbations and contexts
Simulation
VCBench ( Mao et al., 2026 )
7 datasets
In-the-wild perturbation response
Simulation
MVCBench ( Li et al., 2026a )
∼ 1.1M profiles
Transcriptomic and morphology outputs
Simulation
Table 1: Recent virtual-cell and scientific-figure benchmarks.
Figure 2: Construction and paired evaluation pipeline. Filtered OmniScience records are expanded into QA pairs and curated; retained items form OmniVCBench (open + MCQ tracks), while coarse QA forms OmniVCTrain for SFT/RAG.
Figure 3: AIVC-Judge. Given the question, response, original figure image, and closed textual evidence, the judge verifies atomic claims against evidence, then scores level-specific dimensions and returns dimension scores, an overall score, and an evidence trace.
MCQ accuracy (%) ↑
AIVC-Judge (1–5) ↑
Model
L1
L2
L3
Avg
L1
L2
L3
Avg
Proprietary models
Claude-Sonnet-4-5 ( Anthropic, 2025 )
45.3
52.1
52.5
49.4
2.55
2.92
2.62
2.68
Grok-4.6 ( xAI, 2026 )
48.6
53.9
49.2
50.3
3.26
3.39
2.56
3.10
GPT-5.6-luna ( OpenAI, 2025 )
47.2
48.3
33.9
43.7
3.25
3.47
2.72
3.16
GPT-5.6-sol ( OpenAI, 2025 )
58.6
60.6
44.2
55.1
3.37
3.68
2.76
3.28
Table 2: Paired results on OmniVCBench : MCQ accuracy (%, left) and AIVC-Judge scores (1–5, right). Within each block, models without Judge scores precede models sorted by Judge Avg (ascending). † denotes the rank-16, two-epoch SFT configuration on OmniVCTrain . Each statistic summarizes performance over individual QA items in the corresponding track.
Figure 4: Evaluation agreement and adaptation gains. (a) MCQ accuracy vs. AIVC-Judge score over eleven models. (b) Qwen3-VL-8B adaptation across L1–L3. (c) L3 dimension profile.
Method
Resource
L1
L2
L3
Avg
Δ
p
Base
–
34.2
37.4
29.6
33.82
–
–
SFT, full corpus
548k, 2 epochs
33.8
36.5
32.7
34.28
+0.46
0.42
SFT, rank 32
100k, 2 epochs
34.6
38.5
34.8
35.79
+1.97
3.2×10−4
Multimodal RAG
548k index, k=3
35.6
39.3
31.2
35.41
+1.60
5.6×10−4
Table 3: Adaptation of Qwen3-VL-8B on the MCQ track. Δ denotes the change from the base model; p is the p -value from the raw paper-block permutation test. Corrected comparisons appear in Appendix B .
Appendix figures & tables52 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Model ID / version
Source and role
GPT-5.6-sol
gpt-5.6-sol
OpenRouter API; main benchmark
GPT-5.6-luna
gpt-5.6-luna
OpenRouter API; main benchmark
Grok-4.6
grok-4.6
OpenRouter API; main benchmark
Claude-Sonnet-4-5
claude-sonnet-4-5
OpenRouter API; main benchmark
DeepSeek-V4-Flash-Vision-Exp
deepseek-v4-flash-vision-exp
Official DeepSeek API; production AIVC-Judge with image input, temperature 0
GPT-5.6-terra
gpt-5.6-terra
OpenRouter API; auxiliary task-type annotation and distractor generation
Appendix
Table 4: Model identifiers, versions, and access routes. API models are accessed through the listed provider; open-weight checkpoints are run locally from HuggingFace.
Level
Required operation
P-E-D alignment
Bloom levels
n
L1
Infer a value, direction, relation, or outcome from displayed evidence
Predict analogue
2, 3, 4
2,576
L2
Explain an observed result through a biologically grounded mechanism
Explain
4
1,763
L3
Propose a testable hypothesis grounded in the displayed evidence
cancer; drug resistance; therapeutic hypothesis; patient-derived model
Appendix
Table 8: Representative terms used to retrieve AIVC-relevant source records from OmniScience.
P-E-D label
Bloom level
Level
Predict
Explain
Discover
Total
B2
B3
B4
B5
Total
L1
2,362 (91.7%)
214
0
2,576
261
773
1,542
0
2,576
L2
0
1,763 (100%)
0
1,763
0
0
1,763
0
1,763
L3
9
9
1,720 (99.0%)
1,738
0
0
0
1,738
1,738
Appendix
Table 9: Interpretation level by P-E-D label (left) and by Bloom level (right; item counts) in the final benchmark. Parentheses give the within-level percentage of the predominant P-E-D label.
Bloom’s taxonomy
P-E-D function
Level
Category
n
Share
Function
Operational meaning
n
Share
2
Understand
261
4.3%
Predict
Evidence-conditioned inference
2,371
39.0%
3
Apply
773
12.7%
Explain
Mechanistic explanation
1,986
32.7%
4
Analyze
3,305
54.4%
Discover
Testable-hypothesis proposal grounded in displayed evidence
1,720
28.3%
5
Evaluate
1,738
28.6%
Appendix
Table 10: Distribution over Bloom’s taxonomy levels (left) and P-E-D functions (right) in the final benchmark.
Unit
OmniVCTrain
OmniVCBench
Released QA rows
548,450
6,077
Unique source records
96,007
2,625
Unique papers
33,802
1,080
Materialized image paths
327,601
4,428
QA per source record (mean / median)
5.71 / 6
2.32 / 2
QA per paper (mean / median)
16.23 / 12
5.63 / 4
Appendix
Table 11: Unit-level statistics of the released corpora.
Subject
OmniVCTrain
OmniVCBench
Biology
75.5%
73.3%
Medicine
21.7%
26.4%
Physics
0.2%
0.4%
Others (train only)
2.6%
0.0%
Appendix
Table 12: Subject metadata distribution (QA-row weights; source-record weights differ by <1 point in each cell).
Table 14: Publication-year distribution of the 1,080 source papers and the 6,077 benchmark items (Crossref, by DOI). Papers are deduplicated by DOI; QA counts keep the per-paper item weight.
Model
2011
2012
2013
2014
2015
2016
2017
GPT-5.6-sol
51.1
53.3
61.0
57.3
55.8
54.1
53.8
Grok-4.6
49.5
55.3
42.7
59.1
57.9
50.4
49.1
Claude-Sonnet-4-5
49.5
43.9
52.6
44.6
48.2
48.6
50.6
GPT-5.6-luna
47.3
38.0
48.9
38.2
41.7
42.0
44.7
Qwen3-VL-8B-SFT
29.1
32.6
36.6
38.2
30.3
35.8
34.8
Qwen3-VL-8B
29.1
32.2
35.8
35.5
30.8
33.7
34.8
Appendix
Table 15: Item-macro MCQ accuracy (%) by source-paper publication year. Year cells cover 110–2,552 items each (Table 14 ); no model shows a monotone year trend.
Species
Assays
Cell types (top 10)
mouse
49.9
microscopy
61.9
HeLa
7.4
human
38.2
immunoblot
41.8
HEK293T
4.1
rat
5.1
genetic perturbation
37.4
HEK293
3.4
bacteria
3.4
qPCR
36.9
U2OS
2.3
fruit fly
2.5
binding interaction
19.4
fibroblasts
2.3
yeast
2.3
flow cytometry
19.4
MDA-MB-231
2.2
Appendix
Table 16: Coverage of the 1,080 source papers (% of papers carrying each label). Labels are multi-valued and not normalized (variants such as MEF/MEFs and MCF-7/MCF7 co-occur), so entries within a block need not sum to 100%.
Model
L1
L2
L3
Avg
GPT-5.6-sol
37.8
39.1
25.6
34.7
LLaVA-OneVision-7B
23.6
25.3
23.3
24.0
Qwen3-VL-2B
22.0
23.0
19.8
21.7
InternVL3.5-2B
11.0
16.1
15.1
13.7
SmolVLM2-2.2B
11.8
16.1
14.0
13.7
Appendix
Table 17: No-image shortcut test on the stratified 300-item sample: MCQ accuracy (%) when models receive only the question and the six options, without the image or any transcribed evidence. The six-option random baseline is 16.7%. For reference, GPT-5.6-sol scores 57.3% with the image on the same items.
Criterion
Requirement
Strict incorrectness
The screening target is an option contradicted by the evidence or incompatible with the requested answer; an alternative hypothesis requires a substantive scientific distinction.
Plausibility
The error should reflect a biologically or visually plausible model failure.
Relevance and grounding
The option must answer the question and refer only to entities or patterns present in the item.
Diversity
The five distractors should represent distinct error modes rather than paraphrases.
Uniqueness
The screening target is one correct option with semantically distinct distractors; subsequent audits examine violations of this target.
Style consistency
Length, grammar, specificity, and units should not reveal the correct answer.
Appendix
Table 18: Criteria for selecting MDHNM distractors.
Distractor generator
Direct
MDHNM
Δ (pp) [95% CI]
Discordant
McNemar p
Grok-4.6
87.7
57.3
+30.3 [ +23.8 , +36.7 ]
111/20
<10−6
Gemini-3.1-flash-lite
95.3
57.3
+38.0 [ +32.2 , +43.9 ]
120/6
<10−6
Kimi-K2.6
93.3
57.3
+36.0 [ +30.1 , +41.9 ]
116/8
<10−6
Appendix
Table 19: Direct-generation distractor ablation on the stratified 300-item sample: MCQ accuracy (%) of the fixed solver (GPT-5.6-sol, temperature 0) on items whose distractors were written directly by each baseline generator, versus the released MDHNM items on the same questions. Discordant pairs are (correct only on direct) / (correct only on MDHNM); p is the exact two-sided McNemar test, and bracketed intervals are source-paper cluster bootstrap 95% CIs (20,000 replicates).
Level
Distractor generator
Direct
MDHNM
Δ (pp)
L1
Grok-4.6
86.6
63.8
+22.8
L1
Gemini-3.1-flash-lite
91.3
63.8
+27.6
L1
Kimi-K2.6
92.9
63.8
+29.1
L2
Grok-4.6
86.2
65.5
+20.7
L2
Gemini-3.1-flash-lite
97.7
65.5
+32.2
L2
Kimi-K2.6
92.0
65.5
+26.4
Appendix
Table 20: Per-level breakdown of the direct-generation ablation (127 L1, 87 L2, 86 L3 items). The MDHNM column is the fixed solver’s accuracy on the released items; it repeats within each level block because it does not depend on the baseline generator.
Distractor arm
GPT-5.6-sol
Grok-4.6
gemini-3.1-flash-lite direct
93.7
94.0
gemini-3.1-flash-lite mined
90.0
91.3
gemini-3.7-flash-high direct
94.3
93.7
gpt-5.6-terra direct
92.0
91.7
original MDHNM
57.3
48.3
Appendix
Table 21: Same-evidence distractor ablation on the stratified 300-item sample: MCQ accuracy (%) of two solvers on five distractor arms. Both generator arms receive identical evidence and budgets; mined arms are conditioned on mined inference errors, direct arms are not. original MDHNM denotes the released items. For the paired gemini-3.1-flash-lite arms, a positive direct-minus-mined difference means the direct distractors are easier.
Table 22: Level-conditioned dimensions used by AIVC-Judge. The recorded L3 dimensions assess source-referenced critique and argumentation.
Component
Reported configuration
Judge
DeepSeek-V4-Flash-Vision-Exp ( deepseek-v4-flash-vision-exp ); single pointwise judge; temperature 0
Response input
question and candidate response
Closed evidence
figure caption, surrounding paper context, and complete reference answer
Visual input
original figure image set, as shown to the answering model
Reference key points
disabled; the complete reference answer provides the scoring anchor
Anonymization
evaluated model identity omitted from the judge input
Appendix
Table 23: Production AIVC-Judge configuration used for the reported open-response scores.
Annotator
Correct/assigned
Overall [95% CI]
L1
L2
L3
Human 1
166/300
55.3 [50.0, 60.7]
55.9
57.5
52.3
Human 2
214/300
71.3 [65.8, 76.7]
76.4
81.6
53.5
Human 3
224/300
74.7 [69.1, 79.5]
75.6
87.4
60.5
Appendix
Table 24: Human MCQ accuracy (%) on the fixed 300-question sample. Level denominators are L1/L2/L3 = 127/87/86. Human 1’s 17 missing answers count as incorrect. Brackets denote 95% source-paper cluster bootstrap confidence intervals.
Annotator pair
Cohen’s κ [95% CI]
Exact agreement
Human 1–Human 2
0.456 [0.391, 0.520]
54.8%
Human 1–Human 3
0.457 [0.391, 0.523]
54.8%
Human 2–Human 3
0.855 [0.805, 0.897]
88.0%
Appendix
Table 25: Human–human agreement on selected MCQ options. Each comparison uses the same 283 complete questions. Brackets denote 95% source-paper cluster bootstrap confidence intervals.
Scorer
GPT-5.6-sol
Qwen3-VL-8B
LLaVA-Med-Mistral-7B
Human 1
3.163
2.186
1.727
Human 2
4.127
2.657
1.917
Human 3
3.933
2.447
1.797
Production judge
3.354
2.257
1.790
Appendix
Table 26: Observed mean raw overall scores for the three response candidates.
Scope
Paired cells
Spearman [95% CI]
MAE
Bias
All, raw overall
884
0.877 [0.855, 0.897]
0.464
+0.188
L1
377
0.909 [0.887, 0.926]
0.389
+0.193
L2
257
0.875 [0.837, 0.905]
0.450
+0.014
L3
250
0.709 [0.637, 0.774]
0.592
+0.360
GPT-5.6-sol
291
0.840 [0.797, 0.874]
0.608
+0.379
Qwen3-VL-8B
296
0.865 [0.823, 0.896]
0.430
+0.169
Appendix
Table 27: Production-judge agreement with the mean of the three human ratings. The unit is a candidate–question response cell. Rows use raw overall scores except the final weighted-score sensitivity. Bias is human minus judge; brackets denote 95% source-paper cluster bootstrap confidence intervals.
Scorer pair
Paired cells
Spearman [95% CI]
QWK [95% CI]
Human 1–Human 2
887
0.820 [0.788, 0.849]
0.738 [0.694, 0.776]
Human 1–Human 3
887
0.835 [0.807, 0.859]
0.800 [0.766, 0.829]
Human 2–Human 3
900
0.953 [0.942, 0.962]
0.944 [0.930, 0.956]
Human 1–judge
884
0.838 [0.810, 0.863]
0.851 [0.821, 0.876]
Human 2–judge
897
0.832 [0.800, 0.860]
0.779 [0.740, 0.815]
Human 3–judge
897
0.854 [0.830, 0.876]
0.837 [0.806, 0.863]
Appendix
Table 28: Pairwise agreement on raw overall scores, using each pair’s available response cells. QWK denotes quadratic-weighted Cohen’s κ on integer 1–5 ratings. Brackets denote 95% source-paper cluster bootstrap confidence intervals.
Check
Cross-set candidates
Confirmed duplicates
Normalized DOI / paper ID (exact)
0
0
Normalized title (exact)
0
0
Title char-TFIDF cosine ≥.90
0
0
Title char-TFIDF cosine ≥.95
0
0
Appendix
Table 29: Source-level cross-set overlap checks between the released OmniVCTrain (33,802 papers) and OmniVCBench (1,080 papers).
Field
Train uniques
Bench uniques
Exact pairs
Near pairs
Question
548,390
6,077
0
0
Caption
96,007
2,625
0
0
Answer (exact only)
544,652
6,077
0
–
Appendix
Table 30: Text-level cross-set overlap between the released corpora. Near pairs require Jaccard ≥.80 or containment ≥.90 under bottom-12 signature retrieval.
Perturbation
Retrieval
Acceptance
Question: append 10% words
100%
100%
Question: delete one word
100%
98.2%
Question: substitute one word
100%
87.8%
Caption: append 10% words
100%
100%
Caption: delete one word
100%
100%
Caption: substitute one word
100%
100%
Appendix
Table 31: Synthetic-text sensitivity of the overlap detector ( n=500 perturbations per row): signature-retrieval recall and acceptance recall (Jaccard ≥.80 or containment ≥.90 ).
Table 33: Synthetic-image sensitivity of the perceptual-hash detector ( n=300 perturbations per row).
Model
Item-level [95% CI]
Paper-weighted [95% CI]
MCQ accuracy (%)
GPT-5.6-sol
55.093 [53.768, 56.444]
56.033 [54.267, 57.794]
Grok-4.6
50.304 [48.862, 51.729]
51.492 [49.738, 53.286]
Claude-Sonnet-4-5
49.366 [47.937, 50.768]
50.123 [48.313, 51.889]
GPT-5.6-luna
43.739 [42.415, 45.135]
44.478 [42.678, 46.314]
Qwen3-VL-8B-SFT
34.277 [33.018, 35.566]
34.111 [32.444, 35.821]
Appendix
Table 34: Item-level and paper-weighted scores with source-paper cluster bootstrap 95% CIs, computed from unrounded item-level files (6,077 items, 1,080 papers). Top block: MCQ accuracy (%), sixteen configurations; bottom block: AIVC-Judge (1–5), eleven models.
Scope
Pearson
Spearman
Published full-track, item
0.938 ( p=6×10−5 )
0.964 ( p=2×10−5 )
Published full-track, paper
0.945
0.955
Model-specific matched, item
0.936
0.964
Model-specific matched, paper
0.943
0.955
Strict common, item
0.935 [0.919, 0.946] ( p=8×10−5 )
0.964 [0.927, 0.973] ( p=3×10−5 )
Strict common, paper
0.943 [0.923, 0.956]
0.955 [0.927, 0.973]
Appendix
Table 35: MCQ–open-response model-level correlation ( n=11 models) by item set and weighting. Bracketed values are source-paper cluster bootstrap 95% CIs; p values are permutation-based.
Excluded family
n
Pearson
Spearman
OpenAI GPT
9
0.956
0.967
xAI Grok
10
0.932
0.964
Anthropic Claude
10
0.971
0.964
Qwen3-VL (incl. SFT)
8
0.942
0.929
Qwen2.5-VL
10
0.939
0.939
InternVL
10
0.943
0.952
Appendix
Table 36: Leave-one-family-out MCQ–open-response correlation (strict common items, paper-weighted; n = remaining models).
Setting
Data
Epochs
Rank
Avg
Δ
p
Full corpus
548k
2
16
34.28
+0.46
0.42
Rank 32, 100k subset
100k
2
32
35.79
+1.97
3.2×10−4
Appendix
Table 37: Principal SFT configurations for Qwen3-VL-8B. Raw paper-block permutation tests compare each setting with the 33.82% base model on all 6,077 items; Holm-corrected values over the original adaptation comparison family are quoted in the main text.
Retrieval signal
L1
L2
L3
Avg
Δ
p
Joint image–question
35.6
39.3
31.2
35.41
+1.60
5.6×10−4
Image only
34.1
37.5
28.9
33.62
−0.20
.64
Text only
34.1
37.3
28.5
33.40
−0.41
.31
Appendix
Table 38: Retrieval-signal ablation for Qwen3-VL-8B with k=3 . The base accuracy is 33.82%. p : raw paper-block permutation test vs. base.
Model
Base
+ RAG
Δ
p
Qwen3-VL-8B
33.82
35.41
+1.60
5.6×10−4
Qwen3-VL-2B
21.69
23.19
+1.50
2.9×10−3
InternVL3.5-8B
30.28
31.04
+0.76
.10
Appendix
Table 39: Joint multimodal RAG across base models. Improvements are positive across all three models but vary in statistical strength; the InternVL3.5-8B change does not reach significance. p : raw paper-block permutation test vs. the corresponding base model.
Pilot
Simulator / data
Prediction target
Evidence choice
Cases / readouts
Gene interactions
GEARS ( Roohani et al., 2023 ) ; Norman K562 CRISPRa ( Norman et al., 2019 )
non-additive double-gene response AB−A−B per program
high-state fraction change under each intervention
one non-target readout in 96 reserved cells
5 / 30
Appendix
Table 40: The four evidence-acquisition pilots. Each case admits exactly one additional measurement; readouts are the scored prediction fields. The response adapter is the calibration layer updated by the acquired measurement.
Figure 5: Shared pilot structure. Each pilot runs predict, explain and select, measure, then update and test. Three pilots revise predictions through a response adapter after active evidence acquisition; the microscopy pilot revises its interpretation. These are one-step replay loops: simulator weights and free-text biological mechanisms are not validated or updated.
Pilot
Readouts
Initial
Self-review
After evidence
Corrections
Gene interactions
48
77.1
77.1
89.6
+6 / −0
Drug combinations
30
93.3
80.0
100.0
+2 / −0
Real microscopy
9
66.7
66.7
66.7
0 / −0
Signaling
30
66.7
60.0
70.0
+1 / −0
Appendix
Table 41: Per-pilot accuracy (%) before and after acquiring one chosen measurement. Corrections count wrong-to-right transitions against zero right-to-wrong transitions in every pilot.
Figure 6: Per-pilot outcomes. Correct-outcome rate under the initial prediction, self-review with the same evidence, and revision after one chosen measurement, for each of the four pilots. Case counts and readout counts differ across pilots; results are descriptive and support no pooled accuracy or causal-discovery claim.
Figure 7: Real microscopy evidence. Reference compounds of the microscopy pilot (BBBC021), shown as DNA, tubulin, and actin channels; test compounds are disjoint from these references. The task reads real fluorescence pixels, not synthetic bar plots.
L1
L2
L3
Model
C
ES
F
CC
MG
EC
J
A
EW
GPT-5.6-sol
3.35
3.35
3.88
3.72
3.63
3.85
4.04
1.92
1.67
GPT-5.6-luna
3.23
3.23
3.70
3.49
3.46
3.66
3.94
1.92
1.67
Grok-4.6
3.23
3.25
3.64
3.38
3.36
3.62
3.87
1.71
1.53
Claude-Sonnet-4-5
2.52
2.55
3.07
2.99
2.97
3.10
3.64
1.99
1.73
Qwen3-VL-8B-SFT †
2.50
2.55
2.71
2.25
2.34
2.50
3.01
1.44
1.33
Appendix
Table 42: AIVC-Judge dimension scores (1–5; higher is better) by interpretation level. L1: C = correctness, ES = evidence support. L2: F = faithfulness, CC = causal completeness, MG = mechanistic granularity, EC = evidence consistency. L3: J = judgment correctness, A = argumentation quality, EW = evidence weighing. † denotes SFT on OmniVCTrain .
Model
Single
Multi
Δ (pp)
GPT-5.6-sol
54.5
55.7
+1.2
Claude-Sonnet-4-5
45.4
53.2
+7.8
Qwen3-VL-8B
33.6
34.0
+0.4
InternVL3.5-8B
29.6
30.9
+1.3
LLaVA-OneVision-7B
25.0
26.1
+1.1
Appendix
Table 43: MCQ accuracy (%) on single- versus multi-subfigure items (five representative models of the eleven with open-response coverage).
Figure 8: Case A exhibit (L1 / Predict). Source figure, item, reference answer, and six MCQ options; the keyed option (C) is highlighted in green.
Figure 9: Case B exhibit (L1 / Predict). Source figure, item, reference answer, and six MCQ options; the keyed option (B) is highlighted in green.
Figure 10: Case C exhibit (L1 / Predict). Source figure, item, reference answer, and six MCQ options; the keyed option (E) is highlighted in green.
Figure 11: Case D exhibit (L2 / Explain). Source figure, item, reference answer, and six MCQ options; the keyed option (F) is highlighted in green.
Figure 12: Case E exhibit (L2 / Explain). Source figure, item, reference answer, and six MCQ options; the keyed option (A) is highlighted in green.
Figure 13: Case F exhibit (L2 / Explain). Source figure, item, reference answer, and six MCQ options; the keyed option (C) is highlighted in green.
Figure 14: Case G exhibit (L3 / Discover). Source figure, item, reference answer, and six MCQ options; the keyed option (D) is highlighted in green.
Figure 15: Case H exhibit (L3 / Discover). Source figure, item, reference answer, and six MCQ options; the keyed option (B) is highlighted in green. Option F’s knockdown prediction contradicts its own mechanism claim, as discussed below.
Figure 16: Case I exhibit (L3 / Discover). Source figure, item, reference answer, and six MCQ options; the keyed option (D) is highlighted in green. Option B contradicts the integrin → Lck initiation shown in panel h, as discussed below.
Systematic ablations are essential to attribute performance gains in AI Virtual Cells, yet they are rarely performed because biological repositories are under-standardized and tightly coupled to domain-specific data and formats. While recent coding agents can translate ideas into implementations, they typically stop at producing code and lack a verifier that can reproduce strong baselines and rigorously test which components truly matter. We introduce AblateCell, a reproduce-then-ablate agent for virtual cell repositories that closes this verification gap. AblateCell first reproduces reported baselines end-to-end by auto-configuring environments, resolving dependency and data issues, and rerunning official evaluations while emitting verifiable artifacts. It then conducts closed-loop ablation by generating a graph of isolated repository mutations and adaptively selecting experiments under a reward that trades off performance impact and execution cost. Evaluated on three single-cell perturbation prediction repositories (CPA, GEARS, BioLORD), AblateCell achieves 88.9% (+29.9% to human expert) end-to-end workflow success and 93.3% (+53.3% to heuristic) accuracy in recovering ground-truth critical components. These results enable scalable, repository-grounded verification and attribution directly on biological codebases.
Xue Xia, Chengkai Yao, Mingyu Tsoi +10
Shanghai Artificial Intelligence Laboratory · The Hong Kong University of Science and Technology (Guangzhou) · University of California San Diego +4
AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.
Mengran Li, Bo Li, Chengyang Zhang +3
Sun Yat-sen University · University of Macau · Sichuan University +3
Cellular perturbation-response modeling requires coordinated choices of representations, fusion mechanisms, objectives, and training procedures. Large language models (LLMs) can propose executable candidates, but unconstrained revision can produce invalid implementations, change task semantics, or discard useful components. We present CellScientist, a protocol-constrained workflow that converts execution and validation feedback into auditable model-revision trajectories. It records design states and outcomes, routes discrepancies to specific components, and applies local revisions under a fixed task contract. A matched-budget study fixes the candidate language, predictor, fitting, and evaluator: structured revision finds better held-out predictors at small budgets under two LLM backbones. Operational audits link history to fewer repeated proposals, contract checks to contained violations, and discrepancy routing to targeted repairs. Refits of two frozen designs on an independently acquired cohort evaluate external predictive utility. Open-workflow trajectories retain improvements, regressions, and failures, while transcriptomic and single-cell searches extend application to additional response spaces. CellScientist produces both a selected predictor and an inspectable record of its development. Project page: https://limengran98.github.io/CellScientist/.
Mengran Li, Bo Li, Jiaying Wang +7
Center for Artificial Intelligence and Robotics (CAIR), Hong Kong Institute of Science and Innovation (HKISI)