Organizations: Department of Computer Science and Engineering, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Department of Biostatistics, Yale University · College of Pharmacy, The Ohio State University
Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.
Figures & tables
Figure 1: Architecture of PhenoAIR. The system consists of three stages: Externalized reliability calibration (left) precomputes a reliability-annotated phenotype neighborhood and a mechanistic family ontology from the multi-source reference database. Candidate-centric evidence reasoning (center) maintains a shared candidate memory updated by Phenotype Agent (calibrated retrieval) and a Mechanism Agent (structural and pharmacological verdicts). Controller-guided refinement (right) iteratively coordinates the two agents through a rule-based controller, terminating with an Arbiter that produces a canonical or novel MOA. Self-evolving extensions (in blue) handle open-set inference.
Dataset
# Reference
# Test
# MOA classes
Avg Sources/Ref.
Avg Labels
MoA-Novel
63
100
26 (Known)
4.05
1.00
MoA-Verified
63
200
26
4.05
1.19
MoA-Extended
533
600
137
2.64
1.13
Table 1: Benchmark statistics. Our benchmark comprises three settings constructed from the JUMP Cell Painting Consortium and Drug Repurposing Hub.
Dataset
MoA-Verified
MoA-Extended
MoA-Novel
Metrics
Acc
Any-match
Acc
Any-match
Acc
F1 (Novel)
kNN
Random
3.5±1.4
12.3±2.0
1.2±0.3
3.2±0.5
N/A
CellProfiler
9.5±0.0
23.5±0.0
6.8±0.0
12.7±0.0
CellCLIP
14.0±0.0
26.5±0.0
10.2±0.0
16.0±0.0
ECFP
23.0±0.0
41.0±0.0
30.0±0.0
42.3±0.0
GPT 5.1
Structure-only
43.5±1.4
59.0±1.1
39.8±0.3
52.9±0.1
44.2±3.4
53.3±4.0
Table 2: Main results on three datasets. We report accuracy (Acc. %), hit rate (Any-match), and F1 measured against ground-truth MOA labels. All LLM-based methods operate on the same CellCLIP features. Best results are in bold; second-best are underlined. Our method is highlighted in blue.
Dataset
Verified
Extended
Method
Acc
Acc
GPT 5.1
Base LLM
16.4
18.7
PhenoAIR-Base
55.6
43.9
+ Memo.
60.1
44.5
+ Memo. & Cal.
62.6
45.6
Claude 4
Base LLM
26.7
23.3
Table 3: Effect of core components on two datasets. Variants are constructed by incrementally adding each PhenoAIR component to a base LLM, isolating the contribution of reliability calibration, candidate-centric memory with multi-agent reasoning. All variants use the same base LLM and CellCLIP features. Full results are in Appendix.
Figure 2: Self-evolving refinement curves across iterations. Controller rules and arbiter prompts are updated using development-set trajectories after the baseline step. Left two: development set and test set overall accuracy on MoA-Novel. Right: MoA-Extended accuracy on development and test sets.
Setting
Phenotype
Structure side
Acc. (%)
Full PhenoAIR
matched
matched
80.0 ± 1.8
Shuffled structure side
matched
shuffled
63.2 ± 0.7
Shuffled phenotype
shuffled
matched
53.6 ± 1.0
Structure-only agent
removed
matched
50.2 ± 0.7
Table 4: Evidence-shuffling control on MoA-Extended (GPT-5.1, mean ± s.d. over five runs).
Figure 3: (Left) MoA Recall@10 with and without calibration, and accuracy comparison on MoA-Verified, both shown across CellProfiler and CellCLIP features. (Right) Failure mode distribution on MoA-Extended. Inner ring: high-level categories; outer ring: subtype breakdown.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Overall Acc.
Precision (Novel)
Recall (Novel)
F1 (Novel)
GPT 5.1
Structure-only
44.2±3.4
61.5±3.6
47.2±5.0
53.3±4.0
Workflow
58.0±0.3
72.4±2.5
53.2±1.8
61.3±0.8
Single Agent
48.8±2.4
68.8±2.0
66.0±1.8
67.3±1.1
PhenoAIR
79.4±2.2
99.0±1.5
59.8±3.5
74.7±2.9
w/o Self-evolving
66.0±2.0
96.0±3.0
48.0±4.0
64.0±3.0
w/o Calibration
75.5±2.5
93.5±3.0
58.0±4.0
71.5±3.5
Appendix
Table 5: Full results on the MoA-Novel benchmark across two backbone LLMs (GPT-5.1 and Claude 4). We report Overall Accuracy, and Precision, Recall, and F1 of the novel-MOA prediction. Two ablations of PhenoAIR are included: w/o Self-evolving (no offline rule refinement) and w/o Calibration (no externalized reliability calibration). Results are mean ± std over multiple runs with different random seeds.
Dataset
MoA-Verified
MoA-Extended
Method
Acc
Any-match
Acc
Any-match
GPT 5.1
Base LLM
16.4
23.6
18.7
25.3
PhenoAIR-Base
55.6
66.1
43.9
54.4
+ Memo.
60.1
73.0
44.5
55.1
+ Memo. & Cal.
62.6
74.6
45.6
56.2
Claude 4
Base LLM
26.7
31.0
23.3
30.2
Appendix
Table 6: Component ablation of PhenoAIR on MoA-Verified and MoA-Extended across two backbone LLMs. Variants are constructed incrementally: Base LLM predicts MOA from compound information alone; PhenoAIR-Base adds Cell Painting retrieval and multi-agent reasoning; +Memo. adds the candidate-centric evidence memory; +Memo. & Cal. additionally includes the externalized reliability calibration layer. We report Accuracy (Acc) and Any-match (Hit Rate).
Aggregation
Model
Top-1 Acc
Top-5 Acc
F1
Any Overlap
Global
CellProfiler Raw
7.5
30.0
10.3
18.5
CellProfiler Normalized
9.5
29.5
10.2
23.5
CA-MAE Raw
7.5
28.0
9.2
19.5
CA-MAE Normalized
13.0
31.5
10.6
23.5
CellCLIP Raw
14.5
37.5
12.8
24.0
CellCLIP Normalized
14.0
35.5
11.7
26.5
Appendix
Table 7: Retrieval performance on the MoA-Verified benchmark under different feature representations, preprocessing strategies, and cross-source aggregation schemes. Raw features denote the original representations, while normalized features apply batch-effect correction using plate-level normalization and control-based whitening. Global aggregation constructs a single compound representation across sources, whereas per-source voting performs source-specific retrieval followed by cross-source score aggregation.
Aggregation
Model
Top-1 Acc
Top-5 Acc
F1
Any Overlap
Global
CellProfiler Raw
3.3
10.2
3.7
7.5
CellProfiler Normalized
6.8
16.5
6.3
12.7
CA-MAE Raw
3.5
13.7
4.9
10.3
CA-MAE Normalized
6.5
14.7
5.4
12.5
CellCLIP Raw
5.5
16.2
5.7
11.0
CellCLIP Normalized
10.2
19.7
7.1
16.0
Appendix
Table 8: Retrieval performance on the MoA-Extended benchmark under different feature representations, preprocessing strategies, and cross-source aggregation schemes.
Dataset
MoA-Verified
MoA-Extended
MoA-Novel
Metrics
Acc
Any-match
Acc
Any-match
Acc
F1 (Novel)
CellCLIP
62.6±0.7
74.6±0.6
45.6±0.2
56.2±0.3
79.4±2.2
74.7±2.9
CA-MAE
62.0±0.5
74.5±0.4
42.5±0.3
52.8±0.3
77.0±1.5
74.4±1.9
CellProfiler
60.0±0.6
74.1±0.5
42.2±0.3
53.9±0.3
76.6±1.8
72.5±2.5
Appendix
Table 9: Performance of PhenoAIR across different Cell Painting feature representations.
Method
LLM calls/query
Input tokens/query
Output tokens/query
Wall time/query
Base LLM
1
0.3k
180
3.0s
Structure-only
1
0.7k
160
3.7s
Workflow
3
11k
800
12.8s
Singe Agent
10
65k
1.6k
31s
PhenoAIR
14
170k
8k
80s
Appendix
Table 10: Per-query inference cost across methods. We report the average number of LLM calls and the cumulative input tokens, output tokens, and wall time across all calls per query. All LLM-based methods use the same base model and identical hardware configurations.
Round
Diagnosed pattern
Proposed / retained update
Evolution decision
0
Novel cases missed because the ground-truth MOA was never surfaced; verified cases also showed false novel overcalls and P/M reinforcement around wrong known labels.
Generated an early STOP rule for low-reliability phenotype cases with a strong mechanistic alternative; also proposed a canonical-survivor Arbiter fallback.
Initial proposal accepted for testing.
1
The STOP rule over-triggered and caused success-to-failure regressions in both core and novel buckets.
Added phenotype expansion rules for rejected canonical candidates and medium-reliability P/M-supported wrong labels.
Disabled the harmful STOP rule and kept the Arbiter fallback temporarily.
2
Expansion alone did not resolve true novel cases because the Arbiter still blocked novel outputs when canonical candidates remained visible.
Introduced final-state novel diagnostics: no supported canonical candidate, top-3 all rejected, and cumulative rejection rate.
Disabled the canonical-survivor fallback and kept the useful medium-reliability expansion rule.
3–4
Cross-round analysis confirmed that correct novel calls were blocked by missing final-state diagnostics and by overly conservative canonical fallback.
Retained final-state novel signals and disabled rules/prompts that prematurely forced canonical outputs or stopped refinement.
Final policy separates weak canonical visibility from genuine canonical support.
Appendix
Table 11: Self-evolving trace on the MoA-Novel benchmark. The evolving scaffold does not simply add rules monotonically; it proposes, evaluates, disables, and refines controller and Arbiter heuristics based on cross-round trajectory changes.
Evolved principle
Failure mode addressed
Operational signal / action
Avoid premature STOP
A STOP rule intended for unresolved novel cases caused regressions by halting recoverable core and novel success cases.
Cell Painting combines multiplexed fluorescent staining, high-content imaging, and quantitative analysis to generate high-dimensional phenotypic readouts to support diverse downstream tasks such as mechanism-of-action (MoA) inference, toxicity prediction, and construction of drug-disease atlases. However, existing workflows are slow, costly and difficult to interpret. Approaches for drug screening modeling predominantly focus on molecular representation learning, while neglecting actual experimental context (e.g., cell line, dosing schedule, etc.), limiting generalization and MoA resolution. We introduce CP-Agent, an agentic multimodal large language model (MLLM) capable of generating mechanism-relevant, human-interpretable rationales for cell morphological changes under drug perturbations. At its core, CP-Agent leverages a context-aware alignment module, CP-CLIP, that jointly embeds high-content images and experimental metadata to enable robust treatment and MoA discrimination (achieving a maximum F1-score of 0.896). By integrating CP-CLIP outputs with agentic tool usage and reasoning, CP-Agent compiles rationales into a structured report to guide experimental design and hypothesis refinement. These capabilities highlight CP-Agent's potential to accelerate drug discovery by enabling more interpretable, scalable, and context-aware phenotypic screening -- streamlining iterative cycles of hypothesis generation in drug discovery.
Yuxin Zhang, Yiyao Li, Ping Shu Ho +3
Department of Electrical and Computer Engineering, The University of Hong Kong · School of Computing and Data Science, The University of Hong Kong · Nvidia AI Technology Center +2
High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into biological narratives, yet their scientific use requires quantitative auditing. We present an evaluation-first, retrieval-augmented interpretation framework for longitudinal Cell Painting morphology, applied to a 9-week RPE-1 time course across five dose rates (0.003--6.0 mGy/hr). Week-matched treated-control morphology deltas are combined with retrieved perturbation neighbors, pathway context, and literature evidence through stable evidence identifiers, enabling an LLM to generate structured, evidence-linked hypotheses that are hierarchically summarized while preserving provenance. We introduce two quantitative auditing tests: V1 citation validity, which verifies that cited evidence identifiers exist in the prompt, and V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features. In our experiments, V1 detected no invalid evidence references, while V2 showed meaningful morphology compatibility that increased with perturbation strength and was positively associated with an independent morphology drift summary. The framework produces auditable, falsifiable biological hypotheses, including an adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates (0.003--0.3 mGy/hr). Current limitations include proxy-based evaluation and the lack of ground-truth mechanism labels.
Gilchan Park, Guang Zhao, Byung-Jun Yoon +1
Brookhaven National Laboratory Upton, New York, USA
AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.
Mengran Li, Bo Li, Chengyang Zhang +3
Sun Yat-sen University · University of Macau · Sichuan University +3