From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles
Organizations: Department of Computer Science and Engineering, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Department of Biostatistics, Yale University · College of Pharmacy, The Ohio State University
Abstract
Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.
Figures & tables
| Dataset | # Reference | # Test | # MOA classes | Avg Sources/Ref. | Avg Labels |
|---|---|---|---|---|---|
| MoA-Novel | 63 | 100 | 26 (Known) | 4.05 | 1.00 |
| MoA-Verified | 63 | 200 | 26 | 4.05 | 1.19 |
| MoA-Extended | 533 | 600 | 137 | 2.64 | 1.13 |
| Dataset | MoA-Verified | MoA-Extended | MoA-Novel | ||||
|---|---|---|---|---|---|---|---|
| Metrics | Acc | Any-match | Acc | Any-match | Acc | F1 (Novel) | |
| kNN | Random | N/A | |||||
| CellProfiler | |||||||
| CellCLIP | |||||||
| ECFP | |||||||
| GPT 5.1 | Structure-only | ||||||
| Dataset | Verified | Extended | |
|---|---|---|---|
| Method | Acc | Acc | |
| GPT 5.1 | Base LLM | ||
| PhenoAIR-Base | |||
| + Memo. | |||
| + Memo. & Cal. | |||
| Claude 4 | Base LLM |
| Setting | Phenotype | Structure side | Acc. (%) |
|---|---|---|---|
| Full PhenoAIR | matched | matched | 80.0 1.8 |
| Shuffled structure side | matched | shuffled | 63.2 0.7 |
| Shuffled phenotype | shuffled | matched | 53.6 1.0 |
| Structure-only agent | removed | matched | 50.2 0.7 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Overall Acc. | Precision (Novel) | Recall (Novel) | F1 (Novel) | |
|---|---|---|---|---|---|
| GPT 5.1 | Structure-only | ||||
| Workflow | |||||
| Single Agent | |||||
| PhenoAIR | |||||
| w/o Self-evolving | |||||
| w/o Calibration |
| Dataset | MoA-Verified | MoA-Extended | |||
|---|---|---|---|---|---|
| Method | Acc | Any-match | Acc | Any-match | |
| GPT 5.1 | Base LLM | ||||
| PhenoAIR-Base | |||||
| + Memo. | |||||
| + Memo. & Cal. | |||||
| Claude 4 | Base LLM | ||||
| Aggregation | Model | Top-1 Acc | Top-5 Acc | F1 | Any Overlap |
|---|---|---|---|---|---|
| Global | CellProfiler Raw | 7.5 | 30.0 | 10.3 | 18.5 |
| CellProfiler Normalized | 9.5 | 29.5 | 10.2 | 23.5 | |
| CA-MAE Raw | 7.5 | 28.0 | 9.2 | 19.5 | |
| CA-MAE Normalized | 13.0 | 31.5 | 10.6 | 23.5 | |
| CellCLIP Raw | 14.5 | 37.5 | 12.8 | 24.0 | |
| CellCLIP Normalized | 14.0 | 35.5 | 11.7 | 26.5 |
| Aggregation | Model | Top-1 Acc | Top-5 Acc | F1 | Any Overlap |
|---|---|---|---|---|---|
| Global | CellProfiler Raw | 3.3 | 10.2 | 3.7 | 7.5 |
| CellProfiler Normalized | 6.8 | 16.5 | 6.3 | 12.7 | |
| CA-MAE Raw | 3.5 | 13.7 | 4.9 | 10.3 | |
| CA-MAE Normalized | 6.5 | 14.7 | 5.4 | 12.5 | |
| CellCLIP Raw | 5.5 | 16.2 | 5.7 | 11.0 | |
| CellCLIP Normalized | 10.2 | 19.7 | 7.1 | 16.0 |
| Dataset | MoA-Verified | MoA-Extended | MoA-Novel | |||
|---|---|---|---|---|---|---|
| Metrics | Acc | Any-match | Acc | Any-match | Acc | F1 (Novel) |
| CellCLIP | ||||||
| CA-MAE | ||||||
| CellProfiler | ||||||
| Method | LLM calls/query | Input tokens/query | Output tokens/query | Wall time/query |
|---|---|---|---|---|
| Base LLM | 1 | 0.3k | 180 | 3.0s |
| Structure-only | 1 | 0.7k | 160 | 3.7s |
| Workflow | 3 | 11k | 800 | 12.8s |
| Singe Agent | 10 | 65k | 1.6k | 31s |
| PhenoAIR | 14 | 170k | 8k | 80s |
| Round | Diagnosed pattern | Proposed / retained update | Evolution decision |
|---|---|---|---|
| 0 | Novel cases missed because the ground-truth MOA was never surfaced; verified cases also showed false novel overcalls and P/M reinforcement around wrong known labels. | Generated an early STOP rule for low-reliability phenotype cases with a strong mechanistic alternative; also proposed a canonical-survivor Arbiter fallback. | Initial proposal accepted for testing. |
| 1 | The STOP rule over-triggered and caused success-to-failure regressions in both core and novel buckets. | Added phenotype expansion rules for rejected canonical candidates and medium-reliability P/M-supported wrong labels. | Disabled the harmful STOP rule and kept the Arbiter fallback temporarily. |
| 2 | Expansion alone did not resolve true novel cases because the Arbiter still blocked novel outputs when canonical candidates remained visible. | Introduced final-state novel diagnostics: no supported canonical candidate, top-3 all rejected, and cumulative rejection rate. | Disabled the canonical-survivor fallback and kept the useful medium-reliability expansion rule. |
| 3–4 | Cross-round analysis confirmed that correct novel calls were blocked by missing final-state diagnostics and by overly conservative canonical fallback. | Retained final-state novel signals and disabled rules/prompts that prematurely forced canonical outputs or stopped refinement. | Final policy separates weak canonical visibility from genuine canonical support. |
| Evolved principle | Failure mode addressed | Operational signal / action |
|---|---|---|
| Avoid premature STOP | A STOP rule intended for unresolved novel cases caused regressions by halting recoverable core and novel success cases. | Disable over-triggered STOP rules; prefer additional phenotype inspection when evidence remains ambiguous. |
| Separate canonical visibility from canonical support | The Arbiter fallback treated any visible canonical candidate as enough to suppress novel outputs. | Allow novel only using terminal diagnostics rather than candidate visibility alone. |
| Expose final-state novel diagnostics | Per-round signals were insufficient because true novel patterns only became clear after full refinement. | Use final-state signals such as no supported canonical candidate, top-3 all rejected, and high cumulative rejection rate. |