Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
Authors: Laura van Weesep, Riccardo Tedoldi, Jens Sjölund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, +1 more
Organizations: Molecular AI, Discovery Sciences R&D, AstraZeneca · Department of Information Technology, Uppsala University · Robotics, Perception & Learning, KTH Royal Institute of Technology · Science for Life Laboratory, Stockholm, Sweden · Drug Metabolism and Pharmacokinetics, Research and Early Development, Cardiovascular, Renal and Metabolism (CVRM), BioPharmaceuticals R&D, AstraZeneca · Department of Computer Science and Engineering, Chalmers University of Technology and University of Gothenburg, Sweden
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's ∼2 million bioassays is critically sparse, 36% lacking an assay format, 89% a BioAssay type, and >99.9% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.
Figures & tables
Disagreement pattern
Count
%
Sent for review
LLMs support the BARD label against ChEMBL
17
40%
4
ChEMBL and BARD agree and LLMs disagree with both
10
23%
10
ChEMBL and BARD disagree and LLMs disagree with both (novel prediction)
2
5%
2
No BARD label available; ChEMBL-only ground truth
14
33%
1
Total
43
100%
17
Table 2: Categorization of the 43 ChEMBL assays where the majority of LLMs disagreed with the ChEMBL BAO assay format label, and how many of each pattern were sent for senior industrial curator review (Table 13 ).
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Non-empty
% of Total
Unique Values
% Unique (of non-empty)
AID
1 994 310
100.0%
1 994 310
100.00%
Comment
1 769 245
88.7%
647 580
36.60%
Description
1 770 567
88.8%
173 040
9.77%
Name
1 770 568
88.8%
1 559 216
88.06%
Deposit Date
1 910 340
95.8%
1 979
0.10%
Assay Format
1 278 238
64.1%
3
0.00%
Appendix
Table 5 : Assay coverage and uniqueness for description, comment, and name fields (Total AIDs = 1 994 310)
Assay Format
With Comment
Total
Coverage
Percentage
Organism-based
609 321
663 381
91.85%
52.96%
Cell-based
541 113
614 738
88.02%
47.03%
Biochemical
101
119
84.87%
0.01%
Total
1 150 535
1 278 238
90.01%
100.00%
Appendix
Table 6 : Distribution of assay formats available after bulk-downloading and processing PubChem BioAssay, broken down by commented vs. total counts and comment coverage. With Comment counts assays that have a description, comment, or name; Total counts all assays in the category; Coverage is the annotation rate per category (With Comment / Total); Percentage is each category’s share of the commented subset.
BAO Assay Format
With Comment
Total
Coverage
Percentage
cell-based format
112
119
94.12%
44.80%
biochemical format: protein format: single protein format
91
103
88.35%
36.40%
biochemical format: protein format: protein complex format
27
34
79.41%
10.80%
cell based format
9
12
75.00%
3.60%
biochemical format: protein format: Single protein format
5
6
83.33%
2.00%
organism-based format
4
4
100.00%
1.60%
Appendix
Table 7 : Distribution of BAO Assay Formats available after bulk-downloading and processing PubChem BioAssay, broken down by commented vs. total counts and comment coverage. With Comment counts assays that have a description, comment, or name; Total counts all assays in the category; Coverage is the annotation rate per category (With Comment / Total); Percentage is each category’s share of the commented subset.
BAO Detection Technology
With Comment
Total
Coverage
Percentage
AlphaLISA: fluorescence intensity
1
1
100.00%
0.31%
absorbance
2
2
100.00%
0.63%
fluorescence: flow cytometry
37
37
100.00%
11.64%
fluorescence: fluorescence intensity
112
115
97.39%
35.22%
fluorescence: fluorescence polarization
21
24
87.50%
6.60%
fluorescence: fret: htrf
2
3
66.67%
0.63%
Appendix
Table 8 : Distribution of BAO Detection Technologies available after bulk-downloading and processing PubChem BioAssay, broken down by commented vs. total counts and comment coverage. With Comment counts assays that have a description, comment, or name; Total counts all assays in the category; Coverage is the annotation rate per category (With Comment / Total); Percentage is each category’s share of the commented subset.
BioAssay Type
With Comment
Total
Coverage
Percentage
Biochemical
1 866
3 155
59.14%
0.90%
Biochemical ∣ Cell-based
18
22
81.82%
0.01%
Biochemical ∣ Cell-based ∣ In vivo
1
1
100.00%
0.00%
Biochemical ∣ Cell-based ∣ Toxicity
2
3
66.67%
0.00%
Biochemical ∣ In vitro
24
48
50.00%
0.01%
Biochemical ∣ In vivo
68
68
100.00%
0.03%
Appendix
Table 9 : Distribution of BioAssay Types available after bulk-downloading and processing PubChem BioAssay, broken down by commented vs. total counts and comment coverage. With Comment counts assays that have a description, comment, or name; Total counts all assays in the category; Coverage is the annotation rate per category (With Comment / Total); Percentage is each category’s share of the commented subset.
Figure 1 : Figure was made by downloading the BAO complete OWL and creating the subtree in PowerPoint.
Figure 2 : Figure was made by downloading the BAO complete OWL and creating the subtree in PowerPoint.
Parameter
GPT-4o
Gemini 3.6 Flash
Claude Sonnet 4.6
Provider / route
Azure OpenAI
Vertex AI (via AI Gateway)
AWS Bedrock (via AI Gateway)
Endpoint / API
chat.completions
OpenAI-compatible passthrough
bedrock-runtime.converse
Model / deployment ID
gpt-4o
google/ gemini-3.6-flash
us.anthropic.claude-sonnet-4-6
Temperature
0.0
0.0
0.0
Max output tokens
provider default
provider default
provider default
JSON-mode enforced?
yes ( json_object )
yes ( json_object )
no (prompt-only)
Appendix
Table 10: Runtime configuration for the three proprietary API-served models. Temperature was fixed at 0.0 across all models to obtain deterministic (greedy) decoding for reproducibility. Maximum output length was left at the provider default: the response is a compact JSON object with three short fields, well below any provider cap.
Parameter
Gemma 4 31B
GPT-OSS 20B
Gemma 3 27B
Llama 3.3 70B
Ollama model tag
gemma4:31b
gpt-oss:20b
gemma3:27b
llama3.3:70b
Context window ( num_ctx )
9000
9000
9000
9000
Max output ( num_predict )
8192
8192
8192
8192
Temperature
0.0
0.0
0.0
0.0
Seed
42
42
42
42
Structured output
JSON-Schema
JSON-Schema
JSON-Schema
JSON-Schema
Appendix
Table 11: Runtime configuration for the four open-weights models served via Ollama v0.17.4 . All open-weights runs used JSON-Schema-constrained decoding (Ollama’s format= parameter), which enforces the enum of valid BAO labels at generation time.
Figure 11
Figure 5 : Pairwise Cohen’s κ between the three proprietary LLMs (Claude Sonnet 4.6, Gemini 3.6 Flash, GPT-4o) and four open source LLMs (GPT-OSS 20B, Gemma 3 27B, Gemma 4 41B, Llama 70B) on the BAO detection technology task. Agreement patterns mirror the assay format results: near-perfect agreement.
AID
ChEMBL
BARD (if diff.)
Top pred.
Comment
ndis.
nal.
1984
biochemical
cell-based
cell-based
LLMs support BARD label
7
0
2121
biochemical
cell-based
cell-based
LLMs support BARD label
7
0
2217
cell-based
biochemical
biochemical
LLMs support BARD label
7
0
2613
tissue-based
whole-cell lysate (subclass of cell-free)
cell-free
LLMs support BARD label
7
0
588382
organism-based
cell-based
cell-based
LLMs support BARD label
7
0
588766
cell-free
single-protein (part of biochemical)
biochemical
LLMs support BARD label
7
0
Appendix
Table 12: Manual inspection of the 43 assays where the majority of LLMs disagreed with the ChEMBL BAO assay format label. All labels are BAO assay formats; the “ format” suffix is omitted for space. ndis. counts models disagreeing with ChEMBL, nal. counts models agreeing with ChEMBL; seven models were attempted per row unless noted otherwise. Rows shaded in gray were sent for, and received, senior industrial curator review (Table 13 ).
AID
ChEMBL
Top pred.
Short title
Expert format
Expert reason
Comment
1470
biochemical
cell-based
Discovery of novel allosteric modulators of the M1 muscarinic receptor: Agonist NMS binding at M1
cell-free, but could also be biochemical
In protocol: “Membranes were prepared from M1-expressing CHO cells”; format appears to be cell membranes but could be classified as biochemical as well
Expert finds multiple labels plausible
492958
cell-based
organism-based
Counterscreen for AddAB inhibitors: absorbance-based bacterial cell-based high throughput dose response assay for inhibitors of bacterial viability
organism-based
see short title; you might even classify it as “organism based” since these are bacteria cells: later adapted when we discussed definition
Expert agrees with LLM majority: ChEMBL/BARD label incorrect (expert initially leaned towards ChEMBL, but adapted after discussion)
588769
cell-free
biochemical
Late stage assay provider results from the probe development effort to identify inhibitors of plasma platelet activating factor acetylhydrolase (pPAFAH): fluorescence-based dose response biochemical gel-based competitive Activity-Based Protein Profiling (ABPP) assay for HTS compounds
biochemical
fluorescence-based dose response biochemical gel-based competitive Activity-Based Protein Profiling (ABPP) assay for HTS compounds
Expert agrees with LLM majority: ChEMBL/BARD label incorrect
1913
biochemical
cell-free
Luminescence-based dose response biochemical high throughput screening assay for inhibitors of the Heat Shock Protein 90 (HSP90)
biochemical
see short title (it is using a “reticulocyte lysate” so classifying it as “cell-free” makes some sense as well)
Expert agrees with ChEMBL: LLM majority plausible
2693
cell-based
organism-based
Fluorescence Cell-Based Dose Screen to Determine Inhibitors of S. cerevisiae Viability
organism-based
see short title (but could be defined as organism based since it is an organism, S. cerevisiae), but later adapted when we discussed definition
Expert agrees with LLM majority: ChEMBL/BARD label incorrect (expert initially leaned towards ChEMBL, but adapted after discussion)
488745
organism-based
cell-based
Quantitative high throughput screen for delayed death inhibitors of the malarial parasite plastid, 96 hour incubation
cell-based
in protocol: “… Four microliters of infected erythrocytes …”
Expert agrees with LLM majority: ChEMBL label incorrect (LLMs agree with BARD)
Appendix
Table 13 : Assays sent for expert review, with ChEMBL label, top LLM prediction, and expert annotation.
Model
Prediction
Correct?
Cited evidence
Claude
organism-based
✓
“…parasites cultured in the presence of test compounds…measuring susceptibility of the living malarial parasite organism.”
Gemini
organism-based
✓
“The susceptibility of the Dd2 Plasmodium falciparum line against novel small molecules will be determined using a SYBR green-based fluorescence assay.”
Gemma 4 31B
organism-based
✓
“The susceptibility of the Dd2 Plasmodium falciparum line…Parasites are cultured in the presence of serial dilutions of test compounds.”
Gemma 3 27B
cell-based
✗
“A cell-based HTS for delayed death inhibitors of the malarial parasite plastid…Parasites are cultured in the presence of serial dilutions of test compounds…”
GPT-OSS 20B
cell-based
✗
“Parasites are cultured in the presence of serial dilutions of test compounds…using a SYBR green-based fluorescence assay in 384-well plates”
GPT-4o
cell-based
✗
“…explicitly indicates the use of living cells (parasites) in the experimental setup.”
Appendix
Table 14: Model predictions for AID 504631, illustrating the cell-based/organism-based confusion for single-celled eukaryotic organisms. True label: organism-based format .
AID
BAO label (third-party)
Majority pred.
Bioassay annotation (PubChem)
Comment
nmaj. / nal.
588725
spectrophotometry
radiometry
scintillation counting (part of radiometry)
LLMs agree with BARD
7/0
588768
fluorescence
spectrophotometry
absorbance (part of spectrophotometry)
LLMs agree with BARD
7/0
588782
fluorescence
luminescence
luminescence
LLMs agree with BARD
7/0
602192
fluorescence
luminescence
chemiluminescence (part of luminescence)
LLMs agree with BARD
7/0
602194
fluorescence
luminescence
chemiluminescence (part of luminescence)
LLMs agree with BARD
7/0
602231
spectrophotometry
isometric tension recording
label-free method (BARD)
LLMs disagree with all sources
7/0
Appendix
Table 15: Manual inspection of the 34 physical detection method assays where the majority of models disagreed with the third-party BAO detection label in the PubChem bulk download. All labels are BAO physical detection methods; the “ method” suffix is omitted for space. nmaj. is the number of models supporting the majority prediction; nal. is the number of predictions that aligned with the third-party BAO label instead. Seven models were attempted per row unless noted otherwise. Rows shaded in gray were additionally reviewed by a senior industrial curator (see Table 16 for per-model evidence spans).
AID
3rd-party
Majority pred.
Short title
Expert format
Expert reason
623968
spectrophotometry
fluorescence
Inhibitors of the Hepatitis C Virus non-structural protein 3 helicase (NS3)
spectrophotometry
Similar work to the CHRM probe-development assays but “late stage,” so more information is given, e.g. “RNA concentration is determined by reading absorbance at 260 nm”
588778
fluorescence
spectrophotometry
A High Throughput Screening Assay for Inhibitors of Bacterial Motility in Vibrio cholerae
fluorescence (spectroscopy also fine after seeing LLM evidence)
Protocol: “…fluorescence intensity was determined by reading…”
602418
luminescence
fluorescence
Summary of probe development efforts to identify inverse agonists of LRH-1 (NR5A2)
cannot define
A summary of various efforts per the title; no information on what was actually done for each effort
588816
fluorescence
simple measurement
Identify agonists of the human cholinergic receptor, muscarinic 1 (CHRM1)
cannot be defined
No information; this is pre-work to identify a probe, not really an assay in itself
602170
spectrophotometry
visual observation
Probe development efforts to identify activators of Methionine sulfoxide reductase A (MsrA)
cannot be defined
As above (probe-development pre-work)
624103
fluorescence
simple measurement
Identify agonists of the human cholinergic receptor, muscarinic 5 (CHRM5)
cannot be defined
As above
Appendix
Table 16: Detection-method assays sent for expert review, with third-party (PubChem) label, LLM majority prediction, and expert annotation.
AID
3rd-Party Label
Expert Label
Model
Prediction
Conf.
Evidence Span
588816
fluorescence method
cannot be defined
Llama 3.3 70B
simple measurement ✗
low
No specific detection method mentioned
GPT-4o
visual observation ✗
low
No specific detection method or instrumentation mentioned.
Claude Sonnet 4.6
ERROR ✗
Gemini 3.6 Flash
fluorescence ✗
low
Title indicates probe development for human muscarinic receptor 1 (CHRM1), typically fluorescence-based…
Gemma 3 27B
simple measurement ✗
low
No information provided. Conservative assumption…
Gemma 4 31B
label free ✗
low
probe development efforts to identify agonists of CHRM1
Appendix
Table 17 : Detection technology predictions for selected AIDs across all models.
Figure 6 : Error rates across predicted confidence levels for silver-label predictions. The top panel shows results for the ChEMBL baseline, while the bottom panel shows the corresponding results for the detection task. Each bar represents the proportion of incorrect predictions within a confidence category, with the number of samples shown in parentheses.
Figure 7 : Effect of two prompt perturbations on proprietary-LLM predictions for the ChEMBL assay format subset. Removing the BAO definitions (Without defs) changes far more predictions than reversing the class order (Order swap), and disproportionately so for Gemini 3.6 Flash (10.8%). Bar segments indicate whether each changed prediction moved from a different label to the silver label (green), from the silver label to another label (red), or between two labels that are different from the silver label (orange); percentages above each bar give the total fraction of predictions that changed.
Experiment
Model
N
Parse
Input
Output
Cost
Energy
fails
tokens
tokens
($)
(kWh)
ChEMBL baseline
Claude Sonnet 4.6
1097
29
2,302,201
90,590
9.09
–
Gemini 3.6 Flash
1097
0
2,040,815
54,349
3.47
–
GPT-4o
1097
0
1,998,562
67,554
5.67
–
Gemma 3 27B
1097
0
2,081,404
64,558
–
6.44
Gemma 4 31B
1097
0
2,086,917
63,083
–
6.45
Appendix
Table 18: Per-model token usage, parse failures, and cost/energy accounting across all experimental conditions. Input and completion token counts are experiment-wide totals. Cost is reported for API-accessed (proprietary) models; energy for self-hosted (open-weight) models. “–” indicates not applicable.
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.
Chuhan Shi, Xiaoquan Ren, Sicheng Song +3
Southeast University · East China Normal University · Hong Kong University of Science and Technology
Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational model of cellular behavior that could accelerate biological discovery. One of the most compelling promises of this vision is the ability to perform in silico phenotypic screens, in which a model predicts the effects of cellular perturbations in unseen biological contexts. This task combines heterogeneous textual inputs with diverse phenotypic outputs, making it particularly well-suited to LLMs and agentic systems. Yet, no standard benchmark currently exists for this task, as existing efforts focus on narrower molecular readouts that are only indirectly aligned with the phenotypic endpoints driving many real-world drug discovery workflows. In this work, we present AssayBench, a benchmark for phenotypic screen prediction, built from 1,920 publicly available CRISPR screens spanning five broad classes of cellular phenotypes. We formulate the screen prediction task as a gene rank prediction for each screen and introduce the adjusted nDCG, a continuous metric for comparing performance across heterogeneous assays. Our extensive evaluation shows that existing methods remain far from empirically estimated performance ceilings and zero-shot generalist LLMs outperform biology-specific LLMs and trainable baselines. Optimization techniques such as fine-tuning, ensembling, and prompt optimization can further improve LLM performance on this task. Overall, AssayBench offers a practical testbed for measuring progress toward in silico phenotypic screening and, more broadly, virtual cell models.
LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-based indicator that surfaces likely mislabels at 95% precision in the top 200 examples across seven preference and multiple-choice benchmarks using responses from 114 models, outperforming a supervised classifier. We trace these errors to mechanical labeling heuristics, upstream annotation mistakes inherited unchanged from source datasets, and fundamentally ambiguous items without a defensible single label. The same model fit reveals that reward models specialize in stylistic preference rather than factual knowledge, and identifies one frontier reward model that agrees with detected mislabels at 78% accuracy versus 38% for its peers, consistent with benchmark contamination or benchmark-specific over-optimization.