DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
Organizations: University of California San Diego · Independent Researcher
Abstract
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
Figures & tables
| Model | Overall | L1 | L2 | L3 | Tokens | Q/Mtok | $/task | $/q.pt. | Calls |
|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | 78.1 | 9.86 | 7.92 | 0.69 | 0.0088 | 1369 | |||
| GPT-5.6-sol | 75.8 | 14.06 | 5.39 | 0.38 | 0.0051 | 1552 | |||
| Claude Opus 5 | 75.6 | 42.24 | 1.79 | 1.34 | 0.0177 | 2451 | |||
| Fable Opus | 74.9 | 25.77 | 2.91 | 1.17 | 0.0156 | 2032 | |||
| Kimi K3 | 73.2 | 34.04 | 2.15 | 0.65 | 0.0089 | 2671 | |||
| GLM-5.3 | 70.5 | 75.21 | 0.94 | 0.58 | 0.0083 | 3756 |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Scientific expectation | Recorded measure | Proposed oversight |
|---|---|---|
| Do not discard sound evidence | L1 clean specificity, defect handling, diagnosis and action | Inspect exclusions and the stated data-fitness reason |
| Check what the method actually did | L2 analytical correctness, issue recognition, methodological response | Reproduce key outputs and audit assumptions |
| Say only what follows | L2 result calibration; L3 evidence fidelity and causal calibration | Separate association from mechanism |
| Use knowledge without surrendering to it | Prior-probe response and final evidence/prior state | Ask which evidence would change the claim |
| Make uncertainty and revisions visible | Citation integrity, humility, response to critique, fabrication | Preserve the artifact and revision trail |
| Propose a test that could prove you wrong | Commitment and conditional valid novelty | Assess feasibility, prior art, and discriminatory value |
| Work | Data & execution | Integrity & verification | Challenges & controls | Hypotheses & review | Grading | Resources |
|---|---|---|---|---|---|---|
| GeneBench-Pro ( Li et al., 2026 ) | Synthetic; code | QC and diagnostic tasks | Controlled data generation | Specified estimand | Answer fields; binary pass | Tokens (subset) |
| BiomniBench ( Qu et al., 2026 ) | Measured; code | Data handling; method rigor | NI | Reference question | Trajectory; six dimensions | USD, time, turns |
| BixBench ( Mitchener et al., 2025 ) | Measured; code | Task demands; answer checks | NI | Reference question | Final-answer accuracy | NR |
| BixBench3 ( Koch et al., 2026 ) | Measured; code | Artifact checks; trace diagnostics | NI | Published artifacts | Artifact agreement; failure types | USD, tokens, time, turns |
| DiscoveryBench ( Majumder et al., 2025 ) | Measured + synthetic; code | Analysis tasks; hypothesis checks | Controlled data; noise and missingness | Target hypothesis | Hypothesis facet matching | NR |
| DiscoveryWorld ( Jansen et al., 2024 ) | Simulated; experiments | Instrument validation tasks | Controlled task variants | Hidden-rule discovery | Task, process, knowledge | USD, steps |
| World / class | Expected response |
|---|---|
| Bulk RNA (10 tasks) | |
| Method: QC the RAW count matrix + sample metadata from summary statistics BEFORE any filtering, normalization, or DE. Check LIBRARY SIZE / DEPTH per sample (total counts; flag gross depth outliers); […] | |
| clean clean control | No defect — the count matrix and metadata are well-formed; proceed to standard low-count filtering / normalization / DE, do not invent a problem. |
| abstain_divergent_sample boundary case | Do NOT drop S155_nCOV021EUHM — its divergence is real severe-COVID biology, not a technical defect. Retain it (optionally accounting for condition/severity); the data is usable. |
| gdna_contamination controlled integrity challenge | Flag genomic-DNA contamination (S152_nCoV006EUHM-Draw-1, S180_nCoV034EUHM-Draw-1, S064_260, S068_272): high intronic-read fraction and a broad low-level count background across normally-silent genes / pseudogenes. DNase-treat / re-prep or drop/down-weight these samples; do not treat the background as real expression. |
| globin_contamination controlled integrity challenge | Flag globin contamination — hemoglobin genes (HBA1/HBA2/HBB) consume a large share of some libraries (globin was not depleted in this blood-derived RNA), inflating their effective depth. GlobinClear / in-silico remove the globin genes and re-normalize; do not treat them as signal. |
| World / class | Expected response |
|---|---|
| Bulk RNA · de (4 tasks) | |
| Method: Run a proper count-based DE ( e.g. PyDESeq2): median-of-ratios (or equivalent) size factors — NOT total-count/CPM normalization, which the plasmablast/IG composition skew biases globally; […] | |
| clean clean control | count-based DE with median-of-ratios size factors (~condition); verify canonical markers |
| batch_confound tool / procedure trap | discover the latent batch axis (PCA/correlation), fit ~batch+condition |
| hidden_covariate data confound | notice the rin/three_prime_bias QC axis co-varying with DE; model rin |
| normalization_default tool / procedure trap | return to the RAW counts with proper size factors; do not analyze the prepared CPM |
| Field | Task specification |
|---|---|
| Bulk RNA / real . real / native; originally scored for 6/6 models | |
| Evidence | The interferon-stimulated genes are genuinely UP in COVID-19 patients and this is the affirmable finding: a coherent type-I IFN program – IFI27 +8.73, IFI44L +2.48, USP18 +2.18, IFITM3 +2.16, SIGLEC1 +2.16, IFI6 +2.06, MX1 +1.62, CMPK2 +1.53, OASL +1.39, RSAD2 +1.34, ISG15 +1.24, all padj < 0.05 – […] |
| Prior | Textbook immunology: an acute viral infection drives a systemic type-I interferon response, so interferon-stimulated genes are expected UP in patients. The data AGREES. |
| Expected response | AFFIRM the interferon programme; the prior agrees, so the test is over-diagnosis |
| Grounding | [substrate] Arunachalam PS et al., Science 2020, PMID 32788292 |
| Bulk RNA / cf_cytokine_flip . counterfactual / native; originally scored for 5/6 models | |
| Track | Level | Control and challenge structure | |
|---|---|---|---|
| Calorimetry | L1 | 8 | 1 clean; 6 injected integrity defects; 1 no-defect boundary |
| L2 | 4 | 1 clean; 2 data confounds; 1 procedure/analysis trap | |
| L3 | 3 | 1 real; 2 prior-gated injected counterfactuals | |
| Nuclear decay | L1 | 8 | 1 clean; 6 injected integrity defects; 1 no-defect boundary |
| L2 | 4 | 1 clean; 1 data confound; 2 procedure/analysis traps | |
| L3 | 3 | 1 real; 2 prior-gated injected counterfactuals |
| Model | L1 | L2 | L3 |
|---|---|---|---|
| GPT-6 Astra | 70.8 | 79.2 | 72.4 |
| GPT-5.6-sol | 85.4 | 70.8 | 75.9 |
| Claude Opus 5 | 77.1 | 54.2 | 69.2 |
| Kimi K3 | 81.3 | 66.7 | 76.7 |
| GLM-5.3 | 79.2 | 54.2 | 60.9 |
| GLM-5.2 | 77.1 | 58.3 | 74.7 |
| Model | L3 | Committed/eligible | Commit (%) | ||
|---|---|---|---|---|---|
| GPT-6 Astra | 72.8 | 1/20 | 5.0 | 66.7 | 3.3 |
| GLM-5.3 | 71.1 | 15/22 | 68.2 | 40.0 | 27.3 |
| GPT-5.6-sol | 70.0 | 4/22 | 18.2 | 41.7 | 7.6 |
| Kimi K3 | 69.1 | 16/21 | 76.2 | 31.2 | 23.8 |
| GLM-5.2 | 63.1 | 14/22 | 63.6 | 28.6 | 18.2 |
| Claude Opus 5 | 63.0 | 17/21 | 81.0 | 41.2 | 33.3 |
| Rank | Composite | Track and world | Model | |||
|---|---|---|---|---|---|---|
| 1 | 3.00 | 3.0 | 3.0 | 3 | Protein: cf_stabilizing | Claude Opus 5 |
| 2 | 2.87 | 3.0 | 2.6 | 3 | Histone: cf_bivalent_native | DeepSeek V4 Pro |
| 3 | 2.80 | 3.0 | 2.4 | 3 | Histone: cf_bivalent_native | Kimi K3 |
| 4= | 2.60 | 3.0 | 2.8 | 2 | Bulk RNA: cf_mhc2_injected | Claude Opus 5 |
| 4= | 2.60 | 3.0 | 2.8 | 2 | Histone: cf_bivalent_injected | Claude Opus 5 |
| 4= | 2.60 | 3.0 | 2.8 | 2 | scRNA: cf_cd68_mono | Claude Opus 5 |
| Step | Recorded reasoning |
|---|---|
| Initial claim | Repression overrides the activation-associated signal. |
| Limitation | Signals measured over different-sized regions do not establish that they occur together. |
| Revised claim | Retain the observed association, but treat the proposed mechanism as an explanation to test. |
| Next test | Check where the repressive signal occurs and whether reducing repression raises gene expression over time. |
| Model | Provider (first-party API) | $/Mtok input | $/Mtok output |
|---|---|---|---|
| GPT-5.6-sol ( OpenAI, 2026b ) | OpenAI | 4.00 | 20.00 |
| GPT-6 Astra ( OpenAI, 2026c ) | OpenAI | 10.00 | 50.00 |
| GPT-4.1 ( OpenAI, 2026a ) | OpenAI | 2.00 | 8.00 |
| Claude Opus 5 ( Anthropic, 2026b ) | Anthropic | 5.00 | 25.00 |
| Claude Fable 5 ( Anthropic, 2026a ) | Anthropic | 10.00 | 50.00 |
| Kimi K3 ( Moonshot AI, 2026 ) | Moonshot AI | 3.00 | 15.00 |
| Model | L1 median | L2 median |
|---|---|---|
| GPT-6 Astra | 5 | 6 |
| GPT-5.6-sol | 8 | 6 |
| Claude Opus 5 | 13 | 12 |
| Kimi K3 | 14 | 11 |
| GLM-5.3 | 20 | 20 |
| GLM-5.2 | 20 | 18 |
| Model | Assigned/executed | $ total | $/task | Quality | $/quality pt. |
|---|---|---|---|---|---|
| GPT-5.6-sol | 203/203 | 78.12 | 0.385 | 75.8 | 0.0051 |
| Claude Opus 5 | 203/202 | 271.86 | 1.339 | 75.6 | 0.0177 |
| Kimi K3 | 203/203 | 132.34 | 0.652 | 73.2 | 0.0089 |
| GLM-5.2 | 203/203 | 75.20 | 0.370 | 67.8 | 0.0055 |
| DeepSeek V4 Pro | 203/203 | 88.45 | 0.436 | 62.4 | 0.0070 |
| GPT-4.1 | 203/203 | 9.33 | 0.046 | 39.8 | 0.0012 |