Kurate: Scalable Scientific Quality Analysis
Organizations: Kivira Health, U.S.A. · University of Bern, Bern, Switzerland
Abstract
Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.
Figures & tables
| Information Kurate extracts | Trial protocols | Results publications |
|---|---|---|
| Study design and population | ||
| Who was eligible to take part | 30/30 (1.00) | — |
| What the intervention was | 30/30 (1.00) | 29/30 (0.97) |
| What it was compared against | 30/30 (1.00) | 25/30 (0.83) |
| How many participants | 30/30 (1.00) | 38/46 (0.83) |
| How the outcome was defined | 24/30 (0.80) | 26/29 (0.90) |
| Grade | Include | Flag | Retained | Exclude | Total |
|---|---|---|---|---|---|
| A | 71 | 137 | 208 | 72 | 280 |
| B | 100 | 456 | 556 | 378 | 934 |
| C | 44 | 606 | 650 | 775 | 1,425 |
| D | 10 | 570 | 580 | 891 | 1,471 |
| E | 0 | 65 | 65 | 171 | 236 |
| F | 0 | 1 | 1 | 0 | 1 |
| Domain | Mean score | A/B/C | Excluded | |
|---|---|---|---|---|
| Overall corpus | 4,347 | 0.564 | 60.7% | 52.6% |
| Infectious disease / HIV | 349 | 0.602 | 72.5% | 51.3% |
| Neurology / pain | 138 | 0.593 | 63.0% | 44.9% |
| Cardiometabolic | 1,122 | 0.592 | 68.2% | 51.2% |
| Cancer | 843 | 0.580 | 65.5% | 63.0% |
| Resp. / allergy / immune | 152 | 0.569 | 60.5% | 52.0% |
| Dealbreaker criterion | Papers | Excluded papers |
|---|---|---|
| Control group | 124 | 5.4% |
| Comparator validity | 444 | 19.4% |
| Randomization | 712 | 31.1% |
| Blinding and outcome susceptibility | 398 | 17.4% |
| Analysis aligned with research question | 183 | 8.0% |
| ITT or justified alternative | 320 | 14.0% |
| Dimension | Score | Planned or external evidence | Reported evidence | Kurate’s assessment |
|---|---|---|---|---|
| Statistical power | 1.00 | [registry] planned enrollment “9,250 participants” | [report] “88.7% power …20% effect”; “randomly assigned 9361 persons” | Strong support: the paper reports a specific power calculation, and the number of participants randomized exceeded the planned target. |
| Causal identification | 0.67 | [registry] intensive vs standard systolic blood pressure target groups; comparison of two active treatments; limited detail on blinding | [report] “We randomly assigned 9361 persons …” to “intensive” vs “standard” treatment | Partial support: the paper clearly reports random assignment, but evidence on allocation concealment (i.e., preventing those enrolling participants from knowing the upcoming assignment) and on blinding is not fully explicit in the paper. |
| Preregistration | 1.00 | [registry] first submitted 2010-09-20, before study start on 2010-10-01; NCT01206062 | [report] “ClinicalTrials.gov number, NCT01206062”; the same trial is described in the article | Strong support: Kurate links the article to a registration which was submitted before the trial began, and verifies the dates. |
| Selective reporting | 1.00 | [registry] primary outcome “MI …ACS …Stroke …HF …CVD Death”; secondary outcome “All-cause Mortality” | [abstract] same composite primary outcome; “All-cause mortality was also significantly lower …” | Strong support: the main reported outcomes match the preregistered outcomes, and the early stopping of the trial is reported openly. |
| Measurement validity | 1.00 | [registry] primary outcome is a composite of major cardiovascular events | [report] “primary composite outcome was myocardial infarction …stroke …heart failure …” | Strong support: the outcome is a standard clinical measure and is directly relevant to the question the trial addresses. |
| Analysis prespecification | 1.00 | [protocol / registry] registered in advance; protocol available; planned “time to the first occurrence …” analysis using a Cox model (a standard model for time-to-event data) on randomized participants | [results] “stopped early …owing to” lower primary-event rates; the change from the plan is stated explicitly | Strong support: Kurate finds an analysis plan specified in advance, and the major change to that plan (stopping the trial early) is reported explicitly rather than made without disclosure. |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Protocols (SPIRIT) | Results (CONSORT) | |||
| Train | Held-out | Train | Held-out | |
| Documents | 70 | 30 | 70 | 30 |
| Scored items | 566 | 242 | 906 | 370 |
| Questions | 9 | 9 | 15 | 15 |
| Candidate extraction | 0.921 | 0.934 | 0.846 | 0.838 |
| 95% CI [ Wilson, 1927 ] | — | [0.895, 0.959] | — | [0.797, 0.872] |
| Benchmark item | Kurate field | Candidate extraction | Final value |
|---|---|---|---|
| text10a_participants_inclusion | eligibility criteria | 30/30 | 30/30 |
| text11a_intervention_description | intervention name | 30/30 | 30/30 |
| text11a_intervention_description | comparator name | 30/30 | 30/30 |
| text14b_sample_calculation | sample size | 30/30 | 30/30 |
| text20a_statistical_methods_outc | planned analysis methods | 29/29 | 29/29 |
| text31d_sharing_data | data-sharing statement | 5/5 | 5/5 |
| Benchmark item | Kurate field | Candidate extraction | Final value |
|---|---|---|---|
| text14b_sample_calculation | sample size | 23/23 | 20/23 |
| text37a_analysis_numbers | sample size | 23/23 | 18/23 |
| text38a_outcome_results | -value | 22/22 | 21/22 |
| text11a_intervention_description | intervention name | 29/30 | 29/30 |
| text38a_outcome_results | summary statistic | 28/30 | 23/30 |
| text12a_outcomes_definitions | endpoint name | 26/29 | 26/29 |
| Dimension | Short description | Range | Weight |
|---|---|---|---|
| Statistical power | Whether the study was large enough for its main claim, given the planned and analyzed sample sizes. | 1.35 | |
| Causal identification | Whether the design supports a causal comparison, including randomization, allocation safeguards, comparator choice, and blinding where feasible. | 1.35 | |
| Preregistration | Whether the main claim was specified before the relevant data were collected or analyzed. | 1.20 | |
| Selective reporting | Whether the reported outcomes, analyses, and emphasis match the plan made in advance. | 1.20 | |
| Measurement validity | Whether the outcomes are appropriate for the construct, population, intervention, and timepoint. | 1.00 | |
| Analysis prespecification | Whether the main analysis was planned in advance and answers the stated research question. | 1.00 |
| Scoring element | Mapping | Interpretation |
|---|---|---|
| Dimension score | to for most active dimensions, with conflict/funding using a narrower range. | Negative scores are penalties for fundamental flaws or missing evidence, with zero as neutral and positive scores indicating support. |
| Evidence-grade score | Weighted dimension total normalized to 0–1. | Produces the continuous normalized score reported for papers and cohorts. |
| A | normalized score | Strong evidence quality across the active dimensions. |
| B | normalized score and | Adequate-to-strong evidence quality with limited concerns. |
| C | normalized score and | Moderate evidence quality with notable limitations. |
| D | normalized score and | Limited evidence quality with substantial weaknesses. |
| Hard dealbreaker | Short description |
|---|---|
| Control group | The study needs a usable comparison group for the claim being made. A causal claim cannot rest on a paper with no suitable control or comparison. |
| Comparator validity | The comparator must be clear and appropriate. A very weak, poorly described, or inappropriate comparator can make the claimed effect uninterpretable. |
| Randomization | Assignment should be random, or at least not presented as randomized when it was absent, quasi-random, unsupported, or compromised in a way that could bias treatment assignment. |
| Blinding and outcome susceptibility | Lack of blinding is most serious when the primary outcome is subjective or otherwise easy to bias. Lack of blinding is not automatically fatal when blinding is infeasible and the claim is limited accordingly. |
| Analysis aligned with research question | The model, comparator, outcome, timepoint, and analysis population should answer the stated research question. If they do not, the primary claim may not be supported. |
| ITT or justified alternative | The analysis should preserve the randomized comparison, usually by intention-to-treat (ITT), or give a credible reason for an alternative population. Unjustified completer-only, subgroup, or as-treated analyses can undermine the comparison. |