Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
Organizations: Seoul National University · University of Minnesota
Abstract
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work. We benchmark these failures as scientific slop through six measures across Structure, Argument, and Artifacts. We construct SciSlopBench with 390 AI-generated papers, mostly in computer science but spanning the life, social, and natural sciences, each paired with a human-written paper matched by research problem and contribution type. Our measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars. Higher scientific slop accompanies lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025. Reducing these patterns, however, is not as simple as directly optimizing the measures. We therefore propose SciSlopHarness, a harness-level framework that guides a fixed LLM to revise slop only where the experiment records support the change. While standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline without requiring human reference targets. Overall, we demonstrate that AI-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement.
Figures & tables
| Item | Illustration | Unit of analysis | Score | What means | |
|---|---|---|---|---|---|
| Cross-section references | Body sections and labeled objects | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}objects never referenced outside their own section\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all objects\@add@centering}} | No object is ever referred to from another section. | ||
| STRUCTURE | Macro redundancy | Sentences with at least eight tokens | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}sentences half or more copied as 8-grams from an earlier section\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all sentences\@add@centering}} | Every sentence mostly repeats earlier sections. | |
| Argument graph | Key claims in the Introduction | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}shallow claims: nothing earlier leads up to them\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all key claims\@add@centering}} | The argument is flat. Every claim stands alone with no build-up behind it. | ||
| ARGUMENT | Citation isolation | Citation sentences (Intro., Related Work) | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}citations not grouped, compared, or related to other works\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all citations\@add@centering}} | Every citation stands alone. | |
| Figure exposition | Content types in method figures | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}content types not needed to show the method\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all content types in the figure\@add@centering}} | Nothing in the figure shows the method itself. | ||
| ARTIFACTS | Evidence gap | Papers with a body result table | \dfrac{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}papers showing no concrete input, output, or case\@add@centering}}{\parbox{394.98499pt}{\centering\color[rgb]{0,0,0}all papers\@add@centering}} | The paper gives no concrete example at all. |
| Macro redundancy Shaded sentence repeats the Abstract. Slop-aware drops the citations. SciSlopHarness keeps them. | ||
|---|---|---|
| Original | RLVR has emerged as a powerful paradigm… from early work on CodeRL [1] and RLTF [2] to recent breakthroughs like DeepSeek-R1 [3] and DeepSeekMath [4] … | |
| Slop-aware | … Early work in Section [Related Work] demonstrates… from pioneering applications of unit test signals to recent breakthroughs like DeepSeek-R1 and DeepSeekMath … | |
| SciSlopHarness | In code generation, unit tests provide a natural verifier… from early work on CodeRL [1] and RLTF [2] to recent breakthroughs like DeepSeek-R1 [3] and DeepSeekMath [4] … | |
| Cross-section references Table [tau] is unreferenced. Slop-aware scatters references. SciSlopHarness adds one where shaded. | ||
| Original | We presented Delta-Prefill Switching (DPS)… DPS achieves 21–22% speedup over greedy decoding… The method is robust to threshold selection and requires no model modifications. | |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Stratum | Pairs | Mean cosine | Rank 1 | Rank | Rank | Median rank |
|---|---|---|---|---|---|---|
| All pairs | 143 | 0.927 | 37 | 71 | 101 | 4 |
| Cited pool | 120 | 0.928 | 33 | 61 | 85 | 3 |
| Widened pool | 23 | 0.919 | 4 | 10 | 16 | 5 |
| Cross-pair reference | 20,306 | 0.881 |
| Benchmark Property | Scientific text | Whole paper | Tables and figures | Released code | Executed research | Matched human pair | Multiple genres | Adversarial variants |
| RAID ( Dugan et al., 2024 ) | – | – | – | – | ||||
| DetectRL ( Wu et al., 2024 ) | – | – | – | – | ||||
| M4 ( Wang et al., 2024 ) | – | – | – | – | – | |||
| MAGE ( Li et al., 2024 ) | – | – | – | – | – | |||
| CHEAT ( Yu et al., 2023 ) | – | – | – | – | – | |||
| IDMGSP ( Mosca et al., 2023 ) | – | – | – | – |
| Setting | Configuration used for Figure 5 |
|---|---|
| Editor | claude-haiku-4-5-20251001 |
| Reviewer | claude-sonnet-5 , no tools |
| Retirement call | claude-haiku-4-5-20251001 |
| Skill | SciSlop_v0.5.md , staged as SciSlop.md |
| Budget | At most three rounds |
| Measured patterns | All six |
| Pattern | Repair direction | Reported location |
|---|---|---|
| Cross-section references | Refer to an object from one sentence that uses it, placed where the result is interpreted or the method is applied. Leave genuinely local objects unchanged. Add no sentence that only points. | Label, object kind, home section |
| Macro redundancy | Rewrite the recycled sentence so that it adds a detail, consequence, or qualification its section needs, or remove it. Do not paraphrase to disguise repetition. | Section, sentence, source section |
| Argument graph | Move the context that leads up to a shallow claim ahead of it, or move the claim after that context. Do not add, delete, or weaken claims. | Claim sentence and the supporting sentence that currently follows it |
| Citation isolation | State how a cited work relates to another work already cited, in one sentence citing both, using only relations the manuscript already supports. | Section and citing sentence |
| Figure exposition | Remove the expository elements and state them once in the text if needed. Keep every component and connection. The editor lists the phrases to erase and a deterministic step erases them. | Figure, element type, transcribed element text |
| Evidence gap | Never invent an instance. (a) If the experiment records hold a specimen, display it verbatim with its source file. (b) If they hold none, say so where the aggregate result is interpreted: no individual case was retained or inspected. | One manuscript-level observation: silent, acknowledged, or exhibited |
| Check | Condition to report relative to | Class |
|---|---|---|
| Cited works | Addition or removal of a cited bibliography key | Hard |
| Dangling citations | A newly unresolved citation key | Hard |
| Reference targets | A reference to an absent label | Hard |
| Numeric values | An original body numeric token is absent | Hard |
| Body length | Revised-to-original word ratio outside | Hard |
| Labels | An original label is removed | Audit |
| Measure | Original | Base prompting | Claude Code | Reviewer-based | Slop-aware | SciSlopHarness |
|---|---|---|---|---|---|---|
| Macro redundancy | 0.399 | 0.314 | 0.191 | 0.213 | 0.400 | 0.100 |
| Cross-section refs. | 0.201 | 0.193 | 0.180 | 0.141 | 0.660 | 0.078 |
| Argument graph | 0.200 | 0.200 | 0.200 | 0.200 | 0.200 | 0.048 |
| Citation isolation | 0.289 | 0.279 | 0.188 | 0.290 | 0.411 | 0.085 |
| Figure exposition | 0.265 | 0.265 | 0.265 | 0.265 | 0.265 | 0.089 |
| Evidence gap | 0.427 | 0.432 | 0.372 | 0.372 | 0.372 | 0.119 |
| Pattern | Definition | What to do |
|---|---|---|
| Cross-section references | A paper declares objects: its body sections and the labelled figures, tables, equations, and algorithms inside them. An object is unused when no section other than the one that contains it ever refers to it, so the rest of the paper never builds on it. | Where the paper’s argument actually relies on an object, refer to it from the section that relies on it (for example, discuss the result of a table in the section whose claim it supports, and point back to the method’s equation where the experiments use it). Do not add pointers that the surrounding text does not use. |
| Macro redundancy | A sentence is recycled when at least half of it repeats, almost word for word, text that an earlier, different section already contained. | Rewrite each recycled sentence so that it adds what its own section needs (a new detail, a consequence, a qualification), or remove it if the section does not need it. Do not paraphrase merely to disguise the repetition. |
| Citation isolation | A citing sentence is isolated when it cites a single prior work and neither groups it with another work, nor mentions another work elsewhere in the sentence, nor states a relation between two works. A related-work section made of isolated sentences lists prior work instead of positioning it. | Where the paper’s positioning depends on a cited work, state how it relates to other works already cited in the paper: a shared limitation, a difference in assumption or method, or what this paper takes from each. Only use works already cited in the paper; do not add new citations. |
| Evidence gap | The paper reports aggregate results but never displays a single concrete instance of what it counts: an example input and output, a case, a failure, a worked example, or a quoted specimen. | If the paper’s own materials in this directory contain such an instance (for example a prompt, a generated output, a case in the appendix), display one where the reader needs it and refer to it from the text. If no instance is available in the materials, do not invent one; instead say explicitly, in the relevant section, that no instance is shown. |
| Argument graph | A key claim of the Introduction is a sentence asserting superiority, a prior limitation, or a design choice. The claim’s strongest contextual cue is the Introduction sentence that most raises the claim’s plausibility. The claim is declared rather than argued when that cue occurs after the claim, so the Introduction states the conclusion before building toward it. | Reorder so that the context precedes the claim it supports. Move the limitation of prior work, the observation, or the design constraint ahead of the sentence that asserts the contribution, or move the assertion down to follow it. Do not delete claims and do not add new claims. Do not weaken a claim to avoid arguing for it. |
| Figure exposition | A method diagram should depict the mechanism. Its components and how they connect. A box, panel, or caption element is expository when it carries material that belongs to the text rather than the mechanism. Experimental settings, dataset names, numeric results, interpretive claims, a legend that restates the prose, or a step-by-step narration of the pipeline. | Remove the expository material from the method diagram and, where the reader needs it, state it once in the text near the figure. Keep only components of the method and their connections. Remove exactly the elements listed as expository and keep every other element and every arrow the image has. Do not invent components the figure does not show, do not simplify the mechanism to make the drawing easier, and do not drop the figure. |
| Citation isolation Shaded sentence relates CALM to nothing. Slop-aware leaves CALM alone. SciSlopHarness adds InfMem. | ||
|---|---|---|
| Original | For language models, CALM [1] uses confidence-based early exiting at the token level , achieving significant speedups on text generation. | |
| Slop-aware | For language models, confidence-based methods like CALM [1] extend these layer-wise approaches to the token level , achieving significant speedups during text generation… | |
| SciSlopHarness | For language models, CALM [1] uses confidence-based early exiting at the token level, achieving significant speedups on text generation, whereas InfMem [2] learns stopping policies at the chunk level for memory agents, achieving 3.3–5.1 speedup . | |
| Evidence gap No concrete instance appears in the paper. Slop-aware invents one. SciSlopHarness states the gap. | ||
| Original | No example input, output, case, or failure is displayed; the paper reports aggregate accuracies and agreement rates only. | |