I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
Organizations: Universidade da Coruña, CITIC
Abstract
For better or worse, LLMs are by now used routinely for scientific writing.\footnote{This paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments).} Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emph{rather than impressing them}. We study the construction \emph{rather than} in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emph{annoying} and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emph{Annoying} uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.
Figures & tables
| Dataset | Selection | Body text | Papers | Body words | Annotated, by batch | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Median | 1 | 2 | 3 | 4 | Total | ||||
| ACL 2019 | Anthology, all venues | pdftotext | 4,744 | 4,575 | 4,648 | 148 | 100 | 40 | 50 | 338 |
| short ( 7 pp.) | 1,785 | 2,949 | 3,016 | |||||||
| long ( 9 pp.) | 2,590 | 5,775 | 5,673 | |||||||
| arXiv 2026 | cs.CL , *ACL venue and style | detex | 1,821 | 7,626 | 7,167 | 150 | 238 | 180 | 568 | |
| short ( 3.5k words) | 66 | 2,541 | 2,787 | |||||||
| Question | Data | Method | Sec. |
|---|---|---|---|
| Q1: How frequent is the construction? | ACL 2019, arXiv 2026, LLM-written papers; arXiv version pairs | Pattern counts per 1,000 words | 5.1 |
| Q2: Do readers find it annoying ? | 1,000 sampled occurrences | Blind labels from two annotators | 5.2 |
| Q3: What correlates with an annoying use? | Labeled arXiv 2026 occurrences | Straw-man codes, stance ratings, syntax, embeddings | 5.3 |
| Q4: Where does it come from? | Four preference datasets; UltraFeedback completions; pool sentences | Chosen vs. rejected responses; honesty vs. helpfulness instruction; reward-model minimal pairs | 5.4 |
| ID | Annoy. | Straw m. | Auth. | Author | Stage |
|---|---|---|---|---|---|
| A1 | ✓ | ✓ | ✓ | yes | PhD + 5 yrs |
| A2 | ✓ | ✓ | ✓ | yes | PhD + 17 yrs |
| A3 | ✓ | ✓ | yes | PhD student, yr 1 | |
| A4 | ✓ | no | PhD in 2026 | ||
| A5 | (51) | ✓ | yes | PhD student, yr 2 |
| Set | Items | N |
|---|---|---|
| Development | A1 unsure | 20 |
| legitimate, ACL 2019 | 10 | |
| Test | A1 annoying (all) | 41 |
| A1 legitimate (random) | 80 |
| Rate | Precision | ||||
| Type | Pattern | 2019 | 2026 | 2019 | 2026 |
| Papers | 4,744 | 1,821 | |||
| Substitution | rather than | 0.134 | 0.993 | 100 | 100 |
| instead of | 0.165 | 0.105 | 100 | 98 | |
| as opposed to | 0.025 | 0.007 | 100 | 100 | |
| Negation | not X but Y | 0.097 | 0.126 | 58 | 84 |
| Words | Papers | Per 1k words | |||
| Source | Mean | Median | with RT | RT | X, not Y |
| 200 ACL 2019 papers | |||||
| Original | 4,493 | 4,354 | 60 | 0.11 | 0.04 |
| GPT-4o | 1,473 | 1,438 | 20 | 0.08 | 0.00 |
| revised | 1,342 | 1,311 | 17 | 0.07 | 0.01 |
| GPT-5.6-sol | 8,453 | 8,412 | 200 | 1.62 | 0.22 |
| Words | Papers | Per 1k words | |||
| Source | Mean | Median | with RT | RT | X, not Y |
| 45 ACL 2019 papers | |||||
| Original | 4,506 | 4,656 | 9 | 0.07 | 0.06 |
| Tülu 3 8B SFT | 1,296 | 1,393 | 3 | 0.04 | 0.00 |
| revised | 1,269 | 1,295 | 2 | 0.02 | 0.02 |
| Tülu 3 8B DPO | 1,551 | 1,513 | 6 | 0.09 | 0.00 |
| Set | Papers | N | A1 | A2 | Either |
|---|---|---|---|---|---|
| Pool | ACL 2019 | 248 | 4 (1.6) | 0 (0.0) | 4 (1.7) |
| arXiv 2026 | 652 | 75 (11.5) | 54 (8.3) | 116 (17.9) | |
| High-count | 100 | 12 (12.0) | 12 (12.0) | 18 (18.0) | |
| GPT set | ACL 2019 | 40 | 0 (0.0) | 0 (0.0) | 0 (0.0) |
| arXiv 2026 | 180 | 32 (17.9) | 9 (5.0) | 35 (19.6) | |
| GPT-4o | 46 | 4 (8.7) | 0 (0.0) | 4 (8.7) |
| Pool | GPT set | |
| Between annotators | ||
| Items (garbled by neither) | 989 | 398 |
| Raw agreement, three labels | 73.6% | 68.8% |
| Cohen’s , three labels | 0.18 | 0.16 |
| Cohen’s , annoying vs. other | 0.18 | 0.20 |
| Gwet’s AC1, annoying vs. other | 0.86 | 0.82 |
| ACL 2019 | arXiv 2026 | |
| Papers | 4,744 | 1,821 |
| with rather than | 36% | 94% |
| Uses per paper, mean (median) | 0.65 (0) | 8.0 (6) |
| P(at least one annoying use per paper) | ||
| A1 | 1.1% | 51.1% |
| A2 | 0 ( 1.0%) | 42.1% |
| Set | Items | A1: annoying | Gap X Y |
|---|---|---|---|
| ACL 2019 originals | 50 | 2 (4.0) | 1.22 |
| Claude | 163 | 26 (16.0) | 0.81 |
| first draft | 88 | 9 (10.2) ‡ | 0.81 |
| revised | 75 | 17 (22.7) † | 0.81 |
| Sonnet 5.5 | 84 | 16 (19.0) | 0.90 |
| Opus 5.5 | 79 | 10 (12.7) | 0.72 |
| Revised | Disavowal | Stance | |
|---|---|---|---|
| Revised only | 0.91 (.043) | – | – |
| disavowal | 0.58 (.28) | 3.19 ( .001) | – |
| disavowal, stance | 0.62 (.26) | 3.25 ( .001) | 0.16 (.40) |
| Codes / labels | Annoying | Legit. | |
|---|---|---|---|
| Codes and labels from different people | |||
| A2+A3+A4 / A1 | 0.28 (29) | 0.12 (54) | .007 |
| A3+A4 / A1 | 0.34 (31) | 0.18 (58) | .02 |
| A2 / A1 | 5/36 (14%) | 2/74 (3%) | .04 |
| A3 / A1 | 16/33 (48%) | 18/59 (31%) | .12 |
| A4 / A1 | 6/39 (15%) | 5/76 (7%) | .18 |
| Gap X Y | Y | |||||
| Labels | Ann. | Leg. | Ann. | Leg. | ||
| All items | ||||||
| A1 | +2.21 | +0.86 | .001 | 0.97 | 0.36 | .001 |
| A2 | +1.62 | +0.94 | .006 | 0.73 | 0.40 | .01 |
| Either | +1.87 | +0.76 | .001 | 0.81 | 0.32 | .001 |
| Non-inverted items | ||||||
| Papers | Uses | Gap X Y | Y | Y 0 |
|---|---|---|---|---|
| ACL 2019 | 281 | 0.87 | 0.30 | 42% |
| arXiv 2026 | 828 | 1.06 | 0.45 | 55% |
| GPT-4o | 46 | 1.13 | 0.46 | 57% |
| GPT-5.6-sol | 67 | 1.39 | 0.67 | 64% |
| GPT-6-sol | 66 | 1.73 | 0.86 | 73% |
| Source | Uses | Gap X Y | Y | Y 0 |
|---|---|---|---|---|
| Original papers | 451 | 1.13 | 0.47 | 57% |
| Tülu 3 8B SFT | 23 | 2.30 | 0.91 | 78% |
| revised | 23 | 2.26 | 0.96 | 83% |
| Tülu 3 8B DPO | 21 | 0.14 | 0.19 | 33% |
| revised | 7 | 0.14 | 0.14 | 43% |
| Tülu 3 8B (final) | 26 | 0.85 | 0.62 | 15% |
| Set | Ann. | Disavowal | Other | |
|---|---|---|---|---|
| Claude set | A1 | 26/42 (62) | 22/268 (8) | .001 |
| Pool sample | A1 | 27/39 (69) | 53/200 (26) | .001 |
| A2 | 13/39 (33) | 43/200 (22) | .15 | |
| GPT set | A1 | 25/48 (52) | 31/350 (9) | .001 |
| A2 | 10/48 (21) | 17/350 (5) | .001 |
| Source | Uses | Disavowals (%) |
|---|---|---|
| GPT comparison set | ||
| ACL 2019 papers | 40 | 0 (0.0) |
| arXiv 2026 papers | 179 | 19 (10.6) |
| GPT-4o | 46 | 0 (0.0) |
| GPT-5.6-sol | 67 | 19 (28.4) |
| GPT-6-sol | 66 | 10 (15.2) |
| Pattern | A1 | A2 | Either | |||
|---|---|---|---|---|---|---|
| Items | 87 / | 598 | 66 / | 582 | 134 / | 486 |
| Initial Rather than | 15 / | 3 ∗† | 8 / | 3 | 12 / | 2 ∗† |
| Subject we | 30 / | 13 ∗† | 17 / | 15 | 25 / | 13 ∗ |
| X adjective | 9 / | 4 | 20 / | 4 ∗† | 13 / | 3 ∗† |
| X adj. modifier | 2 / | 1 | 6 / | 0 ∗† | 4 / | 0 ∗ |
| X prep. object | 16 / | 20 | 3 / | 23 ∗† | 12 / | 22 ∗ |
| Unit | A1 | A2 | Either |
|---|---|---|---|
| Items | 87 / 598 | 66 / 582 | 134 / 486 |
| Chance | 13% | 10% | 22% |
| X | 20% ∗ | 22% ∗ | 35% ∗ |
| Y | 25% ∗‡ | 18% ∗ | 35% ∗‡ |
| Sentence | 22% ∗ | 16% ∗ | 33% ∗ |
| Paragraph | 14% | 9% | 23% |
| Measure | Passages or labels | A1 | A2 |
| LLM-leaning | ACL 2019 (14) | 7% | 0% |
| arXiv 2026 (109) | 47% | 27% | |
| high-count (19) | 68% | 53% | |
| Hunch vs. annoyance, | own labels | 0.01 | 0.18 |
| other’s labels | 0.21 | 0.02 |
| Measure | Passages or labels | A3 | A5 |
|---|---|---|---|
| LLM-leaning, % | ACL 2019 | 41 / 21 | 24 / 11 |
| arXiv 2026 | 48 / 20 | 46 / 25 | |
| GPT | 37 / 22 | 43 / 22 | |
| GPT vs. ACL, AUC | all GPT | .47 | .59 |
| from ACL papers | .42 | .54 | |
| With vs. filler, AUC | .65 | .62 |
| UF | Tülu 3 | HS2 | HS3 | |
|---|---|---|---|---|
| Rater | GPT-4 | GPT-4o | human | human |
| Pairs | 61k | 270k | 7k | 22k |
| rather than | 0.99 | 1.09 | 1.37 | 1.24 |
| not only…but | 1.39 | 1.14 | 1.41 | 1.13 |
| X, not Y | 1.24 | 0.90 | 1.14 | 1.35 |
| not X but Y | 0.94 | 0.81 | 1.06 | 1.25 |
| ACL 2019 | ||||
|---|---|---|---|---|
| Reward model | O D | C D | O C | O C |
| Qwen3 0.6B | 0.88 (89) | 0.62 (76) | 0.26 (71) | 0.15 (62) |
| Qwen3 1.7B | 0.83 (84) | 0.54 (73) | 0.29 (67) | 0.07 (51) |
| Qwen3 8B | 1.54 (89) | 0.91 (75) | 0.63 (71) | 0.26 (58) |
| Llama-3.1 8B | 2.46 (76) | 1.33 (64) | 1.13 (70) | 0.96 (67) |
| Reward model | A1 | A2 |
|---|---|---|
| Annoying / legitimate | 40 / 261 | 26 / 264 |
| Qwen3 0.6B | 0.26 | 0.09 |
| Qwen3 1.7B | 0.36 | 0.20 |
| Qwen3 8B | 0.88 | 0.43 |
| Llama-3.1 8B | 1.69 | 0.72 |
| Pattern | Honesty | Truthfulness | Calibration |
|---|---|---|---|
| rather than | 1.15 | 1.08 | 0.85 |
| not X but Y | 1.34 | 0.97 | 0.93 |
| instead of | 1.13 | 1.07 | 0.79 |
| not only…but | 0.68 | 0.84 | 0.72 |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Count |
|---|---|
| cs.CL submissions, 2026 12 12 12 Submitted 2026-01-01 through 2026-09-11. | 19,485 |
| Comment mentions an *ACL venue | 1,985 |
| 2 arXiv versions | 819 (41%) |
| ACL 2019 | arXiv 2026 | |
| (N=2,590) | (N=1,748) | |
| Limitations section | 13 (0.5%) | 910 (52.1%) |
| Ethics Statement | 10 (0.4%) | 381 (21.8%) |
| Author Contributions | 8 (0.3%) | 9 (0.5%) |
| Mean body words | 5,775 | 7,847 |
| minus the above | 5,765 | 6,599 |
| Quantity (long-format papers) | Mean/paper |
| ACL 2019 parenthetical citations | 25.3 |
| ACL 2019 reference-list entries | 36.1 |
| arXiv 2026 [redacted] tokens | 148.1 |
| Section | Mean words |
|---|---|
| abstract | 158 |
| introduction | 986 |
| related work | 944 |
| preliminaries | 164 |
| method | 408 |
| experiments | 805 |
| Model (snapshot) | Access | Decoding | Max output | Run |
|---|---|---|---|---|
| GPT-4o | OpenAI API | 0.8, top- 0.92 | 16,384 | Sep 30 |
| GPT-5.6-sol | OpenAI API | API defaults | 16,384 r | Sep 30–Oct 2 |
| GPT-6-sol | OpenAI API | API defaults | 16,384 r | Sep 30–Oct 1 |
| Tülu 3 8B SFT ( Llama-3.1-Tulu-3-8B-SFT ) | local GPU, fp16 | 0.8, top- 0.92 | 16,384 | Oct 6–7 |
| Tülu 3 8B DPO ( Llama-3.1-Tulu-3-8B-DPO ) | local GPU, fp16 | 0.8, top- 0.92 | 16,384 | Oct 6–7 |
| Tülu 3 8B final ( Llama-3.1-Tulu-3-8B ) | local GPU, fp16 | 0.8, top- 0.92 | 16,384 | Oct 6–7 |
| Model | Source | 13-gram % | Longest | RT copied |
|---|---|---|---|---|
| GPT-4o | ACL 2019 | 0.27 / 0.02 | 10 / 0 | 0 / 44 |
| arXiv 2026 | 0.08 / 0.04 | 12 / 9 | 0 / 133 | |
| GPT-5.6-sol | ACL 2019 | 0.04 / 0.03 | 11 / 11 | 0 / 6,312 |
| arXiv 2026 | 0.01 / 0.01 | 12 / 11 | 0 / 7,178 | |
| GPT-6-sol | ACL 2019 | 0.03 / 0.02 | 10 / 10 | 0 / 4,455 |
| arXiv 2026 | 0.02 / 0.01 | 12 / 11 | 0 / 4,930 |
| Claude | GPT-6-sol | Both | |
|---|---|---|---|
| Uses coded as disavowals (of 947) | 147 | 148 | 129 |
| , disavowal against the rest | 0.85 | ||
| , four classes | 0.73 | ||
| Codes | Disav. | with A1 | Annoying | AUC |
|---|---|---|---|---|
| A1 | 36 | – | 26/36 (72) | 0.83 |
| Claude | 39 | 0.47 | 22/39 (56) | 0.70 |
| GPT-6-sol | 41 | 0.38 | 23/41 (56) | 0.70 |
| Both coders | 25 | 0.51 | 18/25 (72) | 0.73 |
| Claude, broader definition | 47 | 0.53 | 26/47 (55) | 0.72 |
| GPT-6-sol, broader definition | 52 | 0.51 | 28/52 (54) | 0.73 |
| Instructions | Claude | GPT-6-sol |
|---|---|---|
| Original | 0.56 | 0.32 |
| Broader definition | 0.61 | 0.54 |
| Broader, with 30 of A1’s codes as examples | 0.56 | 0.58 |
| Disavowal | Other | ||||
|---|---|---|---|---|---|
| Stance rater | n | Gap | n | Gap | |
| GPT-6-sol | 129 | 1.11 | 818 | 1.29 | 0.07 |
| Claude agents | 81 | 0.86 | 468 | 1.21 | 0.09 |
| Change per 1k words | RT sentences | |||
|---|---|---|---|---|
| Model | RT | X, not Y | removed | added |
| Haiku 4.5 | 0.16 (0.007) | 0.01 (0.262) | 27/399 (0) | 42/574 (3) |
| Sonnet 5.5 | 0.01 (0.701) | 0.23 (<.001) | 39/695 (8) | 26/857 (4) |
| Opus 5.5 | 0.16 (<.001) | 0.25 (<.001) | 92/637 (21) | 55/927 (14) |
| v1 | Latest | |
| Mean rate per 1,000 words | 0.921 | 0.988 |
| Total occurrences | 5,394 | 6,471 |
| Median body length (words) | 6,747 | 7,572 |
| Paired difference | 0.067 | |
| 95% CI | 0.042–0.093 | |
| Wilcoxon | ||
| Quantity | Count |
|---|---|
| Paper pairs (v1 + latest) | 807 |
| Mined diff hunks with phrase | 3,550 |
| phrase added | 1,156 |
| phrase kept, context reworded | 1,818 |
| phrase removed from hunk | 576 |
| surgical removal | 86 |
| Type | Pattern | 2019 | 2026 | v1 | latest |
|---|---|---|---|---|---|
| Substitution | rather than | 0.134 | 0.993 | 0.921 | 0.988 |
| instead of | 0.165 | 0.105 | 0.109 | 0.109 | |
| as opposed to | 0.025 | 0.007 | 0.007 | 0.006 | |
| Negation | not X but Y | 0.097 | 0.126 | 0.129 | 0.125 |
| X, not Y | 0.049 | 0.201 | 0.188 | 0.202 | |
| Emphatic | not only…but | 0.079 | 0.119 | 0.124 | 0.116 |
| Quantity | Count |
|---|---|
| Candidate repositories | 68 |
| one commit to the paper | 36 |
| five or more commits | 10 |
| Diff hunks with the phrase (10 repositories) | 95 |
| phrase added | 24 |
| phrase kept, context reworded | 55 |
| Paper (arXiv id) | Uses | Per 1k words |
|---|---|---|
| 2608.13706 | 69 | 5.84 |
| 2608.06171 | 63 | 4.61 |
| 2606.05743 | 54 | 3.28 |
| 2608.29995 | 46 | 2.01 |
| 2607.14252 | 46 | 2.18 |
| 2604.17022 | 46 | 3.88 |
| Stratum | N | A1 | A2 | Either |
|---|---|---|---|---|
| arXiv 2026, main | 388 | 51 (13.1) | 39 (10.1) | 79 (20.5) |
| v1 | 145 | 14 (9.7) | 9 (6.2) | 22 (15.3) |
| latest | 119 | 10 (8.5) | 6 (5.0) | 15 (12.7) |
| Body text only (without 33 items) | ||||
| ACL 2019 | 245 | 1.7 | 0.0 | |
| arXiv 2026, all | 627 | 11.8 | 8.0 | |
| A1 | A2 | |||
| Measure | Batch 1 | Batch 2 | Batch 1 | Batch 2 |
| Annoying , ACL 2019 | 0 / 143 | 4 / 100 | 0 / 144 | 0 / 99 |
| Annoying , arXiv 2026 | 11.5% | 11.6% | 5.6% | 10.0% |
| Gap X Y, annoying / legitimate | 2.31 / 0.78 ∗ | 2.33 / 0.83 ∗ | 1.79 / 0.91 | 1.43 / 0.90 |
| Y, annoying / legitimate | 1.03 / 0.31 ∗ | 1.00 / 0.34 ∗ | 0.71 / 0.39 | 0.65 / 0.37 |
| Initial Rather than (%) | 12 / 4 ∗ | 17 / 2 ∗ | 8 / 3 | 8 / 3 |
| A1 | A2 | |
| Hidden repeats (80 items) | ||
| Same label on both passes | 61 (76%) | 67 (84%) |
| legitimate unsure | 7 | 3 |
| unsure legitimate | 4 | 8 |
| other annoying | 3 | 2 |
| annoying other | 4 | 0 |
| A1–A2 | A1–A5 | A2–A5 | |
|---|---|---|---|
| Raw agreement | 76% | 26% | 20% |
| Cohen’s | 0.15 | 0.00 | 0.03 |
| Annoying : 1st / 2nd / both | 4 / 1 / 0 | 4 / 35 / 3 | 1 / 35 / 0 |
| Core NLP | Other | OR | ||
|---|---|---|---|---|
| A1 | 69/603 (11%) | 18/132 (14%) | 1.2 | .46 |
| A2 | 46/602 (8%) | 19/132 (14%) | 2.1 | .018 |
| Either | 102/601 (17%) | 31/132 (23%) | 1.5 | .082 |
| Shown text | A1 | A2 | ||
|---|---|---|---|---|
| Another rather than | 9 / 12 | .41 | 7 / 9 | .57 |
| Other antithesis | 17 / 12 | .32 | 10 / 9 | .78 |
| Graphite tell | 11 / 12 | 1.0 | 9 / 9 | .85 |
| Kobak words /100 | 5.1 / 4.8 | .10 | 4.4 / 4.8 | .80 |
| Whole paper | ||||
| Rather than /1k | 1.48 / 1.56 | .46 | 1.61 / 1.53 | .013 |
| Papers | Same-paper pairs | ||
|---|---|---|---|
| A1 | 129 | 11 vs. 8.6 | .25 |
| A2 | 129 | 24 vs. 6.4 | .001 |
| A2 annoying , by A1’s codes elsewhere in the paper | |||
| straw man / none | 12/64 vs. 8/105 | .047 | |
| controlled | OR 1.5 | .51 | |
| Term | A1 | A2 | Either | |||
|---|---|---|---|---|---|---|
| Items | 87 / | 598 | 66 / | 582 | 134 / | 486 |
| intended | 8 / | 1 | 0 / | 1 | 5 / | 1 |
| rather than relying | 8 / | 2 | 2 / | 2 | 5 / | 2 |
| we | 34 / | 18 | 21 / | 21 | 29 / | 19 |
| so | 10 / | 8 | 17 / | 7 | 11 / | 7 |
| rather than the | 3 / | 5 | 0 / | 5 | 2 / | 5 |
| A1 | A2 | Either | |
| Similarity to blatant straw men, AUC | |||
| Y, straw-man codes | .67 ∗ | ||
| Y, annoying | .66 ∗ | .53 | .61 ∗ |
| clause, annoying | .64 ∗ | .61 ∗ | .63 ∗ |
| Annoying in the next five items, % | |||
| after / otherwise | 9 / 11 | 5 / 9 | |
| Source | Rather than | X-ed/Y-ed | Share |
|---|---|---|---|
| ACL 2019 | 3,087 | 4 | 0.13% |
| arXiv 2026 | 14,539 | 82 | 0.56% |
| GPT-4o | 104 | 1 | 0.96% |
| revised | 73 | 1 | 1.4% |
| GPT-5.6-sol | 5,878 | 74 | 1.3% |
| revised | 7,636 | 115 | 1.5% |
| Feature | A1 | A2 | Either | |||
| Items | 56 / | 251 | 27 / | 293 | 72 / | 221 |
| Subject we | 29 / | 9 ∗ | 11 / | 12 | 22 / | 10 ∗ |
| X adjective | 12 / | 4 | 11 / | 3 | 11 / | 2 ∗ |
| X adj. modifier | 2 / | 0 | 0 / | 0 | 1 / | 0 |
| X prep. object | 27 / | 21 | 15 / | 24 | 22 / | 22 |
| than coord. adj. | 12 / | 3 ∗ | 7 / | 3 | 10 / | 2 ∗ |
| Items | Unit | A1 | A2 | Either |
|---|---|---|---|---|
| arXiv 2026 | Items | 32 / 121 | 9 / 153 | 36 / 111 |
| Chance | 20% | 5% | 24% | |
| X | 31% ∗ | 9% | 34% ∗ | |
| Y | 41% ∗ | 24% ∗ | 42% ∗ | |
| Sentence | 41% ∗ | 0% | 44% ∗ | |
| Paragraph | 32% ∗ | 4% | 33% |
| Hunches | Labels | Papers | Passages |
|---|---|---|---|
| A1 | A1 | 0.01 | 0.06 |
| A1 | A2 | 0.21 | 0.26 |
| A2 | A2 | 0.18 | 0.16 |
| A2 | A1 | 0.02 | 0.01 |
| arXiv 2026 | n | A1 | A2 |
| with rather than | 52 | 73% | 50% |
| GPT-6-sol | Favors X | Favors Y | Neither |
|---|---|---|---|
| Inverted | 1 | 24 | 15 |
| Not inverted | 20 | 2 | 18 |
| Spearman | Within 1 | Mean diff. | ||
|---|---|---|---|---|
| X (affirmed) | 0.77 | 0.76 | 98% | 0.12 |
| Y (rejected) | 0.72 | 0.73 | 97% | 0.02 |
| Gap X Y | 0.81 | 0.80 | 81% | 0.14 |
| Inverted (X Y), | 0.71 |
| Uses | GPT-6-sol | Claude | |
|---|---|---|---|
| ACL 2019 originals | 50 | 1.22 | 1.04 |
| Claude, first draft | 85 | 0.81 | 0.86 |
| Claude, revised | 75 | 0.81 | 0.61 |
| GPT-6-sol, first draft | 50 | 1.88 | 1.52 |
| GPT-6-sol, revised | 50 | 1.84 | 1.54 |
| Set | Rater | Gap: ann. / other | AUC | |
|---|---|---|---|---|
| Pool sample | GPT-6-sol | 1.99 / 0.81 | 0.65 | .001 |
| Claude | 1.79 / 0.81 | 0.65 | .001 | |
| mean | 1.89 / 0.81 | 0.66 | .001 | |
| Claude set (A1) | GPT-6-sol | 1.42 / 1.18 | 0.53 | .46 |
| Claude | 1.40 / 0.98 | 0.58 | .08 | |
| mean | 1.41 / 1.08 | 0.55 | .24 |
| A1 | A2 | A3 | A4 | |
|---|---|---|---|---|
| Straw men (of 151) | 42 | 7 | 40 | 12 |
| with A2 | 0.18 | |||
| with A3 | 0.16 | 0.10 | ||
| with A4 | 0.11 | 0.27 | 0.12 |
| Revised | Ctrl. | Position | Earlier | |
|---|---|---|---|---|
| Claude | 0.94 (0.03) | 0.90 (0.05) | 0.35 (0.63) | 0.12 (0.26) |
| GPT-6-sol | 0.51 (0.32) | 0.35 (0.50) | 1.86 (0.15) | 0.13 (0.45) |