Does AI-Generated Scientific Text Follow Human Argumentation Patterns? A CARS-Based Comparison of Research Article Introductions
Organizations: Università della Svizzera italiana (USI)
Abstract
Large language models are moving from helping write up research to helping do it, which makes it important to know how the scientific text they produce differs from human writing. Work on this question has stayed mostly at the surface, using lexical and stylistic cues that light paraphrasing erases. We look instead at rhetorical structure, the sequence of argumentative moves through which a text makes its case. We study research-article introductions under Swales' CARS model, and compare original introductions from published linguistics articles with generated counterparts of the same papers. We find that human-written introductions are more flexible in which moves they use and in what order, while the generated ones are more uniform. Giving the models the CARS definitions makes them more rigid.
Figures & tables
| Original | ChatGPT b | ChatGPT c | Claude b | Claude c | Gemma b | Gemma c | Qwen b | Qwen c | |
| Move 1: Establishing a Territory | 59.5 | 70.1* | 62.9 | 57.5 | 51.1* | 60.6 | 46.8* | 54.4 | 48.3* |
| Claiming centrality (CC) | 18.5 | 21.9* | 20.4 | 19.2 | 17.3 | 19.9 | 17.1 | 16.0 | 15.8 |
| Topic generalization (MTG) | 14.5 | 22.3* | 19.3* | 14.0 | 11.5 | 15.7 | 9.9* | 13.0 | 10.0* |
| Reviewing prev. research (RR) | 26.5 | 25.9 | 23.1* | 24.3 | 22.3* | 25.0 | 19.9* | 25.4 | 22.5* |
| Move 2: Establishing a Niche | 16.0 | 14.6 | 17.9* | 19.9* | 22.8* | 21.1* | 22.1* | 22.3* | 26.6* |
| Continuing a tradition (CT) | 3.6 | 3.8 | 4.3 | 3.8 | 4.1 | 6.6* | 6.0* | 4.6 | 4.3 |
| Source | Top step 3-gram (%) | Cyclic % | Uniq. moves | Uniq. steps |
| Original | RR MTG RR (25) | 66 | 2.67 | 5.3 |
| ChatGPT b | CC MTG RR (57) | 70 | 2.93* | 6.3* |
| ChatGPT c | CC MTG RR (61) | 63 | 3.00* | 6.9* |
| Claude b | RR CC RR (33) | 74 | 2.95* | 6.2* |
| Claude c | CC RR Gap (42) | 51 | 3.00* | 6.6* |
| Gemma b | RR CC RR (37) | 67 | 2.96* | 6.1* |
| Similarity to the original | Within source | |||||
| Source | Edit m | Edit s | Bag m | Bag s | Self s | Str m |
| Original | – | – | – | – | 0.33 | 36 |
| ChatGPT b | 0.62 | 0.41 | 0.83 | 0.63 | 0.43 | 18 |
| ChatGPT c | 0.61 | 0.41 | 0.81 | 0.62 | 0.44 | 9 |
| Claude b | 0.61 | 0.42 | 0.81 | 0.63 | 0.40 | 24 |
| Claude c | 0.61 | 0.43 | 0.80 | 0.61 | 0.46 | 11 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Move | Steps |
| M1 : Establishing a Territory | Claiming centrality (CC); Topic generalization (MTG); Reviewing previous research (RR) |
| M2 : Establishing a Niche | Continuing a tradition (CT); Indicating a gap (Gap); Question raising (QR) |
| M3 : Occupying the Niche | Announcing present research (APR); Announcing principal findings (APF); Indicating structure (IS); Outlining purpose (OP) |
| Classification | Segmentation | |||
| Annotator | step-acc | step-F1 | move-acc | bound-F1 |
| HB | 0.752 | 0.699 | 0.911 | 0.756 |
| AA | 0.694 | 0.664 | 0.901 | 0.725 |
| Class | Human | Annotator |
| M1-CC | 0.70 | 0.46 |
| M1-MTG | 0.63 | 0.55 |
| M1-RR | 0.86 | 0.84 |
| M2-CT | 0.23 | 0.32 |
| M2-Gap | 0.77 | 0.77 |
| M2-QR | 0.42 | 0.35 |
| Source | Words /intro | Segments /intro | Words /segment |
| Original (human) | 615 | 9.4 | 65.1 |
| ChatGPT, basic | 743 | 13.9 | 53.5 |
| ChatGPT, CARS | 745 | 13.7 | 54.4 |
| Claude, basic | 655 | 11.4 | 57.3 |
| Claude, CARS | 653 | 9.9 | 66.1 |
| Gemma, basic | 563 | 10.5 | 53.7 |
| Source | Uniq. moves | Uniq. steps | Cyclic % | |
| Q1 | Original | 2.88 | 5.0 | 75 |
| Generated | 2.99 | 6.1 | 73 | |
| Q2 | Original | 2.65 | 5.7 | 70 |
| Generated | 2.97 | 7.2 | 74 | |
| Q3 | Original | 2.50 | 5.3 | 55 |
| Generated | 2.97 | 6.7 | 66 |
| Condition | Items | Correct |
| All items | 120 | 108 ( ) |
| By move annotation (all items) | ||
| Labelled | 60 | 54 ( ) |
| Unlabelled | 60 | 54 ( ) |
| Generated items, by prompting setting | ||
| Basic | 30 | 24 ( ) |