From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection
Authors: Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early, Steve Graham, Danielle S. McNamara
Organizations: Learning Engineering Institute, Arizona State University, Tempe, Arizona, USA · The Polytechnic School, Arizona State University, Tempe, Arizona, USA · Department of English, Arizona State University, Tempe, Arizona, USA · Mary Lou Fulton College for Teaching and Learning Innovation, Arizona State University, Tempe, Arizona, USA
Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic units. This study evaluates whether LLMs can identify meaningful revision unit boundaries in structured revision operation data and whether they provide value beyond simple non-LLM baselines. Using 113 matched draft--revision pairs from undergraduate writing, expert annotation yielded 4,344 candidate boundaries. We compared zero- and few-shot GPT-5.5 and base Qwen3-32B, parameter-efficient fine-tuning of Qwen3-32B, and majority and proximity-based baselines. Despite receiving revision context and task instructions, no prompted LLM condition outperformed the proximity heuristic (macro-F1 = .825). In contrast, fine-tuned Qwen3-32B using the two context representation achieved the highest macro-F1 (.859), identifying more same-unit relationships while maintaining precision comparable to the heuristic. Deterministic post-processing substantially improved the prompted models but added little benefit to the strongest fine-tuned model. These findings suggest that LLMs can support revision boundary judgment when task-adapted, but general purpose prompting alone may not outperform transparent structural heuristics.
Figures & tables
Operation
Count
Preserve
2,390
Replace
1,015
Insert
646
Delete
405
Move
1
Total
4,457
Table 1: Instance Counts for TRACE Revision Operations
Figure 1: Three examples of MRUs identified in the dataset, illustrating typical variation in their scale and internal structure.
Figure 2: Example transformation of annotated TRACE revision operations into a boundary classification instance using a ±1 context window. The focal pair consists of the two adjacent operations to be classified, with one preceding and one following operation included as local context.
Partition
Draft pairs
False
True
Total
Training
90
2,930
602
3,532
Validation
11
315
56
371
Test
12
364
77
441
Total
113
3,609
735
4,344
Table 2: Distribution of Draft Pairs and Boundary Instances Across Data Partitions
Model
Prompting condition
Context window
Accuracy
Precision
Recall
F1
Macro-F1
Majority baseline
–
–
.825
.000
.000
.000
.452
Proximity heuristic
–
–
.912
.852
.597
.702
.825
GPT-5.5
Zero-shot
±1
.522
.246
.844
.381
.496
GPT-5.5
Few-shot
±1
.626
.292
.805
.429
.575
GPT-5.5
Zero-shot
±2
.574
.273
.870
.416
.540
GPT-5.5
Few-shot
±2
.628
.288
.766
.418
.573
Table 3: Baseline and LLM Performance Before Post-Processing