From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection
Authors: Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early, Steve Graham, Danielle S. McNamara
Organizations: Learning Engineering Institute, Arizona State University, Tempe, Arizona, USA · The Polytechnic School, Arizona State University, Tempe, Arizona, USA · Department of English, Arizona State University, Tempe, Arizona, USA · Mary Lou Fulton College for Teaching and Learning Innovation, Arizona State University, Tempe, Arizona, USA
Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic units. This study evaluates whether LLMs can identify meaningful revision unit boundaries in structured revision operation data and whether they provide value beyond simple non-LLM baselines. Using 113 matched draft--revision pairs from undergraduate writing, expert annotation yielded 4,344 candidate boundaries. We compared zero- and few-shot GPT-5.5 and base Qwen3-32B, parameter-efficient fine-tuning of Qwen3-32B, and majority and proximity-based baselines. Despite receiving revision context and task instructions, no prompted LLM condition outperformed the proximity heuristic (macro-F1 = .825). In contrast, fine-tuned Qwen3-32B using the two context representation achieved the highest macro-F1 (.859), identifying more same-unit relationships while maintaining precision comparable to the heuristic. Deterministic post-processing substantially improved the prompted models but added little benefit to the strongest fine-tuned model. These findings suggest that LLMs can support revision boundary judgment when task-adapted, but general purpose prompting alone may not outperform transparent structural heuristics.
Figures & tables
Operation
Count
Preserve
2,390
Replace
1,015
Insert
646
Delete
405
Move
1
Total
4,457
Table 1: Instance Counts for TRACE Revision Operations
Figure 1: Three examples of MRUs identified in the dataset, illustrating typical variation in their scale and internal structure.
Figure 2: Example transformation of annotated TRACE revision operations into a boundary classification instance using a ±1 context window. The focal pair consists of the two adjacent operations to be classified, with one preceding and one following operation included as local context.
Partition
Draft pairs
False
True
Total
Training
90
2,930
602
3,532
Validation
11
315
56
371
Test
12
364
77
441
Total
113
3,609
735
4,344
Table 2: Distribution of Draft Pairs and Boundary Instances Across Data Partitions
Model
Prompting condition
Context window
Accuracy
Precision
Recall
F1
Macro-F1
Majority baseline
–
–
.825
.000
.000
.000
.452
Proximity heuristic
–
–
.912
.852
.597
.702
.825
GPT-5.5
Zero-shot
±1
.522
.246
.844
.381
.496
GPT-5.5
Few-shot
±1
.626
.292
.805
.429
.575
GPT-5.5
Zero-shot
±2
.574
.273
.870
.416
.540
GPT-5.5
Few-shot
±2
.628
.288
.766
.418
.573
Table 3: Baseline and LLM Performance Before Post-Processing
Draft-verify-revise is a common LLM orchestration pattern for scaling inference-time compute. One LLM drafts, a second critiques the draft and provides feedback, and a third uses that feedback to revise the draft into the final output. As context cascades between stages, LLMs at different stages can resolve a context-dependent expression such as "previous" differently. When that happens, the expression undergoes a deictic shift, a change in what it refers to. This phenomenon was studied with a synthetic dataset of 10 base examples, each rendered in three conditions. Holding the shared components constant, the conditions varied whether the draft stage LLM (the assistant) or the verify stage LLM (the grader) resolved the expression correctly, and how much independent reasoning the revise stage LLM (the meta-evaluator) needed to determine which reading was correct. Six models from three providers were tested across 21 reasoning effort configurations using e-values for sequential testing, in a primary experiment and an ablation experiment that removed error classification labels from the grader's feedback. A separate LLM analyzed the meta-evaluator's stated rationale for each wrong verdict. Balanced accuracy (the unweighted mean of sensitivity and specificity) ranged from 0.156, below chance, to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest reasoning effort level, while Gemini 3 Pro stayed above 0.94 at every level. Gemini 3 Pro at low reasoning effort outscored GPT-5.2 at xhigh reasoning effort for roughly 5% of the cost per trial. When the meta-evaluator erred, it tended to rely on surface cues rather than operational reasoning. Context engineers implementing draft-verify-revise pipelines should be wary of deictic shifts and make the intended referent explicit at each stage.
Obinna I. Ekekezie
Cambridge Health Alliance, Cambridge, MA, United States of America. · Harvard Medical School, Boston, MA, United States of America.
We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems. Draft confidence estimates whether the current answer is correct, but the decision requires estimating the effect of a specified revision. For offline training and evaluation, we grade both the returned draft and its candidate revision under the same correctness judge, which makes repair, harm, and the gap to an oracle observable. We call this paired effect its recoverability, and we train policies to predict it before revision. On 25,870 held-out open-domain questions across three revision setups, a scorer trained on the paired outcome has greater area under the accuracy--revision-rate curve than a matched draft-correctness scorer in all nine Llama setup--seed fits, and gains 0.23--0.68 accuracy points on average at development-selected thresholds, a difference significant across training runs only for dense retrieval. The resulting policy improves on always revising and on average closes more than a third of the oracle gap, although it still applies 38--46% of the harmful revisions. When a draft-free standard-RAG answer is also available, however, choosing between the draft and that answer is stronger by about two points for Llama and four for OLMo, and adding candidate revision as a third option yields no significant gain. Recoverability describes one revision; its value as an available action also depends on the alternatives.
Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup +1
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark. The results show that baselines achieve accuracies of 68.3--93%, and the most cost-effective method is selecting from three parallel samples using either LLM-based or medoid selection, which improves accuracy by 2.2--9.7%. Our code and dataset are available at https://github.com/ntt-dkiku/llm-revision-propagation.