cs.CLSep 27, 2026

From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection

Authors: Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early, Steve Graham, Danielle S. McNamara

Organizations: Learning Engineering Institute, Arizona State University, Tempe, Arizona, USA · The Polytechnic School, Arizona State University, Tempe, Arizona, USA · Department of English, Arizona State University, Tempe, Arizona, USA · Mary Lou Fulton College for Teaching and Learning Innovation, Arizona State University, Tempe, Arizona, USA

Abstract

Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic units. This study evaluates whether LLMs can identify meaningful revision unit boundaries in structured revision operation data and whether they provide value beyond simple non-LLM baselines. Using 113 matched draft--revision pairs from undergraduate writing, expert annotation yielded 4,344 candidate boundaries. We compared zero- and few-shot GPT-5.5 and base Qwen3-32B, parameter-efficient fine-tuning of Qwen3-32B, and majority and proximity-based baselines. Despite receiving revision context and task instructions, no prompted LLM condition outperformed the proximity heuristic (macro-F1 = .825). In contrast, fine-tuned Qwen3-32B using the two context representation achieved the highest macro-F1 (.859), identifying more same-unit relationships while maintaining precision comparable to the heuristic. Deterministic post-processing substantially improved the prompted models but added little benefit to the strongest fine-tuned model. These findings suggest that LLMs can support revision boundary judgment when task-adapted, but general purpose prompting alone may not outperform transparent structural heuristics.

Figures & tables

Explore similar work

CardsList
  1. Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

    Sep 14, 2026Obinna I. EkekezieLarge Language Model WorkflowsAmbiguity

  2. Return or Revise? Learning When Revision Helps Retrieval-Augmented QA

    Sep 24, 2026Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup +1Confidence Calibration