cs.CLSep 18, 2026

Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

Authors: Ruotian Wu, Bill E. Johnson, Gene Saunders, Osama Hamzeh, Ankit Vadehra, Pascal Poupart

Organizations: University of Waterloo · Vector Institute · Scribendi Inc.

Abstract

Reference-based metrics for Grammatical Error Correction (GEC) such as M2^2 and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.

Explore similar work

CardsList
  1. Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

    Sep 14, 2026Hayeong Ryu, Sunhee Jo, Seunguk Yu +1Reward ModelingReference-Free Evaluation

  2. Multi-Dimensional Evaluation of LLMs for Grammatical Error Correction

    May 8, 2026Adnan Labib, Qiao Wang, Yixuan Huang +1LLM EvaluationGrammatical Error Correction

  3. Edit-level Majority Voting Mitigates Over-Correction in LLM-based Grammatical Error Correction

    May 13, 2026Takumi Goto, Yusuke Sakai, Taro WatanabeSelf-Consistency DecodingGrammatical Error Correction