Organizations: School of Information Science and Technology, Guangdong University of Foreign Studies · Guangdong Engineering Research Center of Data Security Governance and Privacy Computing · College of Computer Science and Technology, National University of Defense Technology
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.
Figures & tables
Figure 1: Comparison of w / CoT and w/o CoT models. We train the same model twice separately: once on CoT data and once on non-CoT data.
Figure 2: Overview of VAA-CSEC. Left: a CoT distillation pipeline pairs a Qwen3.5-27B generator with a DeepSeek-V3.2 verifier to produce structured data for SFT. Center: GLPO rolls out N candidates from the policy πθ , scores them with a reward combining format Rf , edit-distance correctness Rc , and an over-correction penalty ρ , and reallocates advantages using the margin between each rollout’s reward and the vote-aggregated group reward. Right: at inference, N stochastic generations are aggregated by majority voting to yield the final correction.
Case
Rc
Exact match
+3.4
Effective & minimal
2.0⋅F0.5
Effective & overcorrects
2.0⋅F0.5−0.8⋅min(ρ,2)
No effective edit
0.0
Degrades source
−1.5⋅min(−R^,1)
Table 1: Five branches of the correctness reward Rc , clipped to [−1.5,3.4] . The total reward used during training is R=Rf+Rc , where Rf=0.3 if the output strictly contains exactly one <think> block and one <answer> block, and 0 otherwise.
Figure 3: Performance comparison of F0.5 scores for SFT and GRPO models across different decoding strategies. SFT_CoT denotes the model fine-tuned via SFT on CoT data, while GRPO_CoT refers to the GRPO-optimized model initialized from the SFT checkpoint.
Dataset
Usage
Train
Val
Test
Before
After
CSED-C
non-CoT
8,682
11,338
1,000
1,000
CoT
11,204
NaSGEC-Exam
non-CoT
4,000
5,818
1,000
2,000
CoT
5,406
Table 2: Dataset statistics across training stages. Only the training set is processed. ‘Before’ reports the raw sentence count, while ‘After’ reports the count after task-specific preprocessing.
Dataset
Method
Model
Decoding
P
R
F0.5
CSED-C
Seq2Seq
mT5-small
–
33.70
5.40
16.50
mT5-base
57.00
19.00
40.70
BART-large
53.80
38.30
49.70
SynGEC
53.00
39.50
49.60
CSEC-LLM
ChatGLM3-6B
–
33.18
27.80
31.94
Baichuan2-7B
35.90
32.95
35.26
Table 3: Main results on CSED-C and NaSGEC-Exam. Underlined bold indicates the best score per metric across all methods within each dataset; bold indicates the best score per metric among the remaining methods of the same method.
VAA-CSEC
P
R
F0.5
Time cost
ΔF0.5
Greedy
39.19
35.50
38.39
∼ 2 min
–
4-vote
39.51
36.02
38.76
∼ 8 min
+0.37
8-vote
42.00
37.67
41.06
∼ 16 min
+2.30
16-vote
47.55
40.58
45.97
∼ 32 min
+4.91
32-vote
49.34
42.15
47.72
∼ 64 min
+1.75
Table 4: Cost-performance trade-off of VAA-CSEC on CSED-C.
Method
P
R
F0.5
VAA-CSEC
49.34
42.15
47.72
w / o CoT
44.27
39.54
43.23
w / o GLPO
48.71
40.88
46.91
w / o CoT and GLPO
44.66
39.99
43.64
Table 5: Ablation study of the proposed components on the CSEC task. “ w / o ” denotes removing the corresponding module from the full model.
Train Set
Test Set
Decoding
P
R
F0.5
CSED-C
CSED-C
greedy
39.19
35.50
38.39
+32 Vote
49.34
42.15
47.72
NaSGEC-Exam
Greedy
27.06
36.46
28.53
+32 Vote
35.22
42.14
36.42
NaSGEC-Exam
CSED-C
Greedy
35.65
26.46
33.33
+32 Vote
50.54
27.88
43.47
Table 6: Cross-domain evaluation results. Each model is trained on one dataset and evaluated on both test sets.
Figure 4: Three categories of Chinese text error correction, illustrating a progression from surface-level character errors to deep semantic errors. Errors in the source sentences are highlighted in red, and the corresponding edits in the target sentences are shown in green.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Template for Data Processing
Figure 6: Training curves of GLPO during reinforcement learning: (a) reward, (b) policy entropy, and (c) generated sequence length.
Figure 7: Training curves of GRPO during reinforcement learning: (a) reward, (b) policy entropy, and (c) generated sequence length.
Figure 8: Two examples of case study
Ratio
Mode
Decoding
P
R
F0.5
non-CoT
w/o thinking
Greedy
48.03
37.37
45.44
+32 vote
48.07
38.19
45.71
CoT
w/ thinking
Greedy
37.28
35.80
36.97
+32 vote
47.73
41.55
46.35
1:1 mix
w/o thinking
Greedy
44.27
38.94
43.09
+32 vote
45.48
40.66
44.43
Appendix
Table 7: Performance comparison of models trained with different CoT to non-CoT data ratios.
N
P
R
F0.5
Agree.
4
48.89
42.68
47.50
47.6%
6
47.88
40.51
46.20
49.9%
8
49.34
42.15
47.72
47.7 %
Appendix
Table 8: Parameter study: effect of the number of rollouts N on CSED-C. “Agree.” reports the average agreement rate of 32-vote inference using the corresponding trained policy.
Grammatical error correction using large language models often suffers from the over-correction issue. To mitigate this, we propose a training-free inference method that performs edit-level majority voting over multiple candidates generated by a single model, without requiring model modifications or additional training. Across nine benchmarks covering English, Czech, German, Ukrainian, Korean, Hindi, and Romanian, the proposed method outperforms both greedy and MBR decoding in most cases. Moreover, it yields stable correction quality regardless of the instruction prompts used. We release two repository supporting GEC datasets loading and LLM inference.
Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade F0.5 by rewriting well-formed spans. While fine-tuning provides an effective solution, it imposes substantial infrastructure demands. We introduce a prompt-based approach that closes the gap to fine-tuned models through three advances in GEC prompting methodology. First, we introduce taxonomy-based instructions to enforce minimal-edit constraints with a comprehensive list of grammatical error rules, equipping the LLM with a bounded, metric-aligned scope of correctable edits, which benefits the strongest models while remaining model-dependent overall. Second, we show that batching multiple uncorrected sentences into a single input context acts as a targeted regularizer against overcorrection, systematically reducing the edit rate across diverse LLM families; we hypothesize this arises from attention dilution effect induced by the bounded capacity of self-attention scores. Finally, LLM-assisted Prompt Optimization refines these instructions. Powered by Gemini 3.1-Pro, our prompt achieves F0.5=78.32 on the BEA-2019 test set - establishing a new prompt-based SOTA while shrinking the gap to the fine-tuned single-model SOTA (Staruch et al., 2025) to a mere 0.38 points. Code, prompts, and outputs are publicly available.
Large language model (LLM) self-correction -- the ability to detect and fix errors in generated outputs -- remains largely ad hoc, relying on generic prompts such as "please reconsider your answer" without systematic error analysis or convergence guarantees. We propose CyberCorrect, a framework that formalizes LLM self-correction as a closed-loop control system grounded in cybernetic theory. The framework models the LLM generator as the plant and introduces a tri-modal Error Detector (combining self-consistency, verbalized confidence, and logic-chain verification) as the sensor. A type-directed Correction Controller generates targeted repair instructions based on diagnosed error categories, while a Convergence Judge determines iteration termination using stability criteria adapted from control theory. We further introduce three control-theoretic evaluation metrics -- convergence rate, overshoot rate, and oscillation rate -- that capture correction dynamics beyond final accuracy. Experiments on our constructed CyberCorrect-Bench (440 reasoning tasks with annotated error types and correction paths) show that CyberCorrect achieves 79.8% final accuracy, improving upon the best existing self-correction method by 6.2 percentage points, while reducing overshoot (erroneous over-correction) by 41% through its convergence control mechanism.
Yuning Wu, Yingmin Liu, Yang Shu
School of Software, Henan University, Kaifeng, China · 2Zhejiang University, Hangzhou, China