Organizations: School of Information Science and Technology, Guangdong University of Foreign Studies · Guangdong Engineering Research Center of Data Security Governance and Privacy Computing · College of Computer Science and Technology, National University of Defense Technology
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.
Figures & tables
Figure 1: Comparison of w / CoT and w/o CoT models. We train the same model twice separately: once on CoT data and once on non-CoT data.
Figure 2: Overview of VAA-CSEC. Left: a CoT distillation pipeline pairs a Qwen3.5-27B generator with a DeepSeek-V3.2 verifier to produce structured data for SFT. Center: GLPO rolls out N candidates from the policy πθ , scores them with a reward combining format Rf , edit-distance correctness Rc , and an over-correction penalty ρ , and reallocates advantages using the margin between each rollout’s reward and the vote-aggregated group reward. Right: at inference, N stochastic generations are aggregated by majority voting to yield the final correction.
Case
Rc
Exact match
+3.4
Effective & minimal
2.0⋅F0.5
Effective & overcorrects
2.0⋅F0.5−0.8⋅min(ρ,2)
No effective edit
0.0
Degrades source
−1.5⋅min(−R^,1)
Table 1: Five branches of the correctness reward Rc , clipped to [−1.5,3.4] . The total reward used during training is R=Rf+Rc , where Rf=0.3 if the output strictly contains exactly one <think> block and one <answer> block, and 0 otherwise.
Figure 3: Performance comparison of F0.5 scores for SFT and GRPO models across different decoding strategies. SFT_CoT denotes the model fine-tuned via SFT on CoT data, while GRPO_CoT refers to the GRPO-optimized model initialized from the SFT checkpoint.
Dataset
Usage
Train
Val
Test
Before
After
CSED-C
non-CoT
8,682
11,338
1,000
1,000
CoT
11,204
NaSGEC-Exam
non-CoT
4,000
5,818
1,000
2,000
CoT
5,406
Table 2: Dataset statistics across training stages. Only the training set is processed. ‘Before’ reports the raw sentence count, while ‘After’ reports the count after task-specific preprocessing.
Dataset
Method
Model
Decoding
P
R
F0.5
CSED-C
Seq2Seq
mT5-small
–
33.70
5.40
16.50
mT5-base
57.00
19.00
40.70
BART-large
53.80
38.30
49.70
SynGEC
53.00
39.50
49.60
CSEC-LLM
ChatGLM3-6B
–
33.18
27.80
31.94
Baichuan2-7B
35.90
32.95
35.26
Table 3: Main results on CSED-C and NaSGEC-Exam. Underlined bold indicates the best score per metric across all methods within each dataset; bold indicates the best score per metric among the remaining methods of the same method.
VAA-CSEC
P
R
F0.5
Time cost
ΔF0.5
Greedy
39.19
35.50
38.39
∼ 2 min
–
4-vote
39.51
36.02
38.76
∼ 8 min
+0.37
8-vote
42.00
37.67
41.06
∼ 16 min
+2.30
16-vote
47.55
40.58
45.97
∼ 32 min
+4.91
32-vote
49.34
42.15
47.72
∼ 64 min
+1.75
Table 4: Cost-performance trade-off of VAA-CSEC on CSED-C.
Method
P
R
F0.5
VAA-CSEC
49.34
42.15
47.72
w / o CoT
44.27
39.54
43.23
w / o GLPO
48.71
40.88
46.91
w / o CoT and GLPO
44.66
39.99
43.64
Table 5: Ablation study of the proposed components on the CSEC task. “ w / o ” denotes removing the corresponding module from the full model.
Train Set
Test Set
Decoding
P
R
F0.5
CSED-C
CSED-C
greedy
39.19
35.50
38.39
+32 Vote
49.34
42.15
47.72
NaSGEC-Exam
Greedy
27.06
36.46
28.53
+32 Vote
35.22
42.14
36.42
NaSGEC-Exam
CSED-C
Greedy
35.65
26.46
33.33
+32 Vote
50.54
27.88
43.47
Table 6: Cross-domain evaluation results. Each model is trained on one dataset and evaluated on both test sets.
Figure 4: Three categories of Chinese text error correction, illustrating a progression from surface-level character errors to deep semantic errors. Errors in the source sentences are highlighted in red, and the corresponding edits in the target sentences are shown in green.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Template for Data Processing
Figure 6: Training curves of GLPO during reinforcement learning: (a) reward, (b) policy entropy, and (c) generated sequence length.
Figure 7: Training curves of GRPO during reinforcement learning: (a) reward, (b) policy entropy, and (c) generated sequence length.
Figure 8: Two examples of case study
Ratio
Mode
Decoding
P
R
F0.5
non-CoT
w/o thinking
Greedy
48.03
37.37
45.44
+32 vote
48.07
38.19
45.71
CoT
w/ thinking
Greedy
37.28
35.80
36.97
+32 vote
47.73
41.55
46.35
1:1 mix
w/o thinking
Greedy
44.27
38.94
43.09
+32 vote
45.48
40.66
44.43
Appendix
Table 7: Performance comparison of models trained with different CoT to non-CoT data ratios.
N
P
R
F0.5
Agree.
4
48.89
42.68
47.50
47.6%
6
47.88
40.51
46.20
49.9%
8
49.34
42.15
47.72
47.7 %
Appendix
Table 8: Parameter study: effect of the number of rollouts N on CSED-C. “Agree.” reports the average agreement rate of 32-vote inference using the corresponding trained policy.