Organizations: College of Applied Science, Shenzhen University, Shenzhen, China · School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China · Harbin Institute of Technology (Shenzhen), Shenzhen, China · The Chinese University of Hong Kong, Hong Kong, China · Pengcheng Laboratory, Shenzhen, China
Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we revisit synthetic data generation for AM from a new perspective and propose a novel adversarial reinforcement learning framework for data synthesis. The proposed framework jointly optimizes the generator and the discriminator in an adversarial loop, in which the generator produces structured AM instances, and the discriminator provides learning signals by distinguishing real data from synthetic candidates. This enables the generator to progressively improve both the structural accuracy of generated argument data while maintaining diversity through adversarial feedback. Extensive experiments demonstrate that the proposed framework consistently improves AM performance on three benchmark datasets in both full-data and low-resource settings, validating its effectiveness and scalability.
Figures & tables
Figure 1: Overview of the proposed framework. The framework consists of two iterative stages: generator optimization (\small{1}⃝–\small{8}⃝) and discriminator refinement (\small{9}⃝–\small{13}⃝). Through alternating adversarial training, the generator and discriminator progressively improve together.
AAEC
AbstRCT
CDCP
AM Model
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
Origin
84.70
75.88
54.16
71.58
69.01
63.34
38.43
56.93
82.12
68.60
31.07
60.60
EDA
85.20
76.62
54.18
72.00
69.74
63.91
40.07
57.91
82.37
68.64
30.97
60.66
FTGA
85.71
77.00
54.60
72.44
70.47
63.98
40.16
58.20
81.89
68.54
32.95
61.13
JTLS
85.60
76.35
54.56
72.17
70.15
64.32
40.64
58.37
82.22
68.69
30.95
60.62
QOS
86.19
76.57
56.75
73.17
72.53
66.08
39.49
59.37
82.72
68.62
33.75
61.70
Table 1: Main experimental results under full (100%) and low-resource (5%) training data settings. Bold numbers denote the best performance among all methods on each dataset. “ Avg. ” is the arithmetic mean of the three F1 scores. The significance test results indicate that all improvements on Avg. are statistically significant with p < 0.05.
AAEC
AbstRCT
CDCP
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
Ours
88.16
79.63
59.87
75.89
74.03
69.45
41.74
61.74
85.32
70.03
38.16
64.50
w/o Disc. Refine
84.97
75.75
57.11
72.61
72.23
66.52
40.83
59.86
83.86
67.50
34.30
61.89
100%
w/o Sim. Penalty
84.44
78.85
59.32
74.20
73.86
68.93
41.25
61.35
84.84
68.85
37.10
63.60
Ours
74.74
58.19
26.88
53.27
65.53
58.30
29.39
51.07
79.10
55.24
11.93
48.76
w/o Disc. Refine
73.02
56.81
25.54
51.79
63.94
55.94
28.91
49.60
77.28
53.27
8.60
46.38
Table 2: Ablation study. Here, we use the ST model for AM.
Figure 2: Performance comparison of different rounds. Here, we use the ST model for AM.
Figure 3: Performance comparison under different numbers of synthetic samples ( K ) using the ST model for AM.
AAEC
AbstRCT
CDCP
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
10%
Origin
69.94
52.12
17.53
46.53
58.36
51.65
23.78
44.60
77.95
51.04
5.92
44.97
Ours
76.54
59.42
24.73
53.56
67.11
59.30
31.23
52.55
81.00
56.19
13.67
50.29
30%
Origin
74.41
60.17
19.04
51.21
62.31
55.78
27.51
48.53
79.36
55.31
12.10
48.92
Ours
80.31
66.47
25.84
57.54
69.01
62.03
33.76
54.93
82.31
61.26
19.10
54.22
50%
Origin
79.41
67.30
36.48
61.06
65.57
59.61
31.27
52.15
79.76
59.87
15.43
51.69
Table 3: Experimental results under different training data settings using the ST model for AM.
(a) On the AAEC dataset.
Dataset
Method
F1span
F1aci
F1ari
Avg.
AAEC
POPri
72.64
54.91
21.62
49.72
SEAL
72.73
56.24
20.14
49.70
SFTSyn
68.35
48.11
15.55
44.00
Ours
74.74
58.19
26.88
53.27
AbstRCT
POPri
62.69
54.82
22.41
46.64
SEAL
63.84
55.86
20.10
46.60
Table 4: Performance comparison of different training methods using ST as the AM model.
Dataset
Method
F1span
F1aci
F1ari
Avg.
AAEC
Ours
74.74
58.19
26.88
53.27
w/ Structure Reward
75.24
59.93
27.07
54.08
AbstRCT
Ours
65.53
58.30
29.39
51.07
w/ Structure Reward
66.47
58.93
31.40
52.27
CDCP
Ours
79.10
55.24
11.93
48.76
w/ Structure Reward
77.95
54.14
12.64
48.24
Table 5: Effect of the structure-aware reward on downstream AM performance using the ST model for AM.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
5% Train
100% Train
Vali.
Test
AAEC
15
289
32
80
AbstRCT
18
350
50
100
CDCP
27
530
58
150
Appendix
Table 6: Statistics of the datasets. The 5% and 100% settings differ only in the size of the training split, while the validation and test sets remain unchanged.
AAEC
AbstRCT
CDCP
Setting
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
QOS+DOS
72.10
56.02
23.45
50.52
63.04
56.22
28.18
49.15
77.61
54.80
7.87
46.76
Qwen3-1.7B+Qwen3-8B
70.78
53.64
21.05
48.49
59.26
54.53
22.47
45.42
75.16
51.62
4.73
43.84
Qwen3-4B+Qwen3-8B
73.16
56.73
25.37
51.75
63.31
56.80
26.27
48.79
77.79
53.11
7.90
46.27
Qwen3-8B+Qwen3-1.7B
72.54
56.74
24.96
51.41
63.55
56.64
27.11
49.10
78.05
53.86
6.82
46.24
Qwen3-8B+Qwen3-4B
73.89
58.01
26.13
52.68
64.99
57.75
29.45
50.73
78.57
54.21
8.23
47.00
Appendix
Table 7: Robustness analysis under different generator–discriminator model combinations using ST as the AM model. Here, Ma + Mb denotes using model Ma as the generator and model Mb as the discriminator. Longformer refers to Longformer-4096. In the main experiments, both the generator and discriminator are based on Qwen3-8B.
Figure 5: RST tree depth distributions of different argument data under the full-data setting.
Dataset
Stage
GRPO Reward
Reward Std.
Grad. Norm
Entropy
D Acc. (Beg.)
D Acc. (End.)
AbstRCT
SFT
–
–
–
–
25.0%
69.1%
Round 1
1.72±0.057
0.35±0.095
0.08±0.044
0.26±0.008
55.0%
80.0%
Round 2
1.79±0.034
0.25±0.101
0.07±0.007
0.25±0.007
63.0%
86.2%
Round 3
1.82±0.032
0.23±0.076
0.07±0.020
0.24±0.016
67.0%
88.4%
CDCP
SFT
–
–
–
–
25.0%
76.4%
Round 1
1.87±0.026
0.13±0.017
0.10±0.013
0.26±0.008
48.4%
81.2%
Appendix
Table 8: Training dynamics of the generator and discriminator across adversarial evolution stages. GRPO Reward and Reward Std. respectively denote the reward obtained during generator optimization and the standard deviation of rewards. Grad. Norm and Entropy denote the gradient norm and policy entropy of the generator, respectively. D Acc. (Beg.) and D Acc. (End.) denote the discriminator accuracy at the beginning and end of each refinement stage. SFT denotes the discriminator initialization stage before adversarial evolution.
Dataset
Methods
Time
Memory
Avg.
AbstRCT
Ours
9.3h
60.3GB
51.07
w/ Longformer
8.6h
51.1GB
48.17
SFTSyn
0.9h
43.4GB
42.26
CDCP
Ours
8.1h
58.2GB
48.76
w/ Longformer
7.9h
47.7GB
45.45
SFTSyn
1.1h
49.8GB
42.72
Appendix
Table 9: Computational cost and downstream AM performance of different synthesis methods. Time denotes the total training time in hours, Memory denotes the peak GPU memory usage in GB, and Avg. denotes the average of the three downstream AM F1 scores. We use ST as the AM model.
Figure 6: Examples of original data and synthetic data.