Organizations: College of Applied Science, Shenzhen University, Shenzhen, China · School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China · Harbin Institute of Technology (Shenzhen), Shenzhen, China · The Chinese University of Hong Kong, Hong Kong, China · Pengcheng Laboratory, Shenzhen, China
Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we revisit synthetic data generation for AM from a new perspective and propose a novel adversarial reinforcement learning framework for data synthesis. The proposed framework jointly optimizes the generator and the discriminator in an adversarial loop, in which the generator produces structured AM instances, and the discriminator provides learning signals by distinguishing real data from synthetic candidates. This enables the generator to progressively improve both the structural accuracy of generated argument data while maintaining diversity through adversarial feedback. Extensive experiments demonstrate that the proposed framework consistently improves AM performance on three benchmark datasets in both full-data and low-resource settings, validating its effectiveness and scalability.
Figures & tables
Figure 1: Overview of the proposed framework. The framework consists of two iterative stages: generator optimization (\small{1}⃝–\small{8}⃝) and discriminator refinement (\small{9}⃝–\small{13}⃝). Through alternating adversarial training, the generator and discriminator progressively improve together.
AAEC
AbstRCT
CDCP
AM Model
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
Origin
84.70
75.88
54.16
71.58
69.01
63.34
38.43
56.93
82.12
68.60
31.07
60.60
EDA
85.20
76.62
54.18
72.00
69.74
63.91
40.07
57.91
82.37
68.64
30.97
60.66
FTGA
85.71
77.00
54.60
72.44
70.47
63.98
40.16
58.20
81.89
68.54
32.95
61.13
JTLS
85.60
76.35
54.56
72.17
70.15
64.32
40.64
58.37
82.22
68.69
30.95
60.62
QOS
86.19
76.57
56.75
73.17
72.53
66.08
39.49
59.37
82.72
68.62
33.75
61.70
Table 1: Main experimental results under full (100%) and low-resource (5%) training data settings. Bold numbers denote the best performance among all methods on each dataset. “ Avg. ” is the arithmetic mean of the three F1 scores. The significance test results indicate that all improvements on Avg. are statistically significant with p < 0.05.
AAEC
AbstRCT
CDCP
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
Ours
88.16
79.63
59.87
75.89
74.03
69.45
41.74
61.74
85.32
70.03
38.16
64.50
w/o Disc. Refine
84.97
75.75
57.11
72.61
72.23
66.52
40.83
59.86
83.86
67.50
34.30
61.89
100%
w/o Sim. Penalty
84.44
78.85
59.32
74.20
73.86
68.93
41.25
61.35
84.84
68.85
37.10
63.60
Ours
74.74
58.19
26.88
53.27
65.53
58.30
29.39
51.07
79.10
55.24
11.93
48.76
w/o Disc. Refine
73.02
56.81
25.54
51.79
63.94
55.94
28.91
49.60
77.28
53.27
8.60
46.38
Table 2: Ablation study. Here, we use the ST model for AM.
Figure 2: Performance comparison of different rounds. Here, we use the ST model for AM.
Figure 3: Performance comparison under different numbers of synthetic samples ( K ) using the ST model for AM.
AAEC
AbstRCT
CDCP
Setting
Method
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
10%
Origin
69.94
52.12
17.53
46.53
58.36
51.65
23.78
44.60
77.95
51.04
5.92
44.97
Ours
76.54
59.42
24.73
53.56
67.11
59.30
31.23
52.55
81.00
56.19
13.67
50.29
30%
Origin
74.41
60.17
19.04
51.21
62.31
55.78
27.51
48.53
79.36
55.31
12.10
48.92
Ours
80.31
66.47
25.84
57.54
69.01
62.03
33.76
54.93
82.31
61.26
19.10
54.22
50%
Origin
79.41
67.30
36.48
61.06
65.57
59.61
31.27
52.15
79.76
59.87
15.43
51.69
Table 3: Experimental results under different training data settings using the ST model for AM.
(a) On the AAEC dataset.
Dataset
Method
F1span
F1aci
F1ari
Avg.
AAEC
POPri
72.64
54.91
21.62
49.72
SEAL
72.73
56.24
20.14
49.70
SFTSyn
68.35
48.11
15.55
44.00
Ours
74.74
58.19
26.88
53.27
AbstRCT
POPri
62.69
54.82
22.41
46.64
SEAL
63.84
55.86
20.10
46.60
Table 4: Performance comparison of different training methods using ST as the AM model.
Dataset
Method
F1span
F1aci
F1ari
Avg.
AAEC
Ours
74.74
58.19
26.88
53.27
w/ Structure Reward
75.24
59.93
27.07
54.08
AbstRCT
Ours
65.53
58.30
29.39
51.07
w/ Structure Reward
66.47
58.93
31.40
52.27
CDCP
Ours
79.10
55.24
11.93
48.76
w/ Structure Reward
77.95
54.14
12.64
48.24
Table 5: Effect of the structure-aware reward on downstream AM performance using the ST model for AM.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
5% Train
100% Train
Vali.
Test
AAEC
15
289
32
80
AbstRCT
18
350
50
100
CDCP
27
530
58
150
Appendix
Table 6: Statistics of the datasets. The 5% and 100% settings differ only in the size of the training split, while the validation and test sets remain unchanged.
AAEC
AbstRCT
CDCP
Setting
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
F1span
F1aci
F1ari
Avg.
QOS+DOS
72.10
56.02
23.45
50.52
63.04
56.22
28.18
49.15
77.61
54.80
7.87
46.76
Qwen3-1.7B+Qwen3-8B
70.78
53.64
21.05
48.49
59.26
54.53
22.47
45.42
75.16
51.62
4.73
43.84
Qwen3-4B+Qwen3-8B
73.16
56.73
25.37
51.75
63.31
56.80
26.27
48.79
77.79
53.11
7.90
46.27
Qwen3-8B+Qwen3-1.7B
72.54
56.74
24.96
51.41
63.55
56.64
27.11
49.10
78.05
53.86
6.82
46.24
Qwen3-8B+Qwen3-4B
73.89
58.01
26.13
52.68
64.99
57.75
29.45
50.73
78.57
54.21
8.23
47.00
Appendix
Table 7: Robustness analysis under different generator–discriminator model combinations using ST as the AM model. Here, Ma + Mb denotes using model Ma as the generator and model Mb as the discriminator. Longformer refers to Longformer-4096. In the main experiments, both the generator and discriminator are based on Qwen3-8B.
Figure 5: RST tree depth distributions of different argument data under the full-data setting.
Dataset
Stage
GRPO Reward
Reward Std.
Grad. Norm
Entropy
D Acc. (Beg.)
D Acc. (End.)
AbstRCT
SFT
–
–
–
–
25.0%
69.1%
Round 1
1.72±0.057
0.35±0.095
0.08±0.044
0.26±0.008
55.0%
80.0%
Round 2
1.79±0.034
0.25±0.101
0.07±0.007
0.25±0.007
63.0%
86.2%
Round 3
1.82±0.032
0.23±0.076
0.07±0.020
0.24±0.016
67.0%
88.4%
CDCP
SFT
–
–
–
–
25.0%
76.4%
Round 1
1.87±0.026
0.13±0.017
0.10±0.013
0.26±0.008
48.4%
81.2%
Appendix
Table 8: Training dynamics of the generator and discriminator across adversarial evolution stages. GRPO Reward and Reward Std. respectively denote the reward obtained during generator optimization and the standard deviation of rewards. Grad. Norm and Entropy denote the gradient norm and policy entropy of the generator, respectively. D Acc. (Beg.) and D Acc. (End.) denote the discriminator accuracy at the beginning and end of each refinement stage. SFT denotes the discriminator initialization stage before adversarial evolution.
Dataset
Methods
Time
Memory
Avg.
AbstRCT
Ours
9.3h
60.3GB
51.07
w/ Longformer
8.6h
51.1GB
48.17
SFTSyn
0.9h
43.4GB
42.26
CDCP
Ours
8.1h
58.2GB
48.76
w/ Longformer
7.9h
47.7GB
45.45
SFTSyn
1.1h
49.8GB
42.72
Appendix
Table 9: Computational cost and downstream AM performance of different synthesis methods. Time denotes the total training time in hours, Memory denotes the peak GPU memory usage in GB, and Avg. denotes the average of the three downstream AM F1 scores. We use ST as the AM model.
Figure 6: Examples of original data and synthetic data.
Arguments are a fundamental aspect of human reasoning, in which claims are supported, challenged, and weighed against one another. We present an end-to-end large language model (LLM)-based system for reconstructing arguments from natural language text into abstract argument graphs. The system follows a multi-stage pipeline that progressively identifies argumentative components, selects relevant elements, and uncovers their logical relations. These elements are represented as directed acyclic graphs consisting of two component types (premises and conclusions) and three relation types (support, attack, and undercut). We conduct two complementary experiments to evaluate the system. First, we perform a manual evaluation on arguments drawn from an argumentation theory textbook to assess the system's ability to recover argumentative structure. Second, we conduct a quantitative evaluation on benchmark datasets, allowing comparison with prior work by mapping our outputs to established annotation schemes. Results show that the system can adequately recover argumentative structures and, when adapted to different annotation schemes, achieve reasonable performance across benchmark datasets. These findings highlight the potential of LLM-based pipelines for scalable argument mining.
Paulo Pirozelli, Victor Hugo Nascimento Rocha, Fabio G. Cozman +1
Universidade de S˜ao Paulo · Center for Artificial Intelligence (C4AI) · Instituto Mau´a de Tecnologia +1
Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics. It requires systems to understand not just keywords but also the context and complex reasoning embedded in a text. Current Argument Mining (AM) techniques identify basic claims and premises, yet they often struggle to capture the richer structural information required by advanced schemas such as the Carneades Argumentation Framework (CAF), which incorporates features such as premise types, proof standards, and argument schemes. We address this limitation by introducing CAF-Gen, an automated multi-agent framework designed to enrich shallow argument structures into CAF-compliant argument models. By employing an iterative Creator-Reviewer pipeline, a creator agent's output is validated by a critical agent to ensure structural integrity. This multi-agent collaboration is crucial for mitigating the structural instability typical of single-pass generative models. Our experiments demonstrate that the iterative feedback loop improves the quality of the resulting data and achieves strong alignment with the original annotations, while producing structurally richer models. Our findings show that the multi-agent system can overcome the limitations of single-pass generation, providing a robust methodology for the automated modeling of formal argumentation.
Jakub Bąba, Jarosław A. Chudziak
Faculty of Electronics and Information Technology, Warsaw University of Technology, Poland
Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.
Prasoon Goyal, Sattvik Sahai, Michael Johnston +14