Argument Mining (AM) is a critical NLP task that remains significantly under-resourced in Arabic. This paper presents \testttSTAR−Ar, a BERT-BiLSTM-CRF architecture for argument discourse detection and classification, as our system for Daleel 2026, the inaugural Arabic argument mining shared task. The task requires the identification and classification of argumentative discourse units (ADUs) in debate and editorial texts.We jointly model these two objectives as a token-level sequence labeling task using a BERT-BiLSTM-CRF architecture that combines contextual transformer embeddings with structural transition constraints to support accurate span detection. \testttSTAR−Ar achieves an F1-score of 72.69 on validation and 73.7 on test data. Our domain-specific analysis shows that models trained exclusively on editorials underperform those trained on debates, a disparity we primarily attribute to the smaller size of the editorial dataset. The code for \testttSTAR−Ar is available at [\faGithubTTLabatDaleel2026](https://github.com/ENTAILab/daleel2026Arabic−Argumentative−Discourse−Mining)
Figures & tables
Figure 1: Overview of STAR-Ar , a BERT–BiLSTM–CRF sequence tagger for Task 2 (ADU span detection and classification) of Daleel 2026.
Split
Genre
Paragraphs
AS
OT
AN
TE
CO
ST
Train
Overall
612
462 (75.5%)
222 (36.3%)
182 (29.7%)
160 (26.1%)
36 (5.9%)
28 (4.6%)
Debate
357
287 (80.4%)
202 (56.6%)
69 (19.3%)
100 (28.0%)
22 (6.2%)
6 (1.7%)
Editorial
255
175 (68.6%)
20 (7.8%)
113 (44.3%)
60 (23.5%)
14 (5.5%)
22 (8.6%)
Dev
Overall
217
157 (72.4%)
63 (29.0%)
68 (31.3%)
76 (35.0%)
12 (5.5%)
9 (4.1%)
Debate
129
106 (82.2%)
52 (40.3%)
33 (25.6%)
43 (33.3%)
7 (5.4%)
1 (0.8%)
Editorial
88
51 (58.0%)
11 (12.5%)
35 (39.8%)
33 (37.5%)
5 (5.7%)
8 (9.1%)
Table 1: Paragraph-level label coverage in the Daleel 2026 dataset. Each value reports the number and percentage of paragraphs containing at least one span of the corresponding label. Percentages are computed with respect to the number of paragraphs in each row and do not sum to 100% because a paragraph may contain multiple labels.
Split
Evaluation
Editorial
Debate
Both
Dev
Editorial
60.80
51.30
66.52
Debate
58.28
75.29
75.04
Both
60.15
69.25
72.69
Test
Editorial
62.35
54.52
63.56
Debate
60.07
76.17
77.84
Both
61.25
70.68
73.74
Table 2: F1-score (%) for evaluation on the Daleel development and test sets. Columns denote the training domain, while rows denote the evaluation split and domain.
Figure 2: Cross-domain span-level confusion matrix. Rows are gold classes; cells show mean character-wise overlap with predictions. The diagonal represents span-level recall; ’NONE’ captures false negatives.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Encoder Model
Task 2 Overall F1 (%)
AraBERTv02-Twitter-b
63.75±0.43
AraBERTv02-Twitter-l
65.82±2.17
MARBERTv2
67.53±1.53
Appendix
Table 3: Ablation Study: Impact of different pre-trained text encoders on overall performance (models trained on the combined dataset).
Figure 3: Paragraph-level dev-set heatmap (cross-domain trained). Rows list gold labels with paragraph support; the diagonal gives recall. Off-diagonal cells show spurious predictions on missed gold paragraphs.
Large Language Models (LLMs) are increasingly assessed and utilized in the field of Argument Mining (AM), thanks to their strong general reasoning capabilities. However, standard training-free models often miss sophisticated details, specifically in contexts where two parts of the text have to be analyzed together. Furthermore, self-correction mechanisms tend to reinforce initial hallucinations in reasoning. Overcoming these limitations typically requires expensive, domain-specific supervised fine-tuning. Recent work has shown that a multi-agent paradigm can address such weaknesses for the component classification task through dialectical refinement with a Proponent-Opponent-Judge architecture, setting a promising direction for training-free approaches in the field. In this paper, we extend and evaluate this framework on the Argument Relation Identification and Classification (ARIC) task, reformulating it as a debate over component pairs. Besides that, we introduce a confidence gating mechanism that enables debating only on the uncertain cases and accepting the initial prediction when confidence is high. On the UKP Argument Annotated Essays v2 corpus, we demonstrate that the selective debate achieves the highest Macro F1 among all training-free methods, while debate over all samples degrades performance below that of one of the baselines. All generative approaches also outperform fine-tuned RoBERTa models on Macro F1, suggesting that the under-representation of the Attack class was more damaging to supervised fine-tuning than to inference-only models. Additionally, our framework produces human-readable debate transcripts, offering interpretability absent from both single-agent and supervised classifiers.
Jakub Bąba, Jarosław A. Chudziak
Faculty of Electronics and Information Technology, Warsaw University of Technology, Poland
Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems. Our code and the datasets used are openly available in the following GitHub repository.
Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles. While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce CLASP-Ar, which reformulates the task as cloze-style masked language modeling. In this approach, the target, predicted sentiment, and text are combined into a single prompt whose [MASK] prediction is restricted to a verbalizer-constrained label vocabulary.
Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler
Text Technology Lab (TTLab), Goethe University Frankfurt