cs.CLSep 24, 2026
SaveTTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)
Organizations: Text Technology Lab (TTLab), Goethe University Frankfurt
Abstract
Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles. While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce , which reformulates the task as cloze-style masked language modeling. In this approach, the target, predicted sentiment, and text are combined into a single prompt whose prediction is restricted to a verbalizer-constrained label vocabulary.
Figures & tables
| Dataset | Property | Value |
| ArabicStance-X | Usage | Pre-training |
| Samples | 11,500 | |
| Targets | 17 | |
| Favor | 4,980 | |
| Against/None | 4,452 / 2,068 | |
| Mawqif-XT | Usage | Main task |
Table 1: Statistics of the datasets used in this work. ArabicStance-X Alkhathlan et al. (2025) is used only for intermediate stance pre-training, while Mawqif-XT Albalawi et al. (2026b) is used for model development and evaluation.
| Track | Favg2 | Favg3 | Overall Accuracy |
|---|---|---|---|
| Track 1 | 0.7136 | 0.5313 | 0.6761 |
| Track 2 | 0.7414 | 0.6564 | 0.7096 |
Table 2: Performance of CLASP-Ar on the final evaluation data of Track 1 (seen targets) and Track 2 (unseen targets).
| Method | Accuracy | ||
|---|---|---|---|
| Baseline | 85.75 | 71.91 | 83.20 |
| Without Pre-training | |||
| Base Setup | 85.95 | 73.62 | 83.94 |
| + Extended None | 86.33 | 70.40 | 83.94 |
| + Class Weights | 85.40 | 74.23 | 83.36 |
| With Pre-training | |||
Table 3: Impact of adding None instances and class weights to CLASP-Ar . Full experiment results with multiple seeds can be found in Appendix B
| Ablation Setting | Accuracy | ||
|---|---|---|---|
| Baseline (None) | 84.90 | 70.83 | 82.61 |
| + Sentiment | 86.28 | 72.36 | 84.38 |
| + Sarcasm | 84.91 | 71.44 | 82.98 |
| + Both | 85.85 | 72.50 | 83.74 |
Table 4: Ablation analysis on the development set reporting , , and Overall Accuracy across 4 settings.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Encoder | Plain | Pretrain | Joint |
|---|---|---|---|
| AraBERT large-twitter † | 84.13 | 85.24 | 84.42 |
| MARBERTv2 | 82.65 | 82.35 | 82.49 |
| CAMeLBERT-mix | 79.50 | 80.18 | 81.15 |
Table 5: Impact of transfer learning strategies on development . Pretrain denotes intermediate fine-tuning on ArabicStance-X followed by training on Mawqif-XT . Joint denotes training on the union of Mawqif-XT and ArabicStance-X . † Original shared-task baseline using the large model.
| Method | Accuracy | |||||
|---|---|---|---|---|---|---|
| Baseline | 89.77 | 81.72 | 44.25 | 85.75 | 71.91 | 83.20 |
| CLASP-Ar | ||||||
| Without Pre-training | ||||||
| Base Setup | 89.99 ± 0.84 | 81.91 ± 1.03 | 48.96 ± 4.60 | 85.95 ± 0.82 | 73.62 ± 1.78 | 83.94 ± 0.87 |
| + Extended None | 89.94 ± 0.41 | 82.73 ± 0.22 | 38.55 ± 10.81 | 86.33 ± 0.30 | 70.40 ± 3.67 | 83.94 ± 0.42 |
| + Class Weights | 89.66 ± 1.29 | 81.14 ± 0.67 | 51.88 ± 2.13 | 85.40 ± 0.86 | 74.23 ± 0.85 | 83.36 ± 1.33 |
Table 6: Overall stance detection performance comparison. Our implementation is presented as mean and standard deviation over 5 runs using different seeds. All values are reported as percentages (%). Baseline uses same encoder as for the CLASP-Ar which is AraBERT-large-twitter
| Method | Covid Vaccine | Digital Trans. | Women Emp. | |||
|---|---|---|---|---|---|---|
| CLASP-Ar | ||||||
| Without Pre-training | ||||||
| Base Setup | 80.78 ± 1.03 | 64.52 ± 2.83 | 86.30 ± 2.89 | 80.99 ± 2.42 | 88.87 ± 1.19 | 70.48 ± 4.44 |
| + Extended None | 82.52 ± 1.00 | 64.63 ± 2.31 | 84.82 ± 2.10 | 74.17 ± 7.77 | 88.15 ± 0.55 | 67.50 ± 4.67 |
| + Class Weights | 78.86 ± 1.71 | 64.38 ± 1.65 | 87.08 ± 0.92 | 81.89 ± 1.53 | 88.82 ± 0.99 | 74.89 ± 1.47 |
Table 7: Target-specific stance detection performance comparison. Results are shown as mean and standard deviation over 5 runs. All values are reported as percentages (%).
| Ablation | Accuracy | |||||
|---|---|---|---|---|---|---|
| None | 89.41 0.30 | 80.39 2.47 | 42.67 1.21 | 84.90 1.38 | 70.83 0.61 | 82.61 1.04 |
| Sentiment | 90.91 0.29 | 81.66 0.69 | 44.51 4.77 | 86.28 0.20 | 72.36 1.73 | 84.38 0.49 |
| Sarcasm | 89.87 0.78 | 79.95 1.48 | 44.50 4.23 | 84.91 0.94 | 71.44 1.57 | 82.98 0.92 |
| Both | 90.28 1.20 | 81.42 1.39 | 45.80 1.03 | 85.85 1.28 | 72.50 1.10 | 83.74 1.35 |
Table 8: Comprehensive evaluation showing label-specific scores ( , , ), macro-averaged scores ( , ), and overall Accuracy across multi-seed experiments on the development set.