Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.
Figures & tables
System
Task adaptation
Tunes backbone
Generates text
T1 RMSE ↓
T2 cF1 ↑
T3 cF1 ↑
SemEval-2026 participants
PAI ( Ruan et al., 2026 )
LoRA + VA alignment
∙
∙
1.0663
57.73
–
TeleAI ( Zhou et al., 2026 )
LoRA + regression head
∙
∙ T2/3
1.0737
55.66
31.26
PALI ( Chen, 2026 )
LoRA adapters
∙
∙
1.1340
57.50
49.20
Takoyaki ( Yamada et al., 2026 )
Retrieval + rules
∘
∙
–
56.20
48.03
nchellwig ( Hellwig et al., 2026 )
LoRA
∙
∙
–
56.55
47.19
Table 1: Test results and task adaptation. T1: micro RMSE over ten corpora; T2/T3: macro cF1 over eight corpora. ∙ yes, ∘ no. Participant aggregates are computed from the per-corpus scores of Yu et al. (2026, Tables 6–8) ; –: not every corpus reported. Best score per column in bold.
Task 1: RMSE ↓
Task 2: cF1 ↑
Task 3: cF1 ↑
Corpus
Ours
Best
Ours
Best
Exact VA
Ours
Best
Cat. acc.
English restaurant
1.2163
1.1035 a
68.21
70.21 f
+5.31
63.48
65.14 f
93.0
English laptop
1.2086
1.2408 a
62.25
63.66 f
+5.54
37.38
42.27 f
60.2
Japanese hotel
0.6454
0.5561 b
50.03
58.37 b
+2.60
37.59
42.52 g
75.3
Japanese finance
0.7296
0.6581 b
–
–
–
–
–
–
Russian restaurant
1.3290
1.2190 c
51.26
57.93 c
+5.64
46.80
55.99 c
91.3
Table 2: Per-corpus test results. Best: best published score for the corpus, from Yu et al. (2026) ; for the aggregate, the best complete-coverage aggregate of Table 1 . Superscripts: (a) LogSigma ( Hikal et al., 2026 ) , (b) TeleAI ( Zhou et al., 2026 ) , (c) PAI ( Ruan et al., 2026 ) , (d) ICT-NLP ( Huang et al., 2026 ) , (e) HUS@NLP-VNU ( Cao et al., 2026 ) , (f) Takoyaki ( Yamada et al., 2026 ) , (g) PALI ( Chen, 2026 ) , (h) nchellwig ( Hellwig et al., 2026 ) , (i) NYCU Speech Lab. Bold: better than Best. Exact VA: gain if our extracted pairs had gold VA. Cat. acc.: category accuracy (%) on predicted pairs that match a gold pair, highlighted below 80. The finance corpora have Task 1 data only.
Variant (test)
Score
Δ
Task 1, micro RMSE ↓
Full system
1.0645
− joint terms (shrinkage)
1.1203
+0.0558
− demonstrations (zero-shot)
1.1015
+0.0370
− calibration (raw scores)
2.0731
+1.0086
Task 2, exact-match pair F1 ↑
Table 3: Component ablations on test (T1 micro over ten corpora, T2/T3 macro over eight). Each variant removes one component and refits the learned postprocessor on the data the final system uses. Removing the T2 checks also removes the extensions they admit. Losses of at least 0.01 RMSE or one F1 point are highlighted.
Corpus
Not proposed
Not selected
Found
Near-miss FP
Eng. rest.
15.5
14.4
70.0
46.6
Eng. laptop
17.1
20.7
62.2
46.7
Jpn. hotel
17.2
34.5
48.3
36.3
Rus. rest.
26.0
21.3
52.7
32.5
Tat. rest.
27.2
28.5
44.4
30.4
Ukr. rest.
27.9
20.3
51.8
34.0
Table 4: Where Task 2 loses gold pairs on test (%): never proposed as a candidate, proposed but not selected, or found; the larger loss per corpus is in bold. Near-miss FP: share of wrongly selected pairs whose aspect and opinion both overlap one gold pair.
System
T2
T3
Loss
PALI ( Chen, 2026 )
57.50
49.20
8.30
Takoyaki ( Yamada et al., 2026 )
56.20
48.03
8.17
nchellwig ( Hellwig et al., 2026 )
56.55
47.19
9.36
TeamLasse ( Strothe, 2026 )
53.43
44.33
9.10
Ours
52.09
44.06
8.03
Table 5: Macro cF1 on Tasks 2 and 3, and the loss when a category is added to the extracted pairs, for the systems that report every corpus of both tasks and score at least as high as ours on both. Smallest loss in bold.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Train rev.
T1 dev rev.
T1 test rev.
T1 test ann.
T2 dev rev.
T2 test rev.
T2 test ann.
English restaurant
2284
200
1000
1504
200
1000
2129
English laptop
4076
200
1000
1421
200
1000
1974
Japanese hotel
1600
200
800
1092
200
800
1443
Japanese finance
1024
200
800
1302
–
–
–
Russian restaurant
1240
56
1072
1637
48
630
1310
Tatar restaurant
1240
56
1072
1637
48
630
1310
Appendix
Table 6: Dataset statistics (rev.: reviews; ann.: annotations). Tasks 2 and 3 share reviews; Task 3 has one more English laptop test annotation.
Year
D
G
Other
n
2008
0
0
1
1
2010
0
0
1
1
2013
0
0
2
2
2014
1
0
3
4
2015
2
1
1
4
2016
6
1
0
7
Appendix
Table 7: Annual paper counts behind Figure 1 . G: the proposed method uses a text-generative model; D: it does not; Other: latent methods or unresolved model use. Years without included papers are omitted.
Level
Valence criterion
Arousal criterion
1
Strongly negative: a severe fault, harsh or contemptuous complaint
Very calm, low energy: the aspect arouses no feeling at all; the writer is indifferent
2
Clearly negative: the aspect is described as bad or disappointing
Calm: the aspect is regarded without emotional charge
3
Moderately negative: real criticism, but not emphatic
Somewhat calm: only the faintest feeling about the aspect
4
Mildly negative: a small complaint or a slight reservation
Mildly calm: a low-energy, subdued feeling
5
Neutral or mixed: no clear polarity, or praise and criticism cancel out
Moderate: an ordinary, middle-of-the-road level of feeling
6
Mildly positive: a small or lukewarm compliment
Moderately activated: the feeling runs a little above ordinary
Appendix
Table 8: Verbatim ordered criteria for the shared Score questions. Level numbers show the benchmark scale; the API criterion indices are 0–8.
Attribute
Description used in category criteria
GENERAL
the entity as a whole, an overall opinion without a more specific attribute
PRICE / PRICES
price, cost or value for money
QUALITY
how well made it is: build quality, reliability, durability, defects; for food and drinks, taste and freshness; for service, how good it is
OPERATION_PERFORMANCE
how well it works in use: speed, power, performance, battery life, responsiveness
USABILITY
ease of use, how easy it is to learn or operate
DESIGN_FEATURES
looks, size, layout, materials, and the features or specifications it has
Appendix
Table 9: Verbatim attribute descriptions in the Task 3 category template. Only categories observed in the eligible training pool are offered as options.
Aspect-Based Sentiment Analysis (ABSA) encompasses seven distinct subtasks, each focusing on different extracted elements. Despite the proven success of generative models in unified aspect sentiment analysis, existing approaches often rely on auto-regressive token-by-token generation without grasping the whole information of the aspect and opinion terms, resulting in boundary insensitivity, particularly in context of multi-word aspect and opinion terms. To address these issues, we present DiffuSent, a non-auto-regressive diffusion framework that systematically formulates all ABSA subtasks as boundary denoising diffusion processes, progressively refining boundaries over noisy states. Furthermore, we introduce a contrastive denoising training strategy which effectively address duplicate predictions with subtle variations introduced by diffusion process. Extensive experiments across 28 settings (7 subtasks x 4 datasets) demonstrate that DiffuSent achieves delivers consistent improvements over the strongest generative and span-based systems. DiffuSent exhibits notable gains on multi-word triplets, achieving an average improvement of +2.48 F1, and maintains robust extraction accuracy in sentences containing multiple sentiment triplets. Moreover, the non-auto-regressive decoding enables substantial efficiency benefits, reaching up to 181 times faster inference than auto-regressive generative baselines
Shu Long, Yanglei Gan, Xuchuan Zhou
University of Electronic Science and Technology of China · Southwest Petroleum University · Southwest Minzu University
This paper presents an approach to the SemEval-2026 Task 3: Dimensional Aspect-Based Sentiment Analysis. We investigate methods for moving beyond traditional categorical sentiment (e.g., positive or negative) to predict fine-grained, real-valued scores for sentiment "valence" (positivity) and "arousal" (intensity). We participate in two subtasks: predicting these scores for given aspects (Subtask 1) and extracting full sets of sentiment details, including aspects, categories, and opinions alongside their scores (Subtask 3). Our approach for the regression task involves a weighted ensemble of transformer-based encoder models. For the Russian language, we further enhance the input by using a large language model (LLM) to generate synthetic sentiment descriptions. For the extraction task, we fine-tune a decoder LLM to perform structured prediction, allowing the system to identify sentiment elements and estimate their numerical scores simultaneously.
Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in the text. Implicit identification typically relies on an auxiliary knowledge source (e.g., a knowledge graph (KG)) linking opinion cues to aspect categories, but for a lower-resource language the practitioner faces a design choice: reuse a mature English KG through multilingual embeddings, or build a smaller native Arabic KG. This paper reports a controlled comparison of the two strategies within a single hybrid pipeline, evaluated on three Arabic benchmarks (M-ABSA, SemEval-2016 Arabic, and HAAD). We further compare two adaptation strategies for the generative extractor that feeds the KG -- zero-shot prompting versus task-specific fine-tuning of an 8B-parameter large language model (LLM). The native Arabic KG (Strategy 2) outperforms the cross-lingual English KG (Strategy 1) by +0.199 micro-F1 on M-ABSA and +0.251 on SemEval-2016, gaining on both precision and recall. Task-specific fine-tuning raises explicit-extraction micro-F1 from <= 0.13 (zero-shot) to 0.66-0.76 on M-ABSA and SemEval-2016 (0.45 on the smaller HAAD), confirming that task adaptation, rather than model scale, is decisive in a morphologically rich language.
Lujain A. Alawwad
Department of Computing and Informatics Saudi Electronic University Riyadh, Saudi Arabia