Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.
Figures & tables
System
Task adaptation
Tunes backbone
Generates text
T1 RMSE ↓
T2 cF1 ↑
T3 cF1 ↑
SemEval-2026 participants
PAI ( Ruan et al., 2026 )
LoRA + VA alignment
∙
∙
1.0663
57.73
–
TeleAI ( Zhou et al., 2026 )
LoRA + regression head
∙
∙ T2/3
1.0737
55.66
31.26
PALI ( Chen, 2026 )
LoRA adapters
∙
∙
1.1340
57.50
49.20
Takoyaki ( Yamada et al., 2026 )
Retrieval + rules
∘
∙
–
56.20
48.03
nchellwig ( Hellwig et al., 2026 )
LoRA
∙
∙
–
56.55
47.19
Table 1: Test results and task adaptation. T1: micro RMSE over ten corpora; T2/T3: macro cF1 over eight corpora. ∙ yes, ∘ no. Participant aggregates are computed from the per-corpus scores of Yu et al. (2026, Tables 6–8) ; –: not every corpus reported. Best score per column in bold.
Task 1: RMSE ↓
Task 2: cF1 ↑
Task 3: cF1 ↑
Corpus
Ours
Best
Ours
Best
Exact VA
Ours
Best
Cat. acc.
English restaurant
1.2163
1.1035 a
68.21
70.21 f
+5.31
63.48
65.14 f
93.0
English laptop
1.2086
1.2408 a
62.25
63.66 f
+5.54
37.38
42.27 f
60.2
Japanese hotel
0.6454
0.5561 b
50.03
58.37 b
+2.60
37.59
42.52 g
75.3
Japanese finance
0.7296
0.6581 b
–
–
–
–
–
–
Russian restaurant
1.3290
1.2190 c
51.26
57.93 c
+5.64
46.80
55.99 c
91.3
Table 2: Per-corpus test results. Best: best published score for the corpus, from Yu et al. (2026) ; for the aggregate, the best complete-coverage aggregate of Table 1 . Superscripts: (a) LogSigma ( Hikal et al., 2026 ) , (b) TeleAI ( Zhou et al., 2026 ) , (c) PAI ( Ruan et al., 2026 ) , (d) ICT-NLP ( Huang et al., 2026 ) , (e) HUS@NLP-VNU ( Cao et al., 2026 ) , (f) Takoyaki ( Yamada et al., 2026 ) , (g) PALI ( Chen, 2026 ) , (h) nchellwig ( Hellwig et al., 2026 ) , (i) NYCU Speech Lab. Bold: better than Best. Exact VA: gain if our extracted pairs had gold VA. Cat. acc.: category accuracy (%) on predicted pairs that match a gold pair, highlighted below 80. The finance corpora have Task 1 data only.
Variant (test)
Score
Δ
Task 1, micro RMSE ↓
Full system
1.0645
− joint terms (shrinkage)
1.1203
+0.0558
− demonstrations (zero-shot)
1.1015
+0.0370
− calibration (raw scores)
2.0731
+1.0086
Task 2, exact-match pair F1 ↑
Table 3: Component ablations on test (T1 micro over ten corpora, T2/T3 macro over eight). Each variant removes one component and refits the learned postprocessor on the data the final system uses. Removing the T2 checks also removes the extensions they admit. Losses of at least 0.01 RMSE or one F1 point are highlighted.
Corpus
Not proposed
Not selected
Found
Near-miss FP
Eng. rest.
15.5
14.4
70.0
46.6
Eng. laptop
17.1
20.7
62.2
46.7
Jpn. hotel
17.2
34.5
48.3
36.3
Rus. rest.
26.0
21.3
52.7
32.5
Tat. rest.
27.2
28.5
44.4
30.4
Ukr. rest.
27.9
20.3
51.8
34.0
Table 4: Where Task 2 loses gold pairs on test (%): never proposed as a candidate, proposed but not selected, or found; the larger loss per corpus is in bold. Near-miss FP: share of wrongly selected pairs whose aspect and opinion both overlap one gold pair.
System
T2
T3
Loss
PALI ( Chen, 2026 )
57.50
49.20
8.30
Takoyaki ( Yamada et al., 2026 )
56.20
48.03
8.17
nchellwig ( Hellwig et al., 2026 )
56.55
47.19
9.36
TeamLasse ( Strothe, 2026 )
53.43
44.33
9.10
Ours
52.09
44.06
8.03
Table 5: Macro cF1 on Tasks 2 and 3, and the loss when a category is added to the extracted pairs, for the systems that report every corpus of both tasks and score at least as high as ours on both. Smallest loss in bold.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Train rev.
T1 dev rev.
T1 test rev.
T1 test ann.
T2 dev rev.
T2 test rev.
T2 test ann.
English restaurant
2284
200
1000
1504
200
1000
2129
English laptop
4076
200
1000
1421
200
1000
1974
Japanese hotel
1600
200
800
1092
200
800
1443
Japanese finance
1024
200
800
1302
–
–
–
Russian restaurant
1240
56
1072
1637
48
630
1310
Tatar restaurant
1240
56
1072
1637
48
630
1310
Appendix
Table 6: Dataset statistics (rev.: reviews; ann.: annotations). Tasks 2 and 3 share reviews; Task 3 has one more English laptop test annotation.
Year
D
G
Other
n
2008
0
0
1
1
2010
0
0
1
1
2013
0
0
2
2
2014
1
0
3
4
2015
2
1
1
4
2016
6
1
0
7
Appendix
Table 7: Annual paper counts behind Figure 1 . G: the proposed method uses a text-generative model; D: it does not; Other: latent methods or unresolved model use. Years without included papers are omitted.
Level
Valence criterion
Arousal criterion
1
Strongly negative: a severe fault, harsh or contemptuous complaint
Very calm, low energy: the aspect arouses no feeling at all; the writer is indifferent
2
Clearly negative: the aspect is described as bad or disappointing
Calm: the aspect is regarded without emotional charge
3
Moderately negative: real criticism, but not emphatic
Somewhat calm: only the faintest feeling about the aspect
4
Mildly negative: a small complaint or a slight reservation
Mildly calm: a low-energy, subdued feeling
5
Neutral or mixed: no clear polarity, or praise and criticism cancel out
Moderate: an ordinary, middle-of-the-road level of feeling
6
Mildly positive: a small or lukewarm compliment
Moderately activated: the feeling runs a little above ordinary
Appendix
Table 8: Verbatim ordered criteria for the shared Score questions. Level numbers show the benchmark scale; the API criterion indices are 0–8.
Attribute
Description used in category criteria
GENERAL
the entity as a whole, an overall opinion without a more specific attribute
PRICE / PRICES
price, cost or value for money
QUALITY
how well made it is: build quality, reliability, durability, defects; for food and drinks, taste and freshness; for service, how good it is
OPERATION_PERFORMANCE
how well it works in use: speed, power, performance, battery life, responsiveness
USABILITY
ease of use, how easy it is to learn or operate
DESIGN_FEATURES
looks, size, layout, materials, and the features or specifications it has
Appendix
Table 9: Verbatim attribute descriptions in the Task 3 category template. Only categories observed in the eligible training pool are offered as options.