Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing
Authors: Muhammad Imran, Ana Ezquerro, Carlos Gómez-Rodríguez, Anders Søgaard, David Vilares
Organizations: Universidade da Coruña, CITIC, Departamento de Ciencias de la Computación y Tecnologías de la Información, Campus de Elviña s/n, 15071, A Coruña, Spain · Graz University of Technology, Institute in Machine Learning and Neural Computation, Rechbauerstraße 12, 8010, Graz, Austria · University of Copenhagen, Department of Computer Science, Lyngbyvej 2, DK-2100, Copenhagen, Denmark
This study addresses the problem of structured sentiment analysis, whose goal is to obtain a fine-grained sentiment graph where the nodes represent spans of sentiment holders, targets, and expressions, while the arcs define the relationships among them. Our proposed approach casts the task as dependency graph parsing, but departs from traditional parsing methods by solving it through sequence labeling. To do so, we leverage recent advances in linearized graph encodings that allow each word in the input to be assigned a label, effectively capturing the structure of the dependency graph. We conducted experiments on seven datasets spanning five languages (English, Spanish, Norwegian, Basque, and Catalan), showing performance competitive with leading, more complex single-model approaches.
Figures & tables
Figure 1 : Head-first transformation of a sentiment graph to a dependency graph [ 6 ] . The sentiment graph has two opinion tuples: O1=(∅ , the hotel , in great condition , + ) and O2=( my experience , the service , ruined , - ) .
Figure 2 : Absolute ( A ), relative ( R ) and bracketing ( B ) encodings for an example extracted from [ 7 ] . Note that the <b> token is needed to identify the roots of the graph.
Figure 3 : Graph example with two crossing arcs in the same direction: (0→6) crosses (3→7) . For the bracketing ( B ) encoding, (3→7) is encoded with a different set of brackets ( /* , >* ). The 6k bit encodings require k=2 , so each plane is encoded with an independent sequence of bits.
Figure 4 : 4k ( B4 k ) and 6k -bit ( B6 k ) encoding example from the graph displayed in Figure 3 . Note that the arcs of the original graph need to be distributed in two different subsets that are independently encoded. Artificial arcs (green dashed lines) are generated in the 4k -bit encoding to satisfy the first assumption of the algorithm (each node needs one and only one parent). Those arcs are ignored for B6 1 and B6 2 . Note that in this case the subsets of the 4k and 6k -bit encoding are the same, but this might not happen for other graphs.
Figure 5 : Visualization of the system at training time. Each module is represented with a different color: purple for the sentiment-to-dependency transformation, green for the graph linearization and blue for the neural network. Hypothetical errors are displayed in red for completeness. The input is a sentiment graph ( 1 ) that is converted into a dependency graph ( 2 ) through Barnes et al. [6] . The input sentence is fed to Mθe and Mθℓ to obtain the hidden embeddings ( 3 ) and the predicted labels ( 4 ). The arcs from 2 are used to pair embeddings and feeding them to Mθr ( 5 ). The network is optimized with the cross-entropy loss of the predicted labels and predicted relations.
Figure 6 : Visualization of the system at inference time. Same color notation as in Figure 5 . The input is a sentence ( 1 ). The embeddings and predicted labels are obtained by the trained system, as depicted in Figure 5 . The decoding process is applied to the predicted labels to recover an unlabeled dependency graph. The r component of each predicted arc is obtained by gathering the corresponding embeddings of 2 and feed them to Mθr . The predicted dependency graph ( 5 ) is transformed into a sentiment graph ( 6 ) through the inverse transformation of [ 6 ] .
#sents
#holders
#targets
#expr
n
∣O∣
∣A∣/n
DS en
2253
69
869
1525
19.99
0.36
0.07
232
8
117
186
18.09
0.42
0.09
318
12
155
251
20.40
0.41
0.08
MPQA en
5581
2925
8625
3147
24.37
0.28
0.10
1979
852
2499
1024
23.91
0.27
0.08
2025
894
2696
958
23.85
0.24
0.08
Table 1 : General statistics of the datasets used in our experimental study. Subscripts denote the language denoted by its ISO-639 code. Different splits in each subrow: train, development and test. Columns n , ∣O∣ and ∣A∣/n denote the average sentence length, number of opinion tuples and ratio of number of arcs and nodes, respectively.
Br.
6k -bit
Ω
1
2
3
1
2
3
DS en
99.0
99.8
99.8
92.9
99.8
99.8
99.8
99.0
99.8
99.8
92.5
99.8
99.8
99.8
MPQA en
92.4
95.5
95.9
83.1
94.2
95.7
95.9
90.4
96.5
96.5
82.3
95.0
96.5
96.5
MB ca
99.1
99.8
99.8
86.1
97.9
99.4
99.8
Table 2 : Coverage (SF 1 ) of each encoding in the test split. Ω represents the SF 1 of the original sentiment to dependency transformation. The first and second subrows for each dataset are for the head-first and head-final approach, respectively.
A
R
Br.
6k -bit
#rels
1
2
3
1
2
3
head-first
DS en
117
134
37
45
45
15
30
37
6
MPQA en
528
615
170
380
403
22
96
140
13
MB ca
303
274
96
138
138
19
55
72
4
MB eu
197
193
59
91
91
18
48
61
4
NoReC no
1047
1371
276
722
778
24
147
243
14
Table 3 : Size of the label space ( ∣L∣ ) for each encoding and representation; and number of unique arc labels ( #rels ). Same notation as in Table 4 .
h.first
h.final
A
R
Br.
6k -bit
Biaf.
A
R
Br.
6k -bit
Biaf.
1
2
3
1
2
3
1
2
3
1
2
3
DS en
37.9
19.8
44.0
41.4
41.4
43.8
46.1
43.3
35.0
42.1
18.6
36.5
39.8
40.9
41.2
38.9
41.2
44.3
MPQA en
17.6
18.6
21.3
17.1
28.5
32.6
28.8
32.3
29.8
16.8
13.1
27.2
17.8
28.5
31.5
34.8
31.6
29.8
MB ca
68.3
57.1
61.6
69.2
70.8
73.0
77.0
71.8
61.1
51.1
47.6
56.8
64.6
62.7
62.8
65.0
67.0
69.3
MB eu
71.3
57.1
60.5
66.4
69.5
72.5
68.6
71.5
62.1
61.4
54.3
55.3
64.6
60.6
65.5
63.0
65.2
70.9
Table 4 : SF 1 score in the test set for the absolute ( A ), relative ( R ), bracketing ( Br ) and 6k -bit encodings. Subcolumns in the bracketing and bit-based encodings denote the value of the hyperparameter k . The best result is highlighted in bold. The average is included in the last row ( μ ).
DS en
MPQA en
MB ca
MB eu
NoReC no
OpeNER en
OpeNER es
μ
* [ 22 ]
49.4
44.7
72.8
73.9
52.9
76.0
72.2
63.1
[ 8 ]
48.5
41.6
72.8
73.9
52.4
76.3
74.2
62.8
[ 27 ]
46.3
40.2
70.9
71.5
53.3
75.6
73.2
61.6
[ 4 ]
41.0
37.5
68.1
72.3
50.4
74.7
73.5
59.6
[ 1 ]
41.4
34.9
65.3
68.0
46.2
69.8
69.2
56.4
[ 52 ]
49.0
35.1
68.4
68.6
49.6
67.6
62.3
56.1
Table 5 : SF 1 comparison of different studies and our best-performing encoding ( 6k -bit with k=2 ). Systems proposed in the SemEval-2022 Task 10 [ 7 ] are marked with an asterisk. Best results highlighted in bold.
Figure 7 : F 1 -score ( y -axis) per displacement ( x -axis) in the head-first (left) and head-final (right) representations, alongside the rug plot of the ground truth displacements. The biaffine baseline (DM) and the encodings are denoted with different acronyms: absolute (A), relative (R), bracketing (B), and 6k -bit (B6) encodings, using subscripts for the value of the hyperparameter k , when applicable.
Dataset
Model
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
MB ca
[ 37 ] OT
48.0
72.5
68.9
-
-
-
65.7
63.3
[ 37 ] LE
60.8
70.8
72.5
-
-
-
64.5
62.1
[ 37 ] NC
56.1
69.8
70.5
-
-
-
63.5
61.7
[ 21 ]
36.8
-
75.5
73.4
-
-
66.4
63.2
[ 32 ]
46.2
74.2
71.0
60.9
64.5
-
59.3
Table 6 : Fine-grained SSA results comparison against state of the art systems and Biaffine baseline: Spans (F 1 ) column is the token-level F 1 -score for the holder, target and expression; T.SF 1 is the targeted F 1 score; UF 1 and LF 1 are the unlabeled and labeled F 1 scores for dependency graph parsing; and S F 1 and NSF 1 are the sentiment F 1 score with or without the sentiment polarity. Our best-performing encoding 6k -bit ( B6 3 ) where subscript denotes the value of the hyperparemeter k . Dash (-) for results not reported in the original paper. Best result highlighted in bold.
Figure 8 : Performance (SF 1 , y-axis) vs inference speed (tokens per second, x-axis) on the head-first representation. The Pareto front is displayed with dashed lines. The biaffine baseline (Biaff) and the encodings are denoted with different acronyms: absolute (A), relative (R), bracketing (B), and 6k -bit (B6), using subscripts for the value of the hyperparameter k , where applicable.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
[ 32 ]
65.8
71.0
76.7
59.6
66.1
-
62.7
[ 37 ] NC
58.9
63.5
73.9
-
-
-
59.8
58.6
[ 37 ] LE
57.6
64.9
72.5
-
-
-
60.0
58.8
[ 37 ] OT
64.2
67.4
73.2
-
-
-
62.5
61.3
[ 21 ]
61.4
-
70.0
74.5
-
-
63.8
62.6
Appendix
Table A2 : Fine-grained results on the MultiBooked (Basque) dataset. Same notation as in Table A .
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
[ 32 ]
63.6
55.3
56.1
31.5
40.0
-
32.7
[ 37 ] NC
60.3
51.8
54.2
-
-
-
42.7
39.3
[ 37 ] LE
64.0
52.3
56.1
-
-
-
43.7
40.4
[ 37 ] OT
65.1
58.3
60.7
-
-
-
47.8
41.6
[ 21 ]
62.5
-
61.4
61.3
-
-
46.7
41.9
Appendix
Table A3 : Fine-grained results on the NoReC dataset.
Spans (F 1 )
T.F 1
Dep.
Sent.
Holder
Target
Exp.
UF 1
LF 1
NSF 1
SF 1
[ 32 ]
50.0
44.8
43.7
30.7
-
34.9
-
27.4
[ 37 ] NC
31.4
35.0
35.1
-
-
-
24.8
22.9
[ 37 ] LE
32.5
38.0
36.2
-
-
-
28.8
27.3
[ 37 ] OT
42.2
40.6
39.3
-
-
-
33.2
31.2
[ 21 ]
47.0
-
44.7
45.0
-
-
37.8
34.1
Appendix
Table A4 : Fine-grained results on the Darmstadt Universities dataset. Same notation as in Table A .
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
[ 32 ]
47.9
50.7
47.8
33.7
37.6
-
20.4
[ 37 ] NC
58.4
60.3
55.8
-
-
-
38.7
28.3
[ 37 ] LE
53.6
53.4
53.4
-
-
-
33.8
27.0
[ 37 ] OT
55.7
64.0
53.5
-
-
-
45.1
34.1
[ 21 ]
46.5
-
51.8
47.6
-
-
29.3
24.1
Appendix
Table A5 : Fine-grained results on the MPQA dataset.
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
h.first
A
66.2
77.7
76.2
53.0
50.0
49.0
41.8
41.2
R
70.4
81.9
74.3
69.4
63.4
61.4
53.8
52.5
B 1
73.1
80.9
75.3
73.8
66.8
65.2
51.7
51.3
B 2
68.8
82.3
76.9
75.6
68.6
66.7
56.1
55.0
B 3
65.7
80.8
75.6
75.8
67.4
65.6
54.0
53.0
Appendix
Table A6 : Fine-grained results on the OpeNER (English) dataset.
Spans (F 1 )
T.SF 1
Dep.
Sent.
Holder
Target
Exp
UF 1
LF 1
NSF 1
SF 1
h.first
A
68.4
74.7
75.2
48.9
45.8
44.5
39.0
38.5
R
75.9
79.7
71.3
63.7
58.5
56.4
47.0
45.6
B 1
67.5
78.1
73.1
68.6
61.8
60.4
49.0
47.9
B 2
70.6
79.7
77.0
71.5
64.8
62.9
57.7
56.5
B 3
69.9
79.1
75.2
70.3
64.5
62.9
56.4
55.4
Appendix
Table A7 : Fine-grained results on the OpeNER (Spanish) dataset.
Encoding
Training
Inference
Token/s
Sentences/s
Token/s
Sentences/s
A
3792.25
224.78
4454.28
268.80
R
3818.01
225.91
4491.96
270.72
B 1
3831.82
226.45
4409.86
266.07
B 2
3796.78
224.57
4502.59
271.32
B 3
3828.62
226.31
4486.74
270.80
Appendix
Table A8 : Model efficiency in terms of training and inference throughput (tokens/s, sentences/s)
Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained structural reasoning. We introduce GHI, a Graphormer-over-Conditioned-Hypergraph-Incidence framework that is designed as an incidence-based structural reasoning layer built on a bipartite topology. GHI represents diverse linguistic and semantic evidence as token--hyperedge incidence relations, allowing different structural signals to be incorporated through a unified interface. Extensive experiments on six standard ABSA benchmarks show that GHI outperforms all baselines on the SemEval domains, and multi-seed evaluations show stable improvements over strong DeBERTa. Further experiments show that with only 247M parameters, GHI approaches the performance of 11B Flan-T5 based methods on the ISE benchmark. Moreover, it demonstrates strong robustness on the challenging ARTS datasets, maintaining highly competitive performance where traditional models degrade. These results demonstrate that compact structural reasoning remains a valuable alternative to scale-driven approaches for fine-grained tasks.
Semantic Role Labeling (SRL) provides an explicit representation of predicate-argument structure, capturing linguistically grounded relations such as who did what to whom. While recent NLP progress has been dominated by large language models (LLMs), these systems often rely on implicit semantic representations, often lacking explicit structural constraints and systematic explanatory mechanisms. Traditionally, SRL systems have often relied on AllenNLP; however, the framework entered maintenance mode in December 2022, limiting compatibility with evolving encoder architectures and modern inference requirements. We revisit structured SRL modeling, introducing a modernized encoder-based framework that preserves explicit predicate-argument structure while enabling inference 10 times faster. Using BERT-base, the model attains comparable predictive performance, and RoBERTa and DeBERTa further improve F1 performance within the same framework. We adopt a dependency-informed diagnostic methodology to characterize span-level inconsistencies and conduct a representation-level analysis of LLM behavior under dependency-informed structural signals. Results indicate that dependency cues primarily improve structural stability. Finally, we illustrate how the framework's explicit predicate-argument structure can support multilingual SRL projection as a downstream application.
Augmenting Transformers with linguistic structures effectively enhances the syntactic generalization performance of language models. Previous work in this direction focuses on syntactic tree structures of languages, in particular constituency tree structures. We propose Graph-Infused Layers Transformer Language Model (GiLT) which leverages dependency graphs for augmenting Transformer language models. Unlike most previous work, GiLT does not insert extra structural tokens in language modeling; instead, it injects structural information into language modeling by modulating attention weights in the Transformer with features extracted from the dependency graph that is incrementally constructed along with token prediction. In our experiments, GiLT with semantic dependency graphs achieves better syntactic generalization while maintaining competitive perplexity in comparison with Transformer language model baselines. In addition, GiLT can be finetuned from a pretrained language model to achieve improved downstream task performance. Our code is released at https://github.com/cookie-pie-oops/GiLT-LM.
Tianyu Huang, Yida Zhao, Chuyan Zhou +1
School of Information Science and Technology, ShanghaiTech University · Shanghai Engineering Research Center of Intelligent Vision and Imaging