End-to-end speech parsing, a task recently proposed, consists in predicting both the transcription and the syntactic tree for a spoken utterance. Existing architectures for speech parsing often utilise intermediate neural networks. In this work, we examine the effectiveness of intermediate neural networks (NN) for parsing, and, specifically, what role do they play. We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing. We demonstrate that intermediate NN units help reduce the representational gap when the pre-trained encoder is frozen. We do a comprehensive evaluation of speech parsing on French, and medium-low resource languages Slovenian and Naija. We further investigate the impact of the training data size and intermediate layers of the pretrained speech encoder on speech parsing.
Figures & tables
Figure 1: Dependency parsing tree from the NaijaNSC treebank. en : Ala-… Alaska Pepper was shocked. The two lines under the wordform line are the sequence of POS and lemma tags respectively.
Dataset
Section
Duration
Trees
Speakers
Train
6.78h
7,268
67
Naija
Dev
0.87h
990
10
Test
0.91h
972
11
Train
4.88h
3,474
464
Slovenian
Dev
0.67h
400
80
Test
0.99h
358
36
Table 1: Data statistics for the Naija-NSC treebank, Slovenian SST, ParisStories, and Orféo
Figure 2: Comparison of the original architecture ( Pupier et al., 2024 ) and our proposed architecture.
Figure 3: Text parsing with frozen text encoder - with and without intermediate NN unit.
Table 2: Evaluation on test dataset with settings described in Section 5.5 . Parameter counts: w Wav2vec, g Speech_Large_fr_114K , f feedforward + LSTM, p parsing module, r Roberta module. W2T refers to Wav2tree. Best scores across end-to-end speech parsing highlighted in bold. Underlined refers to statistical significant results
Text parsing architecture
Intermediate NN
LAS
LAS
LAS
# of parameters
(LSTM)
(Naija)
(Slovenian)
(ParisStories)
( fine-tuned)
Joint fine-tuning
no
83.9
71.3
72.0
559M r +5M p
Joint fine-tuning + LSTM
yes
81.7
66.1
69.8
559M r +5M p +16M i
Frozen pre-trained text model
no
51.5
40.9
44.8
5M r
Frozen pre-trained text model + LSTM
yes
74.6
55.4
62.9
5M p +16M i
Table 3: Comparative analysis for RQ1 - LAS on the test dataset with and without intermediate NN unit for all 3 treebanks with experiments described in Section 5.1.2 . Best scores across sections highlighted in bold .
Similarity between
CKA
CKA
CKA
fine-tuned encoder &
Naija
Slovenian
French
Frozen encoder without
17%
21%
23%
intermediate NN unit
Frozen encoder with
44%
32%
34%
intermediate NN unit
Table 4: Similarity between representations from fine-tuned and frozen text encoder described in Section 6.1
Figure 4: Scaling training data across 2 setups - End-to-End (Baseline) and Frozen ASR - as described in Section 5.2
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
Baseline Wav2tree ( Pupier et al., 2024 )
Simplified Wav2tree (ours)
Epoch
80
80
Batch size
4
4
Tuning parameters
Learning rate (ASR)
0.00005
0.00005
Learning rate (Remaining)
0.0003
0.0003
Optimizer
AdamW
AdamW
Appendix
Table 5: Hyperparameters across Wav2tree architecture ( Pupier et al., 2024 ) and Simplified Wav2tree (ours).
Figure 5: Predicted word boundary and the corresponding uncollapsed sequence of the predicted transcription (in French) "J AI MON COLLÈGUE" ( en: I have my colleague). Symbols ϵ and _ represents respectively the blank and space labels. Indices {1,2,3,4} refer to the word position in the output sequence.
Index
Punctuation presence
WER ↓
UPOS ↑
UAS ↑
LAS ↑
1
Train & Eval data
34.4
76.5
62.8
57.6
2
Training data
32
78.5
68.8
61.8
3
None
31.8
79.1
69.1
61.8
Appendix
Table 6: Simplified W2T results across different variations of Naija dataset. Best scores highlighted in bold .
Figure 6: Gold and predicted transcription and trees for a Naija utterance (sent_id: WAZL_15_MC_Abi-MG__180). The predicted annotation has all UAS, UPOS and LAS as 72.7. The scores are low because not all markups are predicted by ASR - ‘(’, ‘)’ and ‘ < ’(marked as ‘<’) are not predicted by ASR.
English Vocab
Number of words
WER
1
8,991
24.5%
0
1,269
52.5%
Appendix
Table 7: WER for tokens in Naija-NSC treebank.
# of markups
Trees
Avg markups in ground_truth
Avg markups in prediction
LAS with markups
LAS without markups
LAS decrease
[0, 2]
423
1.5
1.9
66.0
66.0
0.0
[3, 4]
232
3.4
3.5
63.0
64.5
-1.4
[5, 6]
137
5.4
4.8
57.1
61.1
-3.9
[7, 12]
139
8.7
7.1
54.2
60.3
-6.1
>=13
41
16.5
12.6
44.0
54.0
-10.0
Appendix
Table 8: Parsing results comparison with and without syntax-prosody segmentation annotations.
Index
Hypothesis
p value
p value
p value
Naija
Slovenian
ParisStories
1
Null
21.6%
0.2%
3.6%
2
Alternative
-
0.1%
1.8%
Appendix
Table 9: Statistical significance testing between Simplified and Baseline architecture
Index
Hypothesis
p value
p value
p value
Naija
Slovenian
ParisStories
1
Null
21.6%
0.2%
3.6%
2
Alternative
-
0.1%
1.8%
Appendix
Table 10: Statistical significance testing between Simplified and Baseline architecture
We challenge the conventional view of neural network pruning as solely a compression technique, demonstrating that one-shot magnitude pruning serves as a powerful implicit regularizer for ASR. Using Whisper-small, we combine gradient- and Fisher-based sensitivity diagnostics with targeted, component-wise pruning. This reveals architectural asymmetries: decoder FFNs are pruning-fragile, whereas decoder self-attention and the last encoder layers contain redundancy that, when removed, improves generalization. Without fine-tuning, pruning 50% of decoder self-attention reduces WER by 2.38% absolute (20.44% relative) on LibriSpeech test-other; pruning the last four encoder layers at 50% instead yields a 1.72% absolute (14.8% relative) improvement. Gains persisted on Common Voice and TED-LIUM datasets. Beyond regularization benefits, our sensitivity-aware approach enables more aggressive one-shot compression. At 40% sparsity, where established global pruning approaches catastrophically fail, our method preserves near-baseline accuracy. This positions pruning as a first-class architectural design tool: knowing where to prune is as important as how much to prune.
Julian Irigoyen, Arthur Söhler, Andreas Søeborg Kirkedal
Speech LLMs are costly to deploy in resource-constrained settings. Layer pruning can cut this cost, but existing scoring metrics transfer poorly to speech LLMs: they assume a decoder-only architecture with homogeneous token sequences, whereas speech LLMs add encoder and adapter components and process multimodal sequences. We propose Speech Block Influence (SBI), the first layer-importance scoring framework designed for speech LLM pruning that consists of two component-specific scores: SBI-Enc measures the effect of encoder-layer removal at the adapter's output to better reflect downstream impact; SBI-Dec measures layer-wise input-output similarity over text-token positions only to avoid audio-token dominance. Across three speech LLMs, SBI improves pruning robustness, with stronger encoder performance at higher pruning rates and more reliable decoder layer selection by scoring text tokens rather than the audio-dominated full sequence. We further find that text-only calibration yields decoder rankings highly correlated with those from speech-text calibration, suggesting a cheaper alternative to measure decoder layer importance.
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers -- the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT -- across ten typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes until sufficient annotated data is available to leverage the representational capacity of pre-trained transformers.