Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.
Figures & tables
AlignBench
TIMIT
Buckeye
Speech style
mixed
read
conversational
Recording
mixed
booth/headset
quiet/headset
Unique speakers
73
630
40
Non-American accent (%)
21.9
0.0
0.0
Clips with >1 speaker (%)
27.4
0.0
0.0
Difficult words (%)
9.6
2.4
5.9
Table 1: Evaluation-corpus statistics.
AlignBench
TIMIT
Buckeye
Speech style
mixed
read
conversational
Recording
mixed
booth/headset
quiet/headset
Unique speakers
73
630
40
Non-American accent (%)
21.9
0.0
0.0
Clips with >1 speaker (%)
27.4
0.0
0.0
Difficult words (%)
9.6
2.4
5.9
Table 1: Evaluation-corpus statistics.
Dataset
ASR
Sub./Del./Ins. ↓
WER ↓
CER ↓
AlignBench
Scribe
0.71 / 0.50 / 0.66
1.88
1.16
Whisper
3.14 / 8.76 / 2.76
14.65
9.78
TIMIT
Scribe
1.12 / 0.35 / 0.23
1.71
0.38
Whisper
2.14 / 0.39 / 0.23
2.76
0.79
Buckeye
Scribe
5.29 / 1.75 / 5.13
12.16
6.56
Whisper
4.69 / 8.02 / 3.38
16.09
10.16
Table 2: ASR errors (%); bold marks the lower WER/CER per dataset.
Dataset
System
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
Mean asym. (ms) ↓
F1 (%) ↑
AlignBench
FuseAlign
0.38 / 0.43 / 1.71
7.27 / 7.18 / 7.74
29.82 / 29.87 / 29.77
17.16 / 20.43 / 32.82
99.82 / 99.28 / 96.83
Gentle
2.11 / 2.04 / 3.67
12.61 / 12.53 / 13.48
35.20 / 35.13 / 35.64
31.43 / 33.68 / 51.46
98.90 / 98.26 / 97.31
MFA
4.15 / 4.89 / 10.82
34.16 / 34.52 / 38.56
64.92 / 65.18 / 67.18
73.00 / 79.80 / 374.26
100.00 / 98.20 / 95.20
Nyra
1.28 / 1.26 / 2.57
29.99 / 29.85 / 28.88
68.84 / 68.77 / 68.02
33.18 / 33.42 / 56.90
100.00 / 99.31 / 96.81
MMS-FA
2.49 / 2.43 / 3.84
46.69 / 46.85 / 45.73
80.74 / 81.14 / 80.73
48.91 / 50.44 / 73.89
98.96 / 99.31 / 96.82
WhisperX
1.41 / 1.66 / 3.57
53.25 / 53.16 / 52.44
84.75 / 84.69 / 84.43
49.04 / 49.84 / 90.38
99.78 / 99.33 / 96.81
Table 3: Main results. Each cell is hand-corrected / Scribe / Whisper for the same audio. ↓ lower is better, ↑ higher is better. We score exact ASR/reference word matches only, with the uniform fallback for unplaced words (Section 3 ). Within each dataset and at each column position, bold marks the best system and every system statistically tied with it: each point estimate lies within the other’s 95% word-level confidence interval. F1 has no interval, so only its best value is bold.
Ablation
Change from the baseline
Mean asym. (ms) ↓
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
–
baseline ( ∼213 M params)
17.16 ± 1.92
0.38 ± 0.13
7.3 ± 0.6
29.8 ± 1.0
1
width = 512, FFN = 2048 ( ∼60 M params)
17.75 ± 2.49
0.35 ± 0.13
7.1 ± 0.6
29.6 ± 1.0
2
24 layers ( ∼415 M params)
17.01 ± 1.91
0.33 ± 0.13
7.4 ± 0.6
30.1 ± 1.0
3
width = 2048, FFN = 8192 ( ∼821 M params)
22.23 ± 4.62
0.33 ± 0.13
7.3 ± 0.6
30.0 ± 1.0
4
window = 1024, hop = 512 samples
17.77 ± 2.48
0.32 ± 0.12
7.3 ± 0.6
29.9 ± 1.0
5
learned character embeddings (instead of frozen Dia)
17.67 ± 1.94
0.40 ± 0.14
7.8 ± 0.6
30.0 ± 1.0
Table 4: Model ablations on AlignBench (hand-corrected transcript). Parameter counts cover only the trainable alignment model and exclude the frozen 252M-parameter Dia encoder. Small ± values are 95% word-level confidence half-widths.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
FuseAlign
Gentle
MFA
Nyra
MMS-FA
WhisperX
AlignBench
21.51 / 25.46 / 38.11
37.59 / 39.74 / 57.44
80.46 / 87.38 / 383.63
37.83 / 38.31 / 61.88
80.71 / 81.35 / 99.57
54.20 / 53.55 / 94.85
TIMIT
27.65 / 27.60 / 27.66
28.82 / 28.63 / 28.69
23.23 / 23.26 / 23.33
23.06 / 23.10 / 23.13
37.17 / 41.06 / 48.05
46.16 / 45.80 / 44.87
Buckeye
30.51 / 27.62 / 37.73
48.24 / 41.57 / 53.81
42.27 / 38.65 / 112.51
27.66 / 25.08 / 43.32
53.22 / 70.87 / 80.32
59.80 / 55.44 / 76.08
Appendix
Table 5: Mean symmetric boundary error (ms) ↓ for every evaluation dataset and aligner. Each cell is hand-corrected / Scribe / Whisper , matching Table 3 ; bold marks the best system at each transcript position within a dataset.
System
Mean asym. (ms) ↓
Mean sym. (ms) ↓
% >25 ms ↓
% >50 ms ↓
% >300 ms ↓
MFA
24.75
28.83
44.70
21.22
0.89
FuseAlign
19.38
26.65
47.81
8.36
0.20
Appendix
Table 6: MFA-style Buckeye utterances. Both systems are rerun on our reconstruction of the segmentation used by McAuliffe et al. (2026) . Mean asymmetric and symmetric errors pool the two boundaries of every scored word; symmetric error retains displacement within silence. A threshold failure means the maximum asymmetric error over a word’s start and end exceeds the stated tolerance. All values use ground-truth transcripts and all-word scoring.
Mean asym. error (ms) ↓
Words >300 ms (%) ↓
Audio sampled (h)
EMA on
EMA off
Difference
EMA on
EMA off
Difference
7,822
26.46
24.53
-1.93
0.84
0.88
+0.04
15,644
18.52
19.43
+0.91
0.43
0.45
+0.01
23,467
20.01
18.64
-1.38
0.38
0.47
+0.09
31,289
18.33
19.63
+1.30
0.41
0.52
+0.11
39,111
17.43
17.87
+0.44
0.35
0.51
+0.16
Appendix
Table 7: Training dynamics on AlignBench with the hand-corrected transcript at every saved checkpoint, with EMA-snapshot label correction on (baseline) and off (ablation 9). We report both mean asymmetric error and the percentage of words with error above 300 ms. Rows are the hours of audio sampled so far; the rule marks where correction starts at 20k steps. Each difference is off minus on, and the best checkpoint in each on/off column is bold.
Protocol
System
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
Mean asym. (ms) ↓
F1 (%) ↑
baseline
FuseAlign
0.38
7.27
29.82
17.16
99.82
Gentle
2.11
12.61
35.20
31.43
98.90
MFA
4.15
34.16
64.92
73.00
100.00
Nyra
1.28
29.99
68.84
33.18
100.00
MMS-FA
2.49
46.69
80.74
48.91
98.96
WhisperX
1.41
53.25
84.75
49.04
99.78
Appendix
Table 8: Protocol conditions on AlignBench with the hand-corrected transcript. FuseAlign always chunks, so its baseline is already chunked and it has no chunking row. Error-rate changes are coloured when their magnitude is at least 20% of the system’s baseline rate; mean-error changes are coloured at 2 ms and F1 changes at 0.5 points. Green improves and red degrades.
Split
Metric
FuseAlign
Gentle
MFA
Nyra
MMS-FA
WhisperX
Accent n=6371/1701
% >300 ms ↓
0.30 / 0.71
2.07 / 2.23
3.53 / 6.47
1.19 / 1.59
2.26 / 3.35
1.41 / 1.41
% >50 ms ↓
7.3 / 7.1
12.2 / 14.1
34.5 / 32.7
30.4 / 28.3
47.7 / 42.7
54.0 / 50.3
% >25 ms ↓
30.2 / 28.6
34.8 / 36.5
65.0 / 64.5
68.8 / 69.0
81.2 / 79.1
85.2 / 83.0
Mean asym. (ms) ↓
17.41 / 16.22
32.65 / 26.86
66.59 / 97.03
32.98 / 33.92
48.18 / 51.66
49.83 / 46.11
F1 (%) ↑
99.83 / 99.76
98.97 / 98.63
100.00 / 100.00
100.00 / 100.00
98.89 / 99.23
99.76 / 99.82
Speaker count n=4851/3221
% >300 ms ↓
0.06 / 0.87
1.20 / 3.48
1.79 / 7.70
0.80 / 1.99
1.61 / 3.82
0.78 / 2.36
Appendix
Table 9: Alignment performance by AlignBench subgroup with hand-corrected transcripts. Each cell is easy / difficult : American/non-American for accent, single-/multi-speaker for speaker count, and ordinary/difficult for word type. The n values give the number of scored words on the two sides. Bold marks the best system and underline the second best separately on each side of each split.
Ablation
Change from the baseline
Mean asym. (ms) ↓
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
–
baseline ( ∼213 M params)
17.16 / 20.43 / 32.82
0.38 / 0.43 / 1.71
7.3 / 7.2 / 7.7
29.8 / 29.9 / 29.8
1
width = 512, FFN = 2048 ( ∼60 M params)
17.75 / 18.01 / 28.17
0.35 / 0.44 / 1.50
7.1 / 7.1 / 7.6
29.6 / 29.5 / 30.0
2
24 layers ( ∼415 M params)
17.01 / 17.04 / 30.90
0.33 / 0.33 / 1.61
7.4 / 7.5 / 7.9
30.1 / 30.3 / 30.3
3
width = 2048, FFN = 8192 ( ∼821 M params)
22.23 / 21.44 / 34.05
0.33 / 0.35 / 1.50
7.3 / 7.2 / 7.7
30.0 / 29.8 / 30.0
4
window = 1024, hop = 512 samples
17.77 / 17.13 / 29.06
0.32 / 0.41 / 1.72
7.3 / 7.2 / 7.9
29.9 / 29.6 / 30.3
5
learned character embeddings (instead of frozen Dia)
17.67 / 18.37 / 27.56
0.40 / 0.45 / 1.70
7.8 / 7.8 / 8.0
30.0 / 29.9 / 30.0
Appendix
Table 10: Model ablations on AlignBench under hand-corrected and noisy transcripts. Each cell is hand-corrected / Scribe / Whisper for the same audio. Lower is better. ASR conditions score exact ASR - reference word matches only, following Section 3 .
Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the encoder must learn the alignment internally without cross-attention or a transducer lattice. In practice, this alignment often forms abruptly in the upper layers, making training sensitive and brittle on long utterances. We propose InterAligner, which adds an intermediate Aligner objective so alignment can form progressively across depth, together with an intermediate CTC loss (InterCTC) to stabilize optimization. On LibriSpeech with a 17-layer Conformer, a final-only Aligner reaches 5.0/7.8 WER (test-clean/other). InterCTC improves to 3.4/6.0, and InterAligner further reduces WER to 3.1/5.6 with the largest gains on long utterances.
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.
Zhiheng Qian, Aini Li, Hai Hu +1
Shanghai Jiao Tong University · City University of Hong Kong · The Hong Kong Polytechnic University +1
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an alignment by construction, whereas attention-based encoder-decoders (AED) and speech large language models (LLMs) do not, and their word timings are usually read off the attention weights instead. All of these signals live on the encoder frame grid, which bounds their temporal precision. We study a generic gradient-based alignment that applies to any differentiable ASR model. We take the gradient of each teacher-forced token log probability with respect to the input, reduce it to a per-frame saliency, and decode the resulting matrix into word boundaries with a single dynamic-programming pass. The method needs no training, no model modification and no alignment heads, works across all model families including the speech LLMs, and aligns on the input grid rather than on the coarser encoder grid. We evaluate it on sixteen models from four families, on read (TIMIT) and spontaneous (Buckeye) speech, each against the model's own native or attention-based alignment. We find that the gradient yields a usable alignment for every model, that it is usually somewhat behind a strong native aligner but better where the native alignment is weak, as for the streaming models, and that its main disadvantage is the cost of one backward pass per token.
Albert Zeyer, Ralf Schlüter, Hermann Ney
Machine Learning and Human Language Technology Group, RWTH Aachen University, Aachen, Germany · AppTek GmbH, Aachen, Germany