Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.
Figures & tables
AlignBench
TIMIT
Buckeye
Speech style
mixed
read
conversational
Recording
mixed
booth/headset
quiet/headset
Unique speakers
73
630
40
Non-American accent (%)
21.9
0.0
0.0
Clips with >1 speaker (%)
27.4
0.0
0.0
Difficult words (%)
9.6
2.4
5.9
Table 1: Evaluation-corpus statistics.
AlignBench
TIMIT
Buckeye
Speech style
mixed
read
conversational
Recording
mixed
booth/headset
quiet/headset
Unique speakers
73
630
40
Non-American accent (%)
21.9
0.0
0.0
Clips with >1 speaker (%)
27.4
0.0
0.0
Difficult words (%)
9.6
2.4
5.9
Table 1: Evaluation-corpus statistics.
Dataset
ASR
Sub./Del./Ins. ↓
WER ↓
CER ↓
AlignBench
Scribe
0.71 / 0.50 / 0.66
1.88
1.16
Whisper
3.14 / 8.76 / 2.76
14.65
9.78
TIMIT
Scribe
1.12 / 0.35 / 0.23
1.71
0.38
Whisper
2.14 / 0.39 / 0.23
2.76
0.79
Buckeye
Scribe
5.29 / 1.75 / 5.13
12.16
6.56
Whisper
4.69 / 8.02 / 3.38
16.09
10.16
Table 2: ASR errors (%); bold marks the lower WER/CER per dataset.
Dataset
System
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
Mean asym. (ms) ↓
F1 (%) ↑
AlignBench
FuseAlign
0.38 / 0.43 / 1.71
7.27 / 7.18 / 7.74
29.82 / 29.87 / 29.77
17.16 / 20.43 / 32.82
99.82 / 99.28 / 96.83
Gentle
2.11 / 2.04 / 3.67
12.61 / 12.53 / 13.48
35.20 / 35.13 / 35.64
31.43 / 33.68 / 51.46
98.90 / 98.26 / 97.31
MFA
4.15 / 4.89 / 10.82
34.16 / 34.52 / 38.56
64.92 / 65.18 / 67.18
73.00 / 79.80 / 374.26
100.00 / 98.20 / 95.20
Nyra
1.28 / 1.26 / 2.57
29.99 / 29.85 / 28.88
68.84 / 68.77 / 68.02
33.18 / 33.42 / 56.90
100.00 / 99.31 / 96.81
MMS-FA
2.49 / 2.43 / 3.84
46.69 / 46.85 / 45.73
80.74 / 81.14 / 80.73
48.91 / 50.44 / 73.89
98.96 / 99.31 / 96.82
WhisperX
1.41 / 1.66 / 3.57
53.25 / 53.16 / 52.44
84.75 / 84.69 / 84.43
49.04 / 49.84 / 90.38
99.78 / 99.33 / 96.81
Table 3: Main results. Each cell is hand-corrected / Scribe / Whisper for the same audio. ↓ lower is better, ↑ higher is better. We score exact ASR/reference word matches only, with the uniform fallback for unplaced words (Section 3 ). Within each dataset and at each column position, bold marks the best system and every system statistically tied with it: each point estimate lies within the other’s 95% word-level confidence interval. F1 has no interval, so only its best value is bold.
Ablation
Change from the baseline
Mean asym. (ms) ↓
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
–
baseline ( ∼213 M params)
17.16 ± 1.92
0.38 ± 0.13
7.3 ± 0.6
29.8 ± 1.0
1
width = 512, FFN = 2048 ( ∼60 M params)
17.75 ± 2.49
0.35 ± 0.13
7.1 ± 0.6
29.6 ± 1.0
2
24 layers ( ∼415 M params)
17.01 ± 1.91
0.33 ± 0.13
7.4 ± 0.6
30.1 ± 1.0
3
width = 2048, FFN = 8192 ( ∼821 M params)
22.23 ± 4.62
0.33 ± 0.13
7.3 ± 0.6
30.0 ± 1.0
4
window = 1024, hop = 512 samples
17.77 ± 2.48
0.32 ± 0.12
7.3 ± 0.6
29.9 ± 1.0
5
learned character embeddings (instead of frozen Dia)
17.67 ± 1.94
0.40 ± 0.14
7.8 ± 0.6
30.0 ± 1.0
Table 4: Model ablations on AlignBench (hand-corrected transcript). Parameter counts cover only the trainable alignment model and exclude the frozen 252M-parameter Dia encoder. Small ± values are 95% word-level confidence half-widths.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
FuseAlign
Gentle
MFA
Nyra
MMS-FA
WhisperX
AlignBench
21.51 / 25.46 / 38.11
37.59 / 39.74 / 57.44
80.46 / 87.38 / 383.63
37.83 / 38.31 / 61.88
80.71 / 81.35 / 99.57
54.20 / 53.55 / 94.85
TIMIT
27.65 / 27.60 / 27.66
28.82 / 28.63 / 28.69
23.23 / 23.26 / 23.33
23.06 / 23.10 / 23.13
37.17 / 41.06 / 48.05
46.16 / 45.80 / 44.87
Buckeye
30.51 / 27.62 / 37.73
48.24 / 41.57 / 53.81
42.27 / 38.65 / 112.51
27.66 / 25.08 / 43.32
53.22 / 70.87 / 80.32
59.80 / 55.44 / 76.08
Appendix
Table 5: Mean symmetric boundary error (ms) ↓ for every evaluation dataset and aligner. Each cell is hand-corrected / Scribe / Whisper , matching Table 3 ; bold marks the best system at each transcript position within a dataset.
System
Mean asym. (ms) ↓
Mean sym. (ms) ↓
% >25 ms ↓
% >50 ms ↓
% >300 ms ↓
MFA
24.75
28.83
44.70
21.22
0.89
FuseAlign
19.38
26.65
47.81
8.36
0.20
Appendix
Table 6: MFA-style Buckeye utterances. Both systems are rerun on our reconstruction of the segmentation used by McAuliffe et al. (2026) . Mean asymmetric and symmetric errors pool the two boundaries of every scored word; symmetric error retains displacement within silence. A threshold failure means the maximum asymmetric error over a word’s start and end exceeds the stated tolerance. All values use ground-truth transcripts and all-word scoring.
Mean asym. error (ms) ↓
Words >300 ms (%) ↓
Audio sampled (h)
EMA on
EMA off
Difference
EMA on
EMA off
Difference
7,822
26.46
24.53
-1.93
0.84
0.88
+0.04
15,644
18.52
19.43
+0.91
0.43
0.45
+0.01
23,467
20.01
18.64
-1.38
0.38
0.47
+0.09
31,289
18.33
19.63
+1.30
0.41
0.52
+0.11
39,111
17.43
17.87
+0.44
0.35
0.51
+0.16
Appendix
Table 7: Training dynamics on AlignBench with the hand-corrected transcript at every saved checkpoint, with EMA-snapshot label correction on (baseline) and off (ablation 9). We report both mean asymmetric error and the percentage of words with error above 300 ms. Rows are the hours of audio sampled so far; the rule marks where correction starts at 20k steps. Each difference is off minus on, and the best checkpoint in each on/off column is bold.
Protocol
System
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
Mean asym. (ms) ↓
F1 (%) ↑
baseline
FuseAlign
0.38
7.27
29.82
17.16
99.82
Gentle
2.11
12.61
35.20
31.43
98.90
MFA
4.15
34.16
64.92
73.00
100.00
Nyra
1.28
29.99
68.84
33.18
100.00
MMS-FA
2.49
46.69
80.74
48.91
98.96
WhisperX
1.41
53.25
84.75
49.04
99.78
Appendix
Table 8: Protocol conditions on AlignBench with the hand-corrected transcript. FuseAlign always chunks, so its baseline is already chunked and it has no chunking row. Error-rate changes are coloured when their magnitude is at least 20% of the system’s baseline rate; mean-error changes are coloured at 2 ms and F1 changes at 0.5 points. Green improves and red degrades.
Split
Metric
FuseAlign
Gentle
MFA
Nyra
MMS-FA
WhisperX
Accent n=6371/1701
% >300 ms ↓
0.30 / 0.71
2.07 / 2.23
3.53 / 6.47
1.19 / 1.59
2.26 / 3.35
1.41 / 1.41
% >50 ms ↓
7.3 / 7.1
12.2 / 14.1
34.5 / 32.7
30.4 / 28.3
47.7 / 42.7
54.0 / 50.3
% >25 ms ↓
30.2 / 28.6
34.8 / 36.5
65.0 / 64.5
68.8 / 69.0
81.2 / 79.1
85.2 / 83.0
Mean asym. (ms) ↓
17.41 / 16.22
32.65 / 26.86
66.59 / 97.03
32.98 / 33.92
48.18 / 51.66
49.83 / 46.11
F1 (%) ↑
99.83 / 99.76
98.97 / 98.63
100.00 / 100.00
100.00 / 100.00
98.89 / 99.23
99.76 / 99.82
Speaker count n=4851/3221
% >300 ms ↓
0.06 / 0.87
1.20 / 3.48
1.79 / 7.70
0.80 / 1.99
1.61 / 3.82
0.78 / 2.36
Appendix
Table 9: Alignment performance by AlignBench subgroup with hand-corrected transcripts. Each cell is easy / difficult : American/non-American for accent, single-/multi-speaker for speaker count, and ordinary/difficult for word type. The n values give the number of scored words on the two sides. Bold marks the best system and underline the second best separately on each side of each split.
Ablation
Change from the baseline
Mean asym. (ms) ↓
% >300 ms ↓
% >50 ms ↓
% >25 ms ↓
–
baseline ( ∼213 M params)
17.16 / 20.43 / 32.82
0.38 / 0.43 / 1.71
7.3 / 7.2 / 7.7
29.8 / 29.9 / 29.8
1
width = 512, FFN = 2048 ( ∼60 M params)
17.75 / 18.01 / 28.17
0.35 / 0.44 / 1.50
7.1 / 7.1 / 7.6
29.6 / 29.5 / 30.0
2
24 layers ( ∼415 M params)
17.01 / 17.04 / 30.90
0.33 / 0.33 / 1.61
7.4 / 7.5 / 7.9
30.1 / 30.3 / 30.3
3
width = 2048, FFN = 8192 ( ∼821 M params)
22.23 / 21.44 / 34.05
0.33 / 0.35 / 1.50
7.3 / 7.2 / 7.7
30.0 / 29.8 / 30.0
4
window = 1024, hop = 512 samples
17.77 / 17.13 / 29.06
0.32 / 0.41 / 1.72
7.3 / 7.2 / 7.9
29.9 / 29.6 / 30.3
5
learned character embeddings (instead of frozen Dia)
17.67 / 18.37 / 27.56
0.40 / 0.45 / 1.70
7.8 / 7.8 / 8.0
30.0 / 29.9 / 30.0
Appendix
Table 10: Model ablations on AlignBench under hand-corrected and noisy transcripts. Each cell is hand-corrected / Scribe / Whisper for the same audio. Lower is better. ASR conditions score exact ASR - reference word matches only, following Section 3 .