Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.
Figures & tables
Dataset
Language
Hours
Samples
# speakers
License
Domain
Cased
Punctuation
Sex labels
SLR 69
Catalan
9.42
4240
36
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 76
Basque
13.86
7136
29
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 77
Galician
10.32
5587
34
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 108
Spanish
10
2507
-
CC BY 4.0
Media/broadcast
No
No
No
SLR 108
Turkish
10
2513
-
CC BY 4.0
Media/broadcast
No
No
No
SLR 94
Portuguese
3.74
871
10
CC BY 4.0
Audiobooks
No
No
Yes
Table 1: Evaluated speech datasets with licensing, transcription, and metadata characteristics. “Sex labels” indicate whether the source provides speaker sex annotations (Partial: available for a subset of samples).
Figure 1: Duration distribution.
Model
Architecture / Type
Parameters
Languages
Owner
License
CU
whisper-large-v3
Encoder-Decoder
∼ 1.5B
99+
OpenAI
Apache-2.0
✓
Voxtral-Mini-3B-2507
Speech-LLM
∼ 3B
8
Mistral AI
Apache-2.0
✓
Voxtral-Mini-4B-Realtime-2602
Causal Speech-LLM
∼ 4B
13
Mistral AI
Apache-2.0
✓
omniASR-CTC-1B-v2
Wav2Vec2 + CTC Head
∼ 1B
1600+
Meta
Apache-2.0
✓
omniASR-LLM-1B-v2
Wav2Vec2 + LLM Decoder
∼ 1B
1600+
Meta
Apache-2.0
✓
seamless-m4t-v2-large
w2v-BERT 2.0 + NLLB Decoder
∼ 2.3B
101
Meta
CC BY-NC 4.0
×
Table 2: Evaluated models’ architecture, number of parameters, language coverage, owner, and license. CU: commercial use permitted.
Model
WER ( ↓ , %)
Median RTF ( ↓ )
Median RTF×( ↑ )
Sub rate ( ↓ , %)
Del rate ( ↓ , %)
Ins rate ( ↓ , %)
scribe_v2
6.45
0.1244
8.0
4.03
0.71
1.71
seamless-m4t-v2-large
8.12
0.0900
11.1
5.97
0.88
1.27
omniASR_LLM_1B_v2
16.64
0.1200
8.3
8.52
6.78
1.34
whisper-large-v3
19.34
0.0872
11.5
12.72
4.38
2.24
omniASR_CTC_1B_v2
21.46
0.0096
103.7
12.37
7.63
1.46
Voxtral-Mini-3B-2507
23.32
0.0801
12.5
17.22
3.52
2.59
Table 3: Overall ASR performance. WER and the substitution, deletion, and insertion rates are macro-averaged across languages and expressed per reference word. The three rates sum to WER and are comparable across systems. Median RTF and RTF× are computed across utterances. Best value per column in bold; Scribe v2 is shown as a proprietary reference and excluded from the comparison.
Figure 2: Performance ( WERmacro ) vs. Efficiency (RTFx).
Figure 3: Model’s WER: (a) across languages and (b) per speaker sex. Models are ordered by macro WER.