Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.
Figures & tables
Dataset
Language
Hours
Samples
# speakers
License
Domain
Cased
Punctuation
Sex labels
SLR 69
Catalan
9.42
4240
36
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 76
Basque
13.86
7136
29
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 77
Galician
10.32
5587
34
CC BY-SA 4.0
Read speech
Yes
Yes
Yes
SLR 108
Spanish
10
2507
-
CC BY 4.0
Media/broadcast
No
No
No
SLR 108
Turkish
10
2513
-
CC BY 4.0
Media/broadcast
No
No
No
SLR 94
Portuguese
3.74
871
10
CC BY 4.0
Audiobooks
No
No
Yes
Table 1: Evaluated speech datasets with licensing, transcription, and metadata characteristics. “Sex labels” indicate whether the source provides speaker sex annotations (Partial: available for a subset of samples).
Figure 1: Duration distribution.
Model
Architecture / Type
Parameters
Languages
Owner
License
CU
whisper-large-v3
Encoder-Decoder
∼ 1.5B
99+
OpenAI
Apache-2.0
✓
Voxtral-Mini-3B-2507
Speech-LLM
∼ 3B
8
Mistral AI
Apache-2.0
✓
Voxtral-Mini-4B-Realtime-2602
Causal Speech-LLM
∼ 4B
13
Mistral AI
Apache-2.0
✓
omniASR-CTC-1B-v2
Wav2Vec2 + CTC Head
∼ 1B
1600+
Meta
Apache-2.0
✓
omniASR-LLM-1B-v2
Wav2Vec2 + LLM Decoder
∼ 1B
1600+
Meta
Apache-2.0
✓
seamless-m4t-v2-large
w2v-BERT 2.0 + NLLB Decoder
∼ 2.3B
101
Meta
CC BY-NC 4.0
×
Table 2: Evaluated models’ architecture, number of parameters, language coverage, owner, and license. CU: commercial use permitted.
Model
WER ( ↓ , %)
Median RTF ( ↓ )
Median RTF×( ↑ )
Sub rate ( ↓ , %)
Del rate ( ↓ , %)
Ins rate ( ↓ , %)
scribe_v2
6.45
0.1244
8.0
4.03
0.71
1.71
seamless-m4t-v2-large
8.12
0.0900
11.1
5.97
0.88
1.27
omniASR_LLM_1B_v2
16.64
0.1200
8.3
8.52
6.78
1.34
whisper-large-v3
19.34
0.0872
11.5
12.72
4.38
2.24
omniASR_CTC_1B_v2
21.46
0.0096
103.7
12.37
7.63
1.46
Voxtral-Mini-3B-2507
23.32
0.0801
12.5
17.22
3.52
2.59
Table 3: Overall ASR performance. WER and the substitution, deletion, and insertion rates are macro-averaged across languages and expressed per reference word. The three rates sum to WER and are comparable across systems. Median RTF and RTF× are computed across utterances. Best value per column in bold; Scribe v2 is shown as a proprietary reference and excluded from the comparison.
Figure 2: Performance ( WERmacro ) vs. Efficiency (RTFx).
Figure 3: Model’s WER: (a) across languages and (b) per speaker sex. Models are ordered by macro WER.
As automatic speech recognition (ASR) systems shift toward multilingual support and low-resource language modeling, phoneme-based layers serve as a critical language-agnostic foundation. However, most evaluations of ASR's demographic biases related to race, age, gender, and accent focus on standard grapheme-based ASR systems with comparatively little emphasis on phoneme-based systems. In this study, we evaluate the performance of WhisperIPA and ZIPA, two state-of-the-art open-source systems that generate International Phonetic Alphabet (IPA) transcriptions. Our evaluation includes existing multilingual speech corpora and demographically annotated English-language corpora, comparing model-generated IPA transcriptions against grapheme-to-phoneme (G2P) systems using both standard phoneme error rate (PER) and a proposed Soft PER metric that tolerates linguistically similar phoneme substitutions. Our analysis examines how performance varies across language, gender, accent, ethnicity, and age, revealing persistent disparities even after accounting for acceptable phonemic variation. These findings, while limited, provide insight into potential sources of bias and inform the development of more inclusive and linguistically robust phoneme-based ASR systems. Our code and data are publicly available.
Code-switching -- the natural alternation between two languages within a single utterance -- remains one of the most challenging and under-studied conditions for automatic speech recognition (ASR). We present a benchmark evaluating five commercial ASR providers across four language pairs: Egyptian Arabic--English, Saudi Arabic (Najdi/Hijazi)--English, Persian (Farsi)--English, and German--English, comprising 300 samples per pair selected by a two-stage pipeline combining heuristic filtering with a GPT-4o and Gemini 1.5 Pro ensemble scorer, reducing LLM costs by ≈91%. We evaluate on both WER and BERTScore, showing that while both metrics agree on the ordinal ranking of systems for all Arabic and Persian pairs (τ=1.0), WER inflates the magnitude of quality gaps by approximately 3× by penalising semantically correct transliteration choices. ElevenLabs Scribe v2 achieves the lowest WER (13.2% overall) and leads on BERTScore (0.936 overall). Difficulty-stratified analysis reveals performance gaps masked by aggregate averages, and BERT embedding projections confirm semantic proximity between reference and hypothesis despite surface-level script differences. The dataset is publicly available at https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.
Sajjad Abdoli, Ghassan Al-Sumaidaee, Clayton W. Taylor +2
Advances in deep learning and end-to-end Automatic Speech Recognition (ASR) have enabled robust multilingual models, but evaluation metrics remain limited in assessing accuracy. Efforts to improve or replace the common metric Word Error Rate (WER) often focus on English, leaving evaluations for low-resource languages under-explored and hindering fair cross-lingual comparisons. We present OpenWER, an open-source implementation that improves WER robustness through language-specific normalisation and compound word detection. A token-based Levenshtein alignment preserves complementary metrics and allows metadata embedding for granular accuracy scores. Our analysis of 52 languages shows absolute WER reductions of up to 25% compared to common libraries. OpenWER contributes to fairness in ASR research by increasing the reliability of WER across diverse languages and enabling more comprehensive accuracy evaluations.
Korbinian Kuhn, Gottfried Zimmermann
Stuttgart Media University, Germany · University of Tübingen, Germany