Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS
Organizations: UX Factory, Inc.
Abstract
We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character error rate (CER) under the corpus protocol against the published 26.64, and 19.00 - 21.52 on unseen wording. Per speaker, CER spans 1.0 to 52.2%, with seen wording lowering CER by 7.0 - 9.0 points and professional delivery and spontaneous speech raising it by 8.5 - 10.5 and 12.7 points. A low-rank adapter with 4.6% of the parameters, trained on 4 to 29 minutes of the user's frontal video, lowers the CER of twelve high-error speakers by 2.13 to 3.58 points, transfers to every camera without loss, and keeps 85% of the full fine-tuning gain at 12% of its cost to other speakers. Cameras above the mouth plane add about six CER points as a constant offset that training on all views keeps small.
Figures & tables
| System (views, decoding) | Params | CER | WER |
|---|---|---|---|
| V-model [ 21 ] (not stated) | 34M | 26.64 | 47.89 |
| M 9 (all views, joint) | 232M | 9.95/12.19 | 20.47/24.08 |
| M 9 (all views, greedy) | 14.06/16.82 | 29.60/34.13 | |
| M 9 , unseen (joint) | 19.00/21.52 | 36.87/40.57 | |
| M 1 (frontal only, joint) | 232M | 9.92/12.25 | 21.08/24.60 |
| M 1 , unseen (joint) | 18.24/20.97 | 36.69/40.14 |
| Speaker | Mode, wording | greedy [95% CI] | joint | |
|---|---|---|---|---|
| ordinary | read, seen | 6,561 | 4.79 [4.6, 5.0] | 1.56 |
| ordinary | read, unseen | 1,463 | 11.82 [11.3, 12.3] | 6.41 |
| professional | read, seen | 474 | 13.31 [12.1, 14.5] | 6.88 |
| professional | read, unseen | 407 | 22.32 [21.0, 23.8] | 15.32 |
| professional | spontaneous | 2,579 | 35.03 [34.5, 35.7] | 29.10 |