Cross-Lingual Self-Supervised Learning for Low-Resource Speech Recognition
Keywords:
Automatic speech recognition (ASR), self-supervised learning (SSL), cross-lingual transfer, low-resource languages, multilingual acoustic modeling, wav2vec 2.0, HuBERT, XLS-R, speech representation learning, connectionist temporal classification (CTC), word error rate (WER), phonetic alignment.Abstract
Performance in automatic speech recognition (ASR) has only reached close to human levels in high-resource languages; extending to low-resource languages is also an important issue since there are small amounts of labelled speech. Self-supervised learning (SSL) has become an encouraging learning model capable of transferring universal acoustic representations of large-scale unlabeled corpora. This paper research seeks to explore the success of cross-lingual transfer in the case of low-resource ASR through the use of SSL models in the case of extreme data scarcity. We compare monolingual and multilingual pretrained models on typologically diverse target languages on 1-hour, 5-hour and 10-hour of labelled data. Particularly, wav2vec 2.0, HuBERT and XLS-R are fine-tuned on a Connectionist Temporal Classification (CTC) structure and evaluated based on word error rate (WER), convergence behaviour, and phonetic error analysis. Through experimental evidence, the efficiency of multilingual pretraining is demonstrated to be far more efficient in facilitating adaptation than monolingual baselines and up to 28 percent smaller relative WER relative to monolingual baselines in ultra-low-resource settings. Representation analysis also shows better phonetic matching between languages which means multi linguistic SSL transfers language independent acoustic features that assist transfer. The results show that cross-lingual self-supervised training is a cost-effective and scalable model of training ASR systems when limited resources are available in that language. These findings emphasise the future of multilingual pretraining acoustics in minimising data dependent and speeding up the implementation of speech technology in the linguistically diverse area.