Multi-Task Learning Framework for Simultaneous Speech Enhancement and Keyword Spotting
Keywords:
Multi-task learning, speech enhancement, keyword spotting, noise-robust speech recognition, deep neural networks, acoustic signal processing, low-SNR speech processing.Abstract
Strong Keyword spotting (KWS) in impaired acoustic conditions has been a primary challenge of real-time speech-driven systems especially with a low signal-to-noise ratio (SNR) signal-noise-disturbance conditions. Traditional cascaded pipelines implement speech enhancement (SE) first, before keyword detection, but the cascaded pipeline approach is associated with error propagation, higher latency, and unnecessarily compute features. The current paper suggests that a common multi-task learning (MTL) framework used to do simultaneous speech enhancement and key-word spotting with a shared acoustic encoder and task-specific streams of the encoder can be developed. Through a combined reconstruction and classification objective achieved by adaptive loss-balancing mechanism, the proposed model acquires noise-resistant representations that can be helpful in both tasks. The experiments on the Google Speech Commands dataset with the addition of real-world noise at 0 at 10 dB SNR prove that the suggested framework has been shown to beat standalone and cascaded baselines. In particular, the model results in as many as 57% absolute KWS accuracy improvement at 0 dB SNR and remaining competitive in terms of their enhancement quality, based on PESQ, STOI, and SI-SNR. More so, feature extraction performs shared features eliminates redundant computation, and real-time inference can be made without making substantial complexities to the model. These findings confirm the usefulness of joint optimization towards resilient speech-triggered system and point to the promise of multi-task learning on efficient and noise-resistant speech processing system.