Deep Metric Learning-Based Speaker Identification under Channel Variability

Authors

  • V.Ramya Assistant professor, Department of CSE, Excel Engineering college, Kumarapalayam, Namakkal Author

Keywords:

Speaker Identification, Deep Metric Learning, Channel Variability, Speaker Embeddings, Channel-Invariant Representation, Robust Speech Processing

Abstract

Public deployment systems that identify speakers often suffer a reduction in performance because of channel variability as a result of microphone diversity, codec compression, bandwidth constraint and external effects. Traditional embedding-based methods used, which are trained by using softmax, mainly optimise inter speaker classification but do not directly encourage channel induced distribution shift resistance. As a result, embedding of the same speaker made in varying channel conditions is usually characterised by higher intra-class discrepancy, which decreases the consistency of the recognition test in mispaired conditions. A channel-aware hostile framework of deep learning in speaker identification due to the variability of channel is proposed in this paper. An objective is reformulated as a triplet to promote cross-channel anchor positive pairing which forces embeddings of identical speakers to stay near-by despite channel differences. We also introduce a term of channel discrepancy regularisation to make a clear decrease of embedding driftiness across different channel environment conditions. The parameter minimally loss encourages compact and channel-invariant clusters of speakers and does not decrease inter-speaker separability. A lot of experiments are realised with the help of controlled channel distortion modelling, which also comprises bandwidth truncation, codec artefacts, and additive noise. Findings indicate a steady enhancement in the strength as the severity of distortion becomes higher. The proposed approach has significant improvements in mismatched identification accuracy and significant decreases in equal error rate when compared to softmax and traditional triplet-loss baselines, notwithstanding performing similarly when the data is clean. These results demonstrate the efficiency of the usage of channel-invariant constraints in the framework of metric learning in order to perform a trustworthy real-life speaker recognition.

Downloads

Download data is not yet available.

Additional Files

Published

2026-07-21

Issue

Section

Articles

How to Cite

V.Ramya. (2026). Deep Metric Learning-Based Speaker Identification under Channel Variability. National Journal of Speech and Audio Signal Processing, 2(4), 38-45. https://ecejournals.in/index.php/NJSAP/article/view/560