Adversarial Speaker Verification

Meng, Zhong, Zhao, Yong, Li, Jinyu, Gong, Yifan

arXiv.org Machine Learning 

In speech area, it has been applied The use of deep networks to extract embeddings for speaker to speech enhancement [11, 12, 13], voice conversion [14], acoustic recognition has proven successfully. However, such embeddings model adaptation [15, 16, 17], noise-robust [18, 19], speakerinvariant are susceptible to performance degradation due to the mismatches [20, 21] automatic speech recognition, speaker model adaptation among the training, enrollment, and test conditions. In this work, we [22] and speech enhancement [23, 11, 24] using gradient reversal propose an adversarial speaker verification (ASV) scheme to learn layer (GRL) [25]. In these works, adversarial learning is the condition-invariant deep embedding via adversarial multi-task used to learn an intermediate representation in a DNN that is invariant training. In ASV, a speaker classification network and a condition to the shift among different conditions (e.g., environments, identification network are jointly optimized to minimize the speaker speakers, SNRs, etc.). To benefit from this, in this work, we propose classification loss and simultaneously mini-maximize the condition adversarial speaker verification (ASV) to suppress the effects loss. The target labels of the condition network can be categorical of condition variability in speaker modeling for robust SV.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found