MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
Fu, Szu-Wei, Yu, Cheng, Hsieh, Tsun-An, Plantinga, Peter, Ravanelli, Mirco, Lu, Xugang, Tsao, Yu
–arXiv.org Artificial Intelligence
The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Objective evaluation metrics which consider human perception can hence serve as a bridge to reduce the gap. Our previously proposed MetricGAN was designed to optimize objective metrics by connecting the metric with a discriminator. Because only the scores of the target evaluation functions are needed during training, the metrics can even be non-differentiable. In this study, we propose a MetricGAN+ in which three training techniques incorporating domain-knowledge of speech processing are proposed. With these techniques, experimental results on the VoiceBank-DEMAND dataset show that MetricGAN+ can increase PESQ score by 0.3 compared to the previous MetricGAN and achieve state-of-the-art results (PESQ score = 3.15).
arXiv.org Artificial Intelligence
Apr-8-2021
- Country:
- North America
- United States > Ohio
- Franklin County > Columbus (0.04)
- Canada > Quebec
- Montreal (0.04)
- United States > Ohio
- Asia
- Taiwan > Taiwan Province
- Taipei (0.04)
- Japan > Honshū
- Kansai > Kyoto Prefecture > Kyoto (0.04)
- Taiwan > Taiwan Province
- North America
- Genre:
- Research Report (0.70)
- Technology: