Generate Natural Sounding Speech from Text in Real-Time NVIDIA Developer Blog
This blog, intended for developers with professional level understanding of Deep Learning, will help you produce a production ready AI text-to-speech model. Converting text into high quality, natural sounding speech in real-time has been a challenging task for decades. State-of-the-art speech synthesis models are based on parametric neural networks1. Text-to-speech (TTS) synthesis is typically done in two steps. The optimized Tacotron2 model2 and the new WaveGlow model1 take advantage of Tensor Cores on NVIDIA Volta and Turing GPUs to convert text into high quality natural sounding speech in real-time.
Sep-12-2019, 19:11:45 GMT