Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language
–arXiv.org Artificial Intelligence
End-to-end text-to-speech synthesis (TTS) can generate highly natural synthetic speech from raw text. However, rendering the correct pitch accents is still a challenging problem for end-to-end TTS. To tackle the challenge of rendering correct pitch accent in Japanese end-to-end TTS, we adopt PnG~BERT, a self-supervised pretrained model in the character and phoneme domain for TTS. We investigate the effects of features captured by PnG~BERT on Japanese TTS by modifying the fine-tuning condition to determine the conditions helpful inferring pitch accents. We manipulate content of PnG~BERT features from being text-oriented to speech-oriented by changing the number of fine-tuned layers during TTS. In addition, we teach PnG~BERT pitch accent information by fine-tuning with tone prediction as an additional downstream task. Our experimental results show that the features of PnG~BERT captured by pretraining contain information helpful inferring pitch accent, and PnG~BERT outperforms baseline Tacotron on accent correctness in a listening test.
arXiv.org Artificial Intelligence
Dec-16-2022
- Country:
- Oceania > Australia
- Queensland > Brisbane (0.04)
- North America
- United States
- New Jersey > Middlesex County
- Piscataway (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Arizona > Maricopa County
- Scottsdale (0.04)
- New Jersey > Middlesex County
- Canada
- Ontario > Toronto (0.04)
- British Columbia > Metro Vancouver Regional District
- Vancouver (0.04)
- United States
- Europe
- United Kingdom > England
- East Sussex > Brighton (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Italy > Tuscany
- Florence (0.04)
- Austria > Styria
- Graz (0.04)
- United Kingdom > England
- Asia
- Oceania > Australia
- Genre:
- Research Report > New Finding (1.00)
- Technology: