Controllable speech synthesis by learning discrete phoneme-level prosodic representations

Open in new window