Efficient Parallel Audio Generation using Group Masked Language Modeling
Jeong, Myeonghun, Kim, Minchan, Lee, Joun Yeop, Kim, Nam Soo
–arXiv.org Artificial Intelligence
We present a fast and high-quality codec language model for parallel audio generation. While SoundStorm, a state-of-the-art parallel audio generation model, accelerates inference speed compared to autoregressive models, it still suffers from slow inference due to iterative sampling. To resolve this problem, we propose Group-Masked Language Modeling~(G-MLM) and Group Iterative Parallel Decoding~(G-IPD) for efficient parallel audio generation. Both the training and sampling schemes enable the model to synthesize high-quality audio with a small number of iterations by effectively modeling the group-wise conditional dependencies. In addition, our model employs a cross-attention-based architecture to capture the speaker style of the prompt voice and improves computational efficiency. Experimental results demonstrate that our proposed model outperforms the baselines in prompt-based audio generation.
arXiv.org Artificial Intelligence
Jan-2-2024
- Country:
- North America > United States
- Minnesota > Hennepin County > Minneapolis (0.14)
- Europe > Italy
- Calabria > Catanzaro Province > Catanzaro (0.04)
- Asia > South Korea
- North America > United States
- Genre:
- Research Report > New Finding (0.34)
- Technology: