LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
Qu, Leyuan, Weber, Cornelius, Wermter, Stefan
–arXiv.org Artificial Intelligence
The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2 which consists of an encoder-decoder architecture and location-aware attention mechanism to map face image sequences to mel-scale spectrograms directly without requiring any human annotations. The proposed LipSound2 model is firstly pre-trained on $\sim$2400h multi-lingual (e.g. English and German) audio-visual data (VoxCeleb2). To verify the generalizability of the proposed method, we then fine-tune the pre-trained model on domain-specific datasets (GRID, TCD-TIMIT) for English speech reconstruction and achieve a significant improvement on speech quality and intelligibility compared to previous approaches in speaker-dependent and -independent settings. In addition to English, we conduct Chinese speech reconstruction on the CMLR dataset to verify the impact on transferability. Lastly, we train the cascaded lip reading (video-to-text) system by fine-tuning the generated audios on a pre-trained speech recognition system and achieve state-of-the-art performance on both English and Chinese benchmark datasets.
arXiv.org Artificial Intelligence
Dec-9-2021
- Country:
- North America > United States
- New York > Monroe County > Rochester (0.04)
- Europe
- United Kingdom > England
- Tyne and Wear > Sunderland (0.04)
- Italy > Calabria
- Catanzaro Province > Catanzaro (0.04)
- Germany
- Hamburg (0.04)
- Berlin (0.04)
- Hesse > Darmstadt Region
- Frankfurt (0.04)
- United Kingdom > England
- Asia > China
- North America > United States
- Genre:
- Research Report (0.50)
- Industry:
- Media (0.46)
- Health & Medicine (0.46)
- Information Technology (0.46)
- Technology: