CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

Jun-18-2024–arXiv.org Artificial Intelligence

This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.6 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer. Transformer-based models (Vaswani et al., 2017) have revolutionized speech processing, leading to significant advancements in automatic speech recognition and text-to-speech technologies (Latif et al., 2023; Prabhavalkar et al., 2023). This shift towards end-to-end systems has opened new avenues in Speech-to-Speech Translation (S2ST) for translating speech across languages.

bleu score, crossvoice, translation, (15 more...)

arXiv.org Artificial Intelligence

Jun-18-2024

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - Pennsylvania > Philadelphia County > Philadelphia (0.04)
- Asia > India
  - NCT > New Delhi (0.04)

Genre:
- Research Report (0.64)

Technology:
- Information Technology > Artificial Intelligence
  - Speech > Speech Recognition (1.00)
  - Natural Language > Machine Translation (1.00)
  - Machine Learning (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found