ReverBERT: A State Space Model for Efficient Text-Driven Speech Style Transfer
Brown, Michael, Martinez, Sofia, Singh, Priya
–arXiv.org Artificial Intelligence
Conventional speech style transfer approaches either rely on reference audio signals [8] or a few labeled style examples to guide the transformation [1]. Recently, text-driven speech style transfer is emerging as a compelling alternative, where stylistic cues and emotional descriptors (e.g., "whispering tone," "excited pitch," "authoritative" voice) come from natural language prompts [4,13]. While promising, text-driven speech style transfer must contend with several challenges: Semantic gap: The descriptive text for speech style can be abstract ( gentle, warm, comedic), complicating a direct alignment to acoustic features. Computational overhead: State-of-the-art generative models, including large Transformers or diffusion-based approaches, often require prohibitively long inference times. Expressive mismatch: Ensuring that the transferred style remains coherent and does not distort the linguistic content demands precise control over the generation process. In computer vision, the recently proposed Stylemamba [14] introduces a state space model (SSM) to perform efficient text-driven image style transfer. Inspired by how they leverage an SSM for local and global style consistency in images, we propose to adapt a state space viewpoint for speech, focusing on the sequence modeling aspect inherent in auditory signals.
arXiv.org Artificial Intelligence
Mar-26-2025
- Country:
- North America > United States
- Oregon (0.04)
- Asia > India
- West Bengal > Kharagpur (0.04)
- North America > United States
- Genre:
- Research Report (0.50)
- Technology: