From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

Sivasubramaniam, Nithyashree

arXiv.org Artificial Intelligence 

Instead of relying on acoustic signals, SSIs exploit non-vocal modalities that capture articulatory and physiological activity underlying speech production. A variety of input sources have been investigated, including surface electromyography (EMG), ultrasound tongue imaging (UTI), electromagnetic articulography (EMA), and visual cues from lip movements. These modalities offer complementary perspectives on articulatory dynamics, positioning SSIs as a promising solution for communication in scenarios where acoustic speech is impaired. Recent advances in SSI have introduced convolutionaltransducer models combined with transformer encoders and auxiliary phoneme prediction losses, achieving notable improvements in intelligibility. The input signals typically undergo feature extraction to capture articulatory or physiological attributes that preserve discriminative information necessary for speech generation.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found