Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
Wang, Qiongqiong, Sailor, Hardik B., Wong, Jeremy H. M., Liu, Tianchi, Sun, Shuo, Zhang, Wenyu, Huzaifah, Muhammad, Chen, Nancy, Aw, Ai Ti
–arXiv.org Artificial Intelligence
--Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralin-guistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. In recent years, large language models (LLMs) have shown remarkable capabilities across a wide range of natural language processing tasks. Building on this success, large speech language models (Speech-LLMs), which extend LLMs with speech inputs, have emerged as a promising direction to enable spoken dialog systems, voice-based assistants, and human-computer interaction [1]-[4]. While these models excel at content-related tasks like speech recognition, these models often exhibit limitations in tasks requiring empathetic reasoning or emotional understanding. Past efforts to improve paralinguistic understanding for LLM can be grouped into: (1) fine-tuning on labeled emotional data [5]-[8], (2) knowledge distillation from paralinguistic teachers [8]-[10], and (3) translating emotional signals into language prompts [11]-[14].
arXiv.org Artificial Intelligence
Aug-12-2025