KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI

Kuroki, So, Kubo, Yotaro, Akiba, Takuya, Tang, Yujin

arXiv.org Artificial Intelligence 

Their ability to continuously listen and respond to user queries opens new possibilities for human-machine interfaces in multiple applications, especially for those that require fast turnaround between humans and machines, such as question answering and brainstorming. Thanks to the recent developments of large transformers, these models can be implemented as direct speech-to-speech (S2S) models [3, 4, 5]. Because these S2S models are realized by a monolithic architecture that does not need to synchronize with other systems, their turnaround time is typically very low, contributing to more natural interaction. Moshi [2] is a pioneering model in that direction with an end-to-end S2S model for full-duplex conversational AI. However, because the inputs and outputs of S2S models are information-rich acoustic signals, auto-regressive modeling presents a fundamental challenge related to model capacity. In other words, unlike in text-only large language models (LLMs), S2S modeling must capture not only the verbal content of speech but also paralin-guistic features, such as speaking style, emotion, etc. This creates a fundamental inefficiency for S2S models with regard to knowledge acquisition. Given a similar model size, a text-only LLM can learn more knowledge because its capacity is dedicated solely to text, Figure 1.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found