VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
Damianos, Dimitrios, Voukoutis, Leon, Paraskevopoulos, Georgios, Katsouros, Vassilis
–arXiv.org Artificial Intelligence
ABSTRACT We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim of building speech-enabled LLMs. Instead of directly using audio embeddings, we explore an intermediate audio-conditioned text space as a more effective mechanism for alignment. Our method operates fully in continuous text representation spaces, fusing Whisper's hidden decoder states with those of an LLM through cross-modal attention, and supports both offline and streaming modes. We introduce V oxKrikri, the first Greek speech LLM, and show through analysis that our approach effectively aligns representations across modalities. Index T erms-- Speech LLMs, modality fusion, continuous latent space, causal masking, ASR 1. INTRODUCTION Large language models (LLMs) have achieved remarkable success in natural language processing, inspiring the development of multimodal LLMs that combine pre-trained language models with modality-specific encoders such as those for audio and images [1, 2, 3, 4, 5].
arXiv.org Artificial Intelligence
Sep-22-2025