TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
Zheng, Haolong, Yegorova, Yekaterina, Hasegawa-Johnson, Mark
–arXiv.org Artificial Intelligence
ABSTRACT Speech foundation models have recently demonstrated the ability to perform Speech In-Context Learning (SICL). Selecting effective in-context examples is crucial for SICL performance, yet selection methodologies remain under-explored. In this work, we propose Text-Embedding KNN for SICL (TICL), a simple pipeline that uses semantic context to enhance off-the-shelf large multimodal models' speech recognition abilities without fine-tuning. Across challenging automatic speech recognition tasks, including accented English, multilingual speech, and children's speech, our method enable model to surpass zero-shot performance up to 84.7% relative WER reduction. Ablation studies are conducted to show the robustness and efficiency of our method.
arXiv.org Artificial Intelligence
Sep-18-2025