Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Huang, Ailin, Li, Bingxin, Wang, Bruce, Wu, Boyong, Yan, Chao, Feng, Chengli, Wang, Heng, Zhou, Hongyu, Wang, Hongyuan, Li, Jingbei, Sun, Jianjian, Wang, Joanna, Chen, Mingrui, Liu, Peng, Miao, Ruihang, Jiang, Shilei, Fei, Tian, You, Wang, Chen, Xi, Yang, Xuerui, Huang, Yechang, Zhang, Yuxiang, Ge, Zheng, Gong, Zheng, Huang, Zhewei, Zhang, Zixin, Wang, Bin, Li, Bo, Ma, Buyun, Miao, Changxin, Wan, Changyi, Xu, Chen, Shi, Dapeng, Hu, Dingyuan, Liu, Enle, Huang, Guanzhe, Yan, Gulin, Hu, Hanpeng, Jia, Haonan, Gong, Jiahao, Wu, Jiaoren, Wu, Jie, Yang, Jie, Lin, Junzhe, Li, Kaixiang, Xia, Lei, Gu, Longlong, Li, Ming, Hao, Nie, Ming, Ranchen, Pang, Shaoliang, Liu, Siqi, Yuan, Song, Cao, Tiancheng, Li, Wen, He, Wenqing, Zhao, Xu, Zhang, Xuelin, Yu, Yanbo, Zhong, Yinmin, Zhou, Yu, Liang, Yuanwei, Lu, Yuanwei, Yang, Yuxiang, Yang, Zidong, Zhang, Zili, Jiao, Binxing, Shum, Heung-Yeung, Chen, Jiansheng, Li, Jing, Zhang, Xiangyu, Zhang, Xinhao, Zhu, Yibo, Jiang, Daxin, Zhou, Shuchang, Hu, Chen

arXiv.org Artificial Intelligence 

Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate natural speech responses directly, hindering seamless audio interactions. To address this, we introduce Step-Audio-AQAA, a fully end-to-end LALM designed for Audio Query-Audio Answer (AQAA) tasks. The model integrates a dual-codebook audio tokenizer for linguistic and semantic feature extraction, a 130-billion-parameter backbone LLM and a neural vocoder for high-fidelity speech synthesis. Our post-training approach employs interleaved token-output of text and audio to enhance semantic coherence and combines Direct Preference Optimization (DPO) with model merge to improve performance. Evaluations on the StepEval-Audio-360 benchmark demonstrate that Step-Audio-AQAA excels especially in speech control, outperforming the state-of-art LALMs in key areas. This work contributes a promising solution for end-to-end LALMs and highlights the critical role of token-based vocoder in enhancing overall performance for AQAA tasks.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found