Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Huang, Ailin, Li, Bingxin, Wang, Bruce, Wu, Boyong, Yan, Chao, Feng, Chengli, Wang, Heng, Zhou, Hongyu, Wang, Hongyuan, Li, Jingbei, Sun, Jianjian, Wang, Joanna, Chen, Mingrui, Liu, Peng, Miao, Ruihang, Jiang, Shilei, Fei, Tian, You, Wang, Chen, Xi, Yang, Xuerui, Huang, Yechang, Zhang, Yuxiang, Ge, Zheng, Gong, Zheng, Huang, Zhewei, Zhang, Zixin, Wang, Bin, Li, Bo, Ma, Buyun, Miao, Changxin, Wan, Changyi, Xu, Chen, Shi, Dapeng, Hu, Dingyuan, Liu, Enle, Huang, Guanzhe, Yan, Gulin, Hu, Hanpeng, Jia, Haonan, Gong, Jiahao, Wu, Jiaoren, Wu, Jie, Yang, Jie, Lin, Junzhe, Li, Kaixiang, Xia, Lei, Gu, Longlong, Li, Ming, Hao, Nie, Ming, Ranchen, Pang, Shaoliang, Liu, Siqi, Yuan, Song, Cao, Tiancheng, Li, Wen, He, Wenqing, Zhao, Xu, Zhang, Xuelin, Yu, Yanbo, Zhong, Yinmin, Zhou, Yu, Liang, Yuanwei, Lu, Yuanwei, Yang, Yuxiang, Yang, Zidong, Zhang, Zili, Jiao, Binxing, Shum, Heung-Yeung, Chen, Jiansheng, Li, Jing, Zhang, Xiangyu, Zhang, Xinhao, Zhu, Yibo, Jiang, Daxin, Zhou, Shuchang, Hu, Chen
–arXiv.org Artificial Intelligence
Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate natural speech responses directly, hindering seamless audio interactions. To address this, we introduce Step-Audio-AQAA, a fully end-to-end LALM designed for Audio Query-Audio Answer (AQAA) tasks. The model integrates a dual-codebook audio tokenizer for linguistic and semantic feature extraction, a 130-billion-parameter backbone LLM and a neural vocoder for high-fidelity speech synthesis. Our post-training approach employs interleaved token-output of text and audio to enhance semantic coherence and combines Direct Preference Optimization (DPO) with model merge to improve performance. Evaluations on the StepEval-Audio-360 benchmark demonstrate that Step-Audio-AQAA excels especially in speech control, outperforming the state-of-art LALMs in key areas. This work contributes a promising solution for end-to-end LALMs and highlights the critical role of token-based vocoder in enhancing overall performance for AQAA tasks.
arXiv.org Artificial Intelligence
Jun-16-2025