SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation

Mu, Da, Zhang, Zhicheng, Yue, Haobo, Wang, Zehao, Tang, Jin, Yin, Jianqin

arXiv.org Artificial Intelligence 

Transformer-based models have demonstrated impressive capabilities. Utilizing State Space Models (SSMs), which establish However, the quadratic complexity of the Transformer's long-range context dependencies with linear computational self-attention mechanism results in computational complexity, is expected to overcome the aforementioned inefficiencies. In this paper, we propose a network architecture limitation. Recently, SSMs, exemplified by Mamba [8], for SELD called SELD-Mamba, which utilizes Mamba, a have demonstrated their effectiveness across various domains, selective state-space model. We adopt the Event-Independent including natural language processing [9], computer Network V2 (EINV2) as the foundational framework and replace vision [10, 11], and speech processing [12-14]. However, its Conformer blocks with bidirectional Mamba blocks the design of effective and efficient models using SSMs for to capture a broader range of contextual information while SELD has yet to be explored.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found