EquiSELD:基于O(3)等变性的声音事件定位与检测网络
EquiSELD: Efficient training of equivariant sound event localization and detection networks
EquiSELD 借 O(3) 对称性做声音事件定位,训练成本更低,成绩还胜过以前的等变网络。
论文提出 EquiSELD,一种处理一阶 Ambisonics(FOA)信号的等变注意力网络。该网络把 FOA 信号拆成 O(3) 不变标量和等变强度向量两条流,用 Multi-ACCDOA 读出层输出不变的事件活动强度和等变的到达方向(DOA)。不同于旋转数据增强只学到近似对称性、此前工作只覆盖 SO(3) 等变的做法,EquiSELD 完整利用了 FOA 的 O(3) 对称性。在含实测房间冲激响应(RIR)的仿真场景和真实录音场景上,EquiSELD 都超过了先前的等变网络,训练成本只是一小部分。它还在仿真真实声景上超过同等规模的非等变 SELD 网络,并在真实声景中取得有竞争力的成绩。
EquiSELD: Efficient training of equivariant sound event localization and detection networks
First-order Ambisonics (FOA) signals exhibit exact O(3) symmetry: The rotation or reflection of the FOA signal modifies the direction of arrival of the sound sources, while preserving the sound sources themselves. Prior attempts to utilize this spatial symmetry of FOA to improve the efficiency and robustness of sound event detection and localization (SELD) systems either learned only an approximation of the symmetry through rotation-based augmentation or relied on computationally expensive methods to integrate equivariance. Furthermore, prior work focused exclusively on SO(3) equivariance , leaving the potential of incorporating O(3) equivariance for SELD tasks unclear. To address these limitations, we developed EquiSELD. This equivariant attention network processes first-order Ambisonics as paired streams of O(3)-invariant scalars and equivariant intensity vectors, producing an invariant activity magnitude and an equivariant DOA with a Multi-ACCDOA readout. To compare the impact of O(3) versus SO(3)-equivariance, we designed a matched SO(3)-only variant. EquiSELD outperforms prior equivariant networks on both simulated scenes with measured RIRs and recordings of real-world sound scenes at a fraction of the training cost. EquiSELD additionally surpasses the performance of non-equivariant SELD networks of a similar size on the simulated real-world sound scenes and achieves competitive performance on the real-world sound scenes.