生态声景分析终于有了一个能处理真实噪声的可靠模型,做生态监测和声学研究的团队可以直接用它做预处理,省去大量人工标注时间。
生态声景由生物声、地声和人类声组成,但现有分析工具难以区分这些成分。CoarseSoundNet 是一个深度学习模型,能在真实被动声学监测条件下区分三类声音。研究发现,加入与目标域相似的 PAM 数据、引入静音类训练、使用类别阈值和时长约束能显著提升性能。案例验证表明,用 CoarseSoundNet 预过滤录音可获得与人工过滤相当的声学指数趋势,适合作为生态声学分析的预处理工具。
CoarseSoundNet: Building a reliable model for ecological soundscape analysis
A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research question in the field of soundscape ecology is how these components interact with each other, specifically how biophony responds to geophony and anthropophony. Nevertheless, as of today, there are not many analytical instruments that enable the distinct quantification of these elements. Recent machine learning (ML) approaches aim to support automated analysis but often rely on task-specific or clean data, limiting generalisation to noisy passive acoustic monitoring (PAM) recordings. This study presents a clear and reproducible structure to build ML models for coarse soundscape classification and introduces CoarseSoundNet, a deep learning model trained to distinguish biophony, geophony, and anthropophony under realistic PAM conditions. We systematically investigate model architectures, the influence of an additional training class, data composition, and evaluation strategies. Our findings suggest that model performance improves with additional PAM data, especially when similar to the target domain, and by introducing an explicit silence class during training. Class-specific decision thresholds and duration-based constraints further enhance performance, particularly for anthropophony and geophony. Error analyses exhibit challenges for anthropophony due to masking effects and confusions for silence and insect sounds for geophony and biophony. Finally, we conduct an ecological case study which shows that pre-filtering recordings with CoarseSoundNet yields acoustic index trends comparable to ground-truth filtering, supporting its use as an effective preprocessing tool for ecoacoustic analyses.