选模型?看这个研究
本研究使用88个eGeMAPS特征,对六个分类群的生物声学嵌入进行线性与非线性回归探针,揭示模型编码的语音特征。结果显示没有单一模型能覆盖全部特征空间,拼接嵌入性能最佳。Loudness特征编码最好(R²=0.76),F0最难恢复(R²=0.33)。通过交叉引用可恢复性与特征显著性(NMI),为模型选择提供数据驱动指导。
Beyond task performance: Decoding bioacoustic embeddings with speech features
Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species or data-scarce domains. Here we reveal which speech-like features are encoded in bioacoustic representations. Using the 88~eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a ``no free lunch'' pattern: no single model captures the full feature space. A concatenated embedding achieves the highest performance, suggesting complementary acoustic space coverage across models. Loudness features are best encoded ($R^2 = 0.76$) while F0 is hardest to recover ($R^2 = 0.33$). By cross-referencing recoverability with per-species feature salience (NMI), we derive data-driven model selection guidance for bioacoustics.