这篇论文提出了TriA Pipeline,用2130小时音频自动生成431类标注数据,在家庭场景任务上准确率提升近4%,做音频分类研究可以看看。
TriA Pipeline是一个自动音频标注流水线,可将各场景音频高效转换为带事件标注的高质量训练数据。利用该流水线构建的TriA数据集包含超过2130小时音频,覆盖431个音频类别。从TriA中划分出先验知识引导子集TriA_GK,并在三个家庭音频分类任务上进行了对比实验。在手动标注数据基础上加入TriA_GK后,平均准确率提升3.97%,Macro-F1提升3.35%,验证了流水线的有效性。
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$utomatic $\textbf{A}$udio $\textbf{A}$nnotation Pipeline--TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriA$_{\mathrm{GK}}$) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA$_{\mathrm{GK}}$, TriA$_{\mathrm{GK}}$ could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA$_{\mathrm{GK}}$ and the TriA Pipeline.