MulTTiPop: 流行音乐多轨转录数据集

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

精选理由

做音乐转录的看过来,MulTTiPop是新的多轨基准,覆盖老歌到现代,现有模型只拿38%得分,挑战很大。

AI 摘要

MulTTiPop数据集包含572段流行音乐,总计3.5小时音频,覆盖1930年代至2000年代的多种风格。数据通过元数据匹配从Lakh MIDI和TheoryTab数据集提取,并手动对齐节拍后调整MIDI时间。评估现有自动音乐转录模型,最佳模型仅达38% Onset F1,显示改进空间巨大。

原文 · arXiv cs.LG

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.