LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
Naoki Koga,Yoshiaki Bando,Keisuke Imoto
TLDR
A LargE-scale Annotator’s labels for sound event Detection (LEAD) dataset is introduced, which is the dataset used to gain a better understanding of the variation in strong labels in sound event detection (SED).
Abstract
In this paper, we introduce a LargE-scale Annotator’s labels for sound event Detection (LEAD) dataset, which is the dataset used to gain a better understanding of the variation in strong labels in sound event detection (SED). In SED, it is very time-consuming to collect large-scale strong labels, and in most cases, multiple workers divide up the annotations to create a single dataset. In general, strong labels created by multiple annotators have large variations in the type of sound events and temporal onset/offset. If the strong labels of SED vary greatly depending on the annotator, the SED model trained on a dataset created by multiple annotators will be biased. Moreover, if annotators differ between training and evaluation data, there is a risk that the model cannot be evaluated correctly. To investigate the variation in strong labels and construct a robust method against the variation, we propose the LEAD dataset, which has 20 annotators for each sound clip. The LEAD dataset consists of strong labels assigned to sound clips from TUT Sound Events 2016/2017, TUT Acoustic Scenes 2016, and URBAN-SED. We also analyze variations in the strong labels in the LEAD dataset and provide insights into the variations.
