作者
Yong Xu, Qiuqiang Kong, Wenwu Wang, Mark D Plumbley
发表日期
2018/4/15
研讨会论文
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
页码范围
121-125
出版商
IEEE
简介
In this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 challenge. The audio clips in this task, which are extracted from YouTube videos, are manually labelled with one or more audio tags, but without time stamps of the audio events, hence referred to as weakly labelled data. Two subtasks are defined in this challenge including audio tagging and sound event detection using this weakly labelled data. We propose a convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) non-linearity applied on the log Mel spectrogram. In addition, we propose a temporal attention method along the frames to predict the locations of each audio event in a …
引用总数
2017201820192020202120222023202433958484842215
学术搜索中的文章
Y Xu, Q Kong, W Wang, MD Plumbley - 2018 IEEE international conference on acoustics …, 2018