CNN Architectures for Large-Scale Audio Classification

Hershey, Shawn, Chaudhuri, Sourish, Ellis, Daniel P. W., Gemmeke, Jort F., Jansen, Aren, Moore, R. Channing, Plakal, Manoj, Platt, Devin, Saurous, Rif A., Seybold, Bryan, Slaney, Malcolm, Weiss, Ron J., Wilson, Kevin

arXiv.org Machine Learning 

ABSTRACT Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We investigate varying the size of both training set and label vocabulary, finding that analogs of the CNNs used in image classification do well on our audio classification task, and larger training and label sets help up to a point. A model using embeddings from these classifiers does much better than raw features on the Audio Set [5] Acoustic Event Detection (AED) classification task. Index Terms-- Acoustic Event Detection, Acoustic Scene Classification, Convolutional Neural Networks, Deep Neural Networks, Video Classification 1. INTRODUCTION Image classification performance has improved greatly with the advent of large datasets such as ImageNet [6] using Convolutional Neural Network (CNN) architectures such as AlexNet [1], VGG [2], Inception [3], and ResNet [4]. We are curious to see if similarly large datasets and CNNs can yield good performance on audio classification problems.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found