Reviews: SoundNet: Learning Sound Representations from Unlabeled Video
–Neural Information Processing Systems
This is a simple but interesting idea, and the results are promising. There are several points that could be improved: 1) What is the motivation for the particular architecture that was chosen? No other architecture is investigated, and this choice seems rather ad hoc. There exist many other neural networks that are being used in audio, whether for automatic speech recognition or for source separation, and I have yet to see such a network architecture. Why not simply use a DNN or LSTM on a frame-wise representation of the signal, e.g., STFT magnitude or log-magnitude?
Neural Information Processing Systems
Jan-20-2025, 14:05:19 GMT
- Technology: