Spectrograms Are Sequences of Patches

Oct-28-2022–arXiv.org Artificial Intelligence

Self-supervised pre-training models have been used successfully in several machine learning domains. However, only a tiny amount of work is related to music. In our work, we treat a spectrogram of music as a series of patches and design a self-supervised model that captures the features of these sequential patches: Patchifier, which makes good use of self-supervised learning methods from both NLP and CV domains. We do not use labeled data for the pre-training process, only a subset of the MTAT dataset containing 16k music clips. After pre-training, we apply the model to several downstream tasks. Our model achieves a considerably acceptable result compared to other audio representation models. Meanwhile, our work demonstrates that it makes sense to consider audio as a series of patch segments.

artificial intelligence, inductive learning, machine learning, (16 more...)

arXiv.org Artificial Intelligence

Oct-28-2022

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - Virginia (0.04)
  - New York > New York County
    - New York City (0.04)

Genre:
- Research Report (0.50)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning
  - Inductive Learning (0.56)
  - Neural Networks (0.47)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found