Masked Feature Modelling: Feature Masking for the Unsupervised Pre-training of a Graph Attention Network Block for Bottom-up Video Event Recognition

Daskalakis, Dimitrios, Gkalelis, Nikolaos, Mezaris, Vasileios

arXiv.org Artificial Intelligence 

Then, randomly selected image patches are masked and both masked and unmasked image patches are In this paper, we introduce Masked Feature Modelling fed into ViTs. The objective of the pretraining procedure (MFM), a novel approach for the unsupervised pre-training involves recovering the original visual tokens based on the of a Graph Attention Network (GAT) block. MFM utilizes masked image patches. This pretraining does not require a pretrained Visual Tokenizer to reconstruct masked features any ground-truth annotations for the employed image corpus. of objects within a video, leveraging the MiniKinetics Subsequently, the pretrained vision encoder can be dataset. We then incorporate the pre-trained GAT block into deployed and fine-tuned for various downstream tasks by a state-of-the-art bottom-up supervised video-event recognition appending lightweight task-specific layers. Although many architecture, ViGAT, to improve the model's starting studies in the image domain have experimented with this point and overall accuracy. Experimental evaluations on masking approach, masking has not yet been thoroughly the YLI-MED dataset demonstrate the effectiveness of MFM explored in the realm of video event recognition or videorelated in improving event recognition performance.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found