Country
VideoMAE: MaskedAutoencodersareData-Efficient LearnersforSelf-SupervisedVideoPre-Training
Transformer [70]has brought significant progress in natural language processing [17,7,54]. The vision transformer [20] also improves a series of computer vision tasks including image classification [66,88], object detection [8,37], semantic segmentation [80], object tracking [13,16], and video recognition [6,3].
NavigatingtheEffectofParametrization forDimensionalityReduction
Parametric dimensionality reduction methods have gained prominence for their ability togeneralize tounseen datasets, anadvantage that traditional approaches typically lack. Despite their growing popularity, there remains a prevalent misconception among practitioners about the equivalence in performance between parametric and non-parametric methods. Here, we showthat these methods are not equivalent - parametric methods retain global structure but lose significant localdetails.