What is the most important stuff in Vision Transformer?

#artificialintelligence 

This blog post describes the paper "Patches Are All You Need?" (Under review, 2021), which was submitted to ICLR2022 (under review as of the end of the Oct.). The ConvMixer proposed in this paper is composed of CNN BN, and unlike previous Vision Transformer series, it can achieve results even on small datasets such as CIFAR. We will then discuss whether patches are really the only important thing with the point of view that the model contains the global information and local information processing mechanisms. In this article, I will explain it according to the following items. The summary of this paper is as follows.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found