
Compared with convolutional neural networks limited bytherelativelysmall receptive fields, the advantage of transformer for visual tasks is the capacity to perceivelong-range dependencies amongallimagepatches,whilethedeficiency is that the local fine-grained information is not fully excavated.