Meet Transformer in Transformer: A Visual Transformer That Captures Structural Information From Images

#artificialintelligence 

A new paper from Huawei, ISCAS and UCAS researchers proposes a novel Transformer-iN-Transformer (TNT) network architecture that outperforms conventional vision transformers on local information preservation and modelling for visual recognition. Transformer architectures were introduced in 2017, and their computational efficiency and scalability quickly made them the de-facto standard for natural language processing (NLP) tasks. Recently, transformers have also begun to show their potential in computer vision (CV) tasks such as image recognition, object detection, and image processing. Most of today's visual transformers view an input image as a sequence of image patches while ignoring intrinsic structural information among the patches -- a deficiency that negatively impacts their overall visual recognition ability. While convolutional neural networks (CNN) remain dominant in CV, transformer-based models have achieved promising performance on visual tasks without an image-specific inductive bias.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found