Understanding when spatial transformer networks do not support invariance, and what to do about it

Finnveden, Lukas, Jansson, Ylva, Lindeberg, Tony

arXiv.org Artificial Intelligence 

Abstract--Spatial transformer networks (STNs) were designed to enable convolutional neural networks (CNNs) to learn invariance to image transformations. STNs were originally proposed to transform CNN feature maps as well as input images. This enables the use of more complex features when predicting transformation parameters. However, since STNs perform a purely spatial transformation, they do not, in the general case, have the ability to align the feature maps of a transformed image with those of its original. STNs are therefore unable to support invariance when transforming CNN feature maps. Our first contribution is to present a simple proof that STNs Spatial transformer networks (STNs) [1] constitute a widely do not enable invariant recognition when transforming CNN used end-to-end trainable solution for CNNs to learn invariance feature maps. We do not claim mathematical novelty of this to image transformations. This makes it part of a fact, which is in some sense intuitive and can be inferred from growing body of work concerned with developing CNNs that more general results (see e.g. The key alternative proof directly applicable to STNs and accessible idea behind STNs is to introduce a trainable module - the with knowledge about basic analysis and some group theory. If such a module successfully learns to proposed explicitly or the question about the ability for invariance align images to a canonical pose, it can enable invariant is ignored for a range of different methods transforming recognition.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found