ViTs for SITS: Vision Transformers for Satellite Image Time Series
Tarasiou, Michail, Chavez, Erik, Zafeiriou, Stefanos
–arXiv.org Artificial Intelligence
In this paper we introduce the Temporo-Spatial Vision Transformer (TSViT), a fully-attentional model for general Satellite Image Time Series (SITS) processing based on the Vision Transformer (ViT). TSViT splits a SITS record into non-overlapping patches in space and time which are tokenized and subsequently processed by a factorized temporo-spatial encoder. We argue, that in contrast to natural images, a temporal-then-spatial factorization is more intuitive for SITS processing and present experimental evidence for this claim. Additionally, we enhance the model's discriminative power by introducing two novel mechanisms for acquisition-time-specific temporal positional encodings and multiple learnable class tokens. The effect of all novel design choices is evaluated through an extensive ablation study. Our proposed architecture achieves state-of-the-art performance, surpassing previous approaches by a significant margin in three publicly available SITS semantic segmentation and classification datasets. All model, training and evaluation codes are made publicly available to facilitate further research.
arXiv.org Artificial Intelligence
Apr-14-2023
- Country:
- Africa (0.04)
- North America > United States
- Kansas (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Europe
- France (0.04)
- Germany > Bavaria
- Upper Bavaria > Munich (0.04)
- Asia
- Uzbekistan (0.04)
- Central Asia (0.04)
- Genre:
- Research Report (0.64)
- Industry:
- Food & Agriculture > Agriculture (0.46)
- Government > Regional Government (0.46)
- Technology: