Scinovex
article Open AccessTop 1% cited

Vision Transformers for Remote Sensing Image Classification

Remote Sensing · 2021 · Vol. 13(3) · pp. 516–516
Yakoub BaziLaila BashmalMohamad Mahmoud Al RahhalReham Al-DayilNaif Al Ajlan

Abstract

In this paper, we propose a remote-sensing scene-classification method based on vision transformers. These types of networks, which are now recognized as state-of-the-art models in natural language processing, do not rely on convolution layers as in standard convolutional neural networks (CNNs). Instead, they use multihead attention mechanisms as the main building block to derive long-range contextual relation between pixels in images. In a first step, the images under analysis are divided into patches, then converted to sequence by flattening and embedding. To keep information about the position, embedding position is added to these patches. Then, the resulting sequence is fed to several multihead attention layers for generating the final representation. At the classification stage, the first token sequence is fed to a softmax classification layer. To boost the classification performance, we explore several data augmentation strategies to generate additional data for training. Moreover, we show experimentally that we can compress the network by pruning half of the layers while keeping competing classification accuracies. Experimental results conducted on different remote-sensing image datasets demonstrate the promising capability of the model compared to state-of-the-art methods. Specifically, Vision Transformer obtains an average classification accuracy of 98.49%, 95.86%, 95.56% and 93.83% on Merced, AID, Optimal31 and NWPU datasets, respectively. While the compressed version obtained by removing half of the multihead attention layers yields 97.90%, 94.27%, 95.30% and 93.05%, respectively.

Remote-Sensing Image ClassificationAdvanced Image and Video Retrieval TechniquesRemote Sensing and Land UseComputer scienceSoftmax functionArtificial intelligencePattern recognition (psychology)EmbeddingTransformerConvolutional neural networkContextual image classificationBlock (permutation group theory)Encoder

Funding

  • King Saud University
Citations
564
FWCI
47.26
field-weighted impact
References
59
Percentile
100%
vs. same field & year
Citations per year
Cited by
Remote Sensing Image Change Detection With Transformers
IEEE Transactions on Geoscience and Remote Sensing · 2021 · 990 citations
References
Face Description with Local Binary Patterns: Application to Face Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2006 · 5,587 citations
AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
IEEE Transactions on Geoscience and Remote Sensing · 2017 · 2,039 citations
Convolutional Neural Networks for Large-Scale Remote-Sensing Image Classification
IEEE Transactions on Geoscience and Remote Sensing · 2016 · 1,088 citations
Remote Sensing Image Scene Classification: Benchmark and State of the Art
Proceedings of the IEEE · 2017 · 2,453 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.