Vision Transformers for Image Recognition and Dense Prediction: From Patch-Based Attention to Hierarchical Architectures
Keywords:
deep learning, Vision Transformer, image recognitionAbstract
Convolutional neural networks dominated computer vision throughout much of the previous decade because their locality, translation equivariance, and hierarchical feature extraction provided effective inductive biases for image processing. By 2021, however, Transformer architectures originally developed for sequence modeling had become increasingly competitive alternatives for visual recognition. This review examines the development of Vision Transformers through 2021, beginning with the adaptation of self-attention to image patches and progressing toward more computationally efficient and hierarchical architectures. The Vision Transformer demonstrated that images could be represented as sequences of fixed-size patches and processed largely through standard Transformer encoders. Subsequent methods addressed the substantial data and computational requirements of the original architecture. Data-efficient image Transformers introduced attention-based knowledge distillation, while Tokens-to-Token ViT incorporated progressive tokenization to better represent local image structure. Pyramid Vision Transformer and Swin Transformer introduced hierarchical feature representations suitable for object detection and semantic segmentation. Convolutional Vision Transformer and LeViT further explored hybridization between convolutional inductive biases and attention mechanisms. This review compares these approaches in terms of token generation, self-attention strategy, computational complexity, hierarchical representation, data efficiency, and suitability for dense prediction. Major limitations involving quadratic attention complexity, training-data requirements, positional representations, and deployment efficiency are discussed. The evidence available by 2021 indicated that Transformers were becoming viable general-purpose computer-vision backbones rather than remaining exclusively language-processing architectures.
References
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, et al. An image is worth 16×16 words: transformers for image recognition at scale. In: International Conference on Learning Representations. 2021.
Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H. Training data-efficient image transformers & distillation through attention. Proc Mach Learn Res. 2021;139:10347-10357.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, et al. Swin Transformer: hierarchical vision Transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 10012-10022.
Touvron H, Cord M, Sablayrolles A, Synnaeve G, Jégou H. Going deeper with image Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 32-42.
Yuan L, Chen Y, Wang T, Yu W, Shi Y, Jiang ZH, et al. Tokens-to-Token ViT: training Vision Transformers from scratch on ImageNet. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 558-567.
Wang W, Xie E, Li X, Fan DP, Song K, Liang D, et al. Pyramid Vision Transformer: a versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 568-578.
Wu H, Xiao B, Codella N, Liu M, Dai X, Yuan L, Zhang L. CvT: introducing convolutions to Vision Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 22-31.
Graham B, El-Nouby A, Touvron H, Stock P, Joulin A, Jégou H, et al. LeViT: a Vision Transformer in ConvNet's clothing for faster inference. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 12259-12269.
Han K, Xiao A, Wu E, Guo J, Xu C, Wang Y. Transformer in Transformer. Adv Neural Inf Process Syst. 2021;34:15908-15919.
Tolstikhin IO, Houlsby N, Kolesnikov A, Beyer L, Zhai X, Unterthiner T, et al. MLP-Mixer: an all-MLP architecture for vision. Adv Neural Inf Process Syst. 2021;34:24261-24272.
Bello I, Zoph B, Le QV, Vaswani A, Shlens J. Attention augmented convolutional networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. p. 3286-3295.
Ramachandran P, Parmar N, Vaswani A, Bello I, Levskaya A, Shlens J. Stand-alone self-attention in vision models. Adv Neural Inf Process Syst. 2019;32:68-80.
Hu J, Shen L, Sun G. Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 7132-7141.
He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. p. 770-778.
Tan M, Le QV. EfficientNet: rethinking model scaling for convolutional neural networks. Proc Mach Learn Res. 2019;97:6105-6114.
Howard A, Sandler M, Chu G, Chen LC, Chen B, Tan M, et al. Searching for MobileNetV3. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. p. 1314-1324.
Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S. End-to-end object detection with Transformers. In: Vedaldi A, Bischof H, Brox T, Frahm JM, editors. Computer Vision—ECCV 2020. Cham: Springer; 2020. p. 213-229.
Cordonnier JB, Loukas A, Jaggi M. On the relationship between self-attention and convolutional layers. In: International Conference on Learning Representations. 2020.
Parmar N, Vaswani A, Uszkoreit J, Kaiser L, Shazeer N, Ku A, Tran D. Image Transformer. Proc Mach Learn Res. 2018;80:4055-4064.