Self-Supervised Visual Representation Learning: Contrastive, Siamese, Clustering, and Self-Distillation Approaches
Keywords:
computer vision, representation learning, BYOL, MoCoAbstract
The dependence of supervised deep learning on large quantities of manually annotated data represents an important constraint in computer vision. Self-supervised learning addresses this limitation by constructing supervisory signals directly from unlabeled observations and learning transferable representations before downstream fine-tuning. Between 2018 and 2021, self-supervised visual learning progressed from handcrafted pretext tasks toward contrastive learning, momentum-based encoders, clustering, Siamese architectures, redundancy reduction, and self-distillation. This review examines the principal approaches available through 2021. Contrastive predictive coding and instance-discrimination techniques established the foundations for learning representations through similarity relationships. Momentum Contrast introduced dynamic dictionaries and momentum encoders, while SimCLR demonstrated the importance of data augmentation, projection heads, and large training batches. Bootstrap Your Own Latent showed that competitive representations could be obtained without explicit negative samples. SwAV combined clustering with swapped prediction between augmented views, whereas SimSiam demonstrated the importance of stop-gradient operations within comparatively simple Siamese networks. Barlow Twins approached representation learning through redundancy reduction, while DINO demonstrated strong emergent properties when self-distillation was combined with Vision Transformers. The review compares these approaches in terms of negative samples, architectural asymmetry, augmentation strategy, memory requirements, representation collapse, and transferability. By the end of 2021, self-supervised learning had become a major alternative to conventional supervised pretraining and provided an increasingly practical route for exploiting large unlabeled image collections.
References
Bengio Y, Courville A, Vincent P. Representation learning: a review and new perspectives. IEEE Trans Pattern Anal Mach Intell. 2013;35(8):1798-1828.
Vincent P, Larochelle H, Bengio Y, Manzagol PA. Extracting and composing robust features with denoising autoencoders. In: Proceedings of the 25th International Conference on Machine Learning. New York: ACM; 2008. p. 1096-1103.
Doersch C, Gupta A, Efros AA. Unsupervised visual representation learning by context prediction. In: Proceedings of the IEEE International Conference on Computer Vision. 2015. p. 1422-1430.
Noroozi M, Favaro P. Unsupervised learning of visual representations by solving jigsaw puzzles. In: Leibe B, Matas J, Sebe N, Welling M, editors. Computer Vision—ECCV 2016. Cham: Springer; 2016. p. 69-84.
Gidaris S, Singh P, Komodakis N. Unsupervised representation learning by predicting image rotations. In: International Conference on Learning Representations. 2018.
Oord A, Li Y, Vinyals O. Representation learning with contrastive predictive coding. arXiv. 2018;1807.03748.
Wu Z, Xiong Y, Yu SX, Lin D. Unsupervised feature learning via non-parametric instance discrimination. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 3733-3742.
He K, Fan H, Wu Y, Xie S, Girshick R. Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. p. 9729-9738.
Chen T, Kornblith S, Norouzi M, Hinton G. A simple framework for contrastive learning of visual representations. Proc Mach Learn Res. 2020;119:1597-1607.
Grill JB, Strub F, Altché F, Tallec C, Richemond P, Buchatskaya E, et al. Bootstrap your own latent: a new approach to self-supervised learning. Adv Neural Inf Process Syst. 2020;33:21271-21284.
Caron M, Misra I, Mairal J, Goyal P, Bojanowski P, Joulin A. Unsupervised learning of visual features by contrasting cluster assignments. Adv Neural Inf Process Syst. 2020;33:9912-9924.
Chen X, He K. Exploring simple Siamese representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021. p. 15750-15758.
Zbontar J, Jing L, Misra I, LeCun Y, Deny S. Barlow Twins: self-supervised learning via redundancy reduction. Proc Mach Learn Res. 2021;139:12310-12320.
Caron M, Touvron H, Misra I, Jégou H, Mairal J, Bojanowski P, Joulin A. Emerging properties in self-supervised Vision Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 9650-9660.
Chen X, Fan H, Girshick R, He K. Improved baselines with momentum contrastive learning. arXiv. 2020;2003.04297.
Tian Y, Krishnan D, Isola P. Contrastive multiview coding. In: Vedaldi A, Bischof H, Brox T, Frahm JM, editors. Computer Vision—ECCV 2020. Cham: Springer; 2020. p. 776-794.
Bachman P, Hjelm RD, Buchwalter W. Learning representations by maximizing mutual information across views. Adv Neural Inf Process Syst. 2019;32:15535-15545.
Misra I, van der Maaten L. Self-supervised learning of pretext-invariant representations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. p. 6707-6717.
Henaff O. Data-efficient image recognition with contrastive predictive coding. Proc Mach Learn Res. 2020;119:4182-4192.
Ericsson L, Gouk H, Hospedales TM. How well do self-supervised models transfer? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021. p. 5414-5423.