Vision-Language Pretraining for Multimodal Artificial Intelligence: Contrastive Alignment, Cross-Modal Fusion, and Generative Learning

Authors

  • Sophie Laurent Author

Keywords:

multimodal Transformer, contrastive learning, image-text retrieval, visual question answering, Flamingo

Abstract

The integration of visual and linguistic information became a major direction in artificial intelligence by 2022. Earlier image-captioning and visual-semantic embedding systems typically relied on task-specific architectures and comparatively limited annotated datasets. Vision-language pretraining shifted this paradigm toward large-scale representation learning from paired image-text corpora, enabling models to transfer across image retrieval, visual question answering, image captioning, classification, grounding, and multimodal generation. This review examines major developments in vision-language representation learning through 2022. Early visual-semantic alignment and image-captioning approaches are reviewed as foundations for subsequent multimodal Transformers. Architectures such as ViLBERT, LXMERT, VisualBERT, UNITER, OSCAR, and ViLT introduced different strategies for fusing image-region representations and linguistic tokens. Contrastive Language-Image Pretraining demonstrated that large-scale image-text contrastive learning could produce transferable visual representations and strong zero-shot image classification. ALIGN extended contrastive pretraining to noisier web-scale image-text datasets, while ALBEF combined contrastive alignment with multimodal fusion. By 2022, BLIP introduced bootstrapped data filtering and unified understanding and generation, while Flamingo demonstrated few-shot multimodal learning through pretrained vision and language components. The review compares dual-encoder and fusion architectures, contrastive and generative objectives, region-based and patch-based visual inputs, zero-shot transfer, dataset scale, and computational requirements. Persistent challenges included noisy web supervision, multimodal hallucination, evaluation reliability, representational bias, fine-grained grounding, and the cost of increasingly large multimodal models.

References

Frome A, Corrado GS, Shlens J, Bengio S, Dean J, Ranzato MA, Mikolov T. DeViSE: a deep visual-semantic embedding model. Adv Neural Inf Process Syst. 2013;26:2121-2129.

Karpathy A, Fei-Fei L. Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015. p. 3128-3137.

Vinyals O, Toshev A, Bengio S, Erhan D. Show and tell: a neural image caption generator. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015. p. 3156-3164.

Anderson P, He X, Buehler C, Teney D, Johnson M, Gould S, Zhang L. Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 6077-6086.

Lu J, Batra D, Parikh D, Lee S. ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Adv Neural Inf Process Syst. 2019;32:13-23.

Tan H, Bansal M. LXMERT: learning cross-modality encoder representations from Transformers. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. 2019. p. 5100-5111.

Li LH, Yatskar M, Yin D, Hsieh CJ, Chang KW. VisualBERT: a simple and performant baseline for vision and language. arXiv. 2019;1908.03557.

Chen YC, Li L, Yu L, El Kholy A, Ahmed F, Gan Z, et al. UNITER: universal image-text representation learning. In: Vedaldi A, Bischof H, Brox T, Frahm JM, editors. Computer Vision—ECCV 2020. Cham: Springer; 2020. p. 104-120.

Li X, Yin X, Li C, Zhang P, Hu X, Zhang L, et al. OSCAR: object-semantics aligned pre-training for vision-language tasks. In: Vedaldi A, Bischof H, Brox T, Frahm JM, editors. Computer Vision—ECCV 2020. Cham: Springer; 2020. p. 121-137.

Kim W, Son B, Kim I. ViLT: Vision-and-Language Transformer without convolution or region supervision. Proc Mach Learn Res. 2021;139:5583-5594.

Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. Proc Mach Learn Res. 2021;139:8748-8763.

Jia C, Yang Y, Xia Y, Chen YT, Parekh Z, Pham H, et al. Scaling up visual and vision-language representation learning with noisy text supervision. Proc Mach Learn Res. 2021;139:4904-4916.

Li J, Selvaraju R, Gotmare A, Joty S, Xiong C, Hoi SCH. Align before fuse: vision and language representation learning with momentum distillation. Adv Neural Inf Process Syst. 2021;34:9694-9705.

Li J, Li D, Xiong C, Hoi S. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. Proc Mach Learn Res. 2022;162:12888-12900.

Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst. 2022;35:23716-23736.

Wang Z, Yu J, Yu AW, Dai Z, Tsvetkov Y, Cao Y. SimVLM: simple visual language model pretraining with weak supervision. In: International Conference on Learning Representations. 2022.

Zhai X, Wang X, Mustalam M, Steiner A, Gritsenko A, Chiu CY, et al. LiT: zero-shot transfer with locked-image text tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. p. 18123-18133.

Singh A, Hu R, Goswami V, Couairon G, Galuba W, Rohrbach M, Kiela D. FLAVA: a foundational language and vision alignment model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. p. 15638-15650.

Yu J, Wang Z, Vasudevan V, Yeung L, Seyedhosseini M, Wu Y. CoCa: contrastive captioners are image-text foundation models. arXiv. 2022;2205.01917.

Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, et al. Zero-shot text-to-image generation. Proc Mach Learn Res. 2021;139:8821-8831.

Published

2022-06-01