Multimodal Large Language Models for General-Purpose Vision-Language Intelligence: Architectures, Instruction Tuning, and Emergent Capabilities

Authors

  • Yuna Takahashi Author

Keywords:

BLIP-2, LLaVA, vision-language model, Multimodal large language models

Abstract

The convergence of large language modeling and vision-language pretraining led to rapid development of multimodal large language models during 2023. Earlier vision-language systems were typically optimized for predefined tasks such as captioning, retrieval, visual question answering, or recognition. Multimodal foundation models instead sought to connect pretrained visual encoders with large language models capable of instruction following and open-ended generation. This review examines the architectural evolution and principal capabilities of multimodal large language models available through 2023. Contrastive Language-Image Pretraining established transferable image-text alignment at scale, while Flamingo demonstrated few-shot multimodal learning by connecting frozen vision and language components. BLIP and BLIP-2 advanced bootstrapped vision-language pretraining, with BLIP-2 employing a lightweight Querying Transformer to bridge frozen image encoders and frozen large language models. LLaVA demonstrated visual instruction tuning using language-model-generated multimodal instruction data, while InstructBLIP systematically extended instruction tuning across diverse vision-language tasks. Concurrent systems including MiniGPT-4, Kosmos-1, PaLI, PaLI-X, and ImageBind explored different combinations of multimodal representation, language generation, and cross-modal transfer. The review analyzes visual encoders, connector modules, cross-attention, instruction tuning, frozen versus jointly trained components, multimodal benchmarks, and computational efficiency. Challenges identified in 2023 included visual hallucination, insufficient spatial reasoning, OCR limitations, weak fine-grained grounding, dataset contamination, safety, evaluation inconsistencies, and the difficulty of distinguishing genuine multimodal reasoning from language-prior exploitation.

References

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.

Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. Proc Mach Learn Res. 2021;139:8748-8763.

Jia C, Yang Y, Xia Y, Chen YT, Parekh Z, Pham H, et al. Scaling up visual and vision-language representation learning with noisy text supervision. Proc Mach Learn Res. 2021;139:4904-4916.

Li J, Selvaraju R, Gotmare A, Joty S, Xiong C, Hoi SCH. Align before fuse: vision and language representation learning with momentum distillation. Adv Neural Inf Process Syst. 2021;34:9694-9705.

Li J, Li D, Xiong C, Hoi S. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. Proc Mach Learn Res. 2022;162:12888-12900.

Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst. 2022;35:23716-23736.

Wang Z, Yu J, Yu AW, Dai Z, Tsvetkov Y, Cao Y. SimVLM: simple visual language model pretraining with weak supervision. In: International Conference on Learning Representations. 2022.

Chen T, Saxena S, Li L, Fleet DJ, Hinton G. Pix2seq: a language modeling framework for object detection. In: International Conference on Learning Representations. 2022.

Chen X, Wang X, Changpinyo S, Piergiovanni AJ, Padlewski P, Salz D, et al. PaLI: a jointly-scaled multilingual language-image model. In: International Conference on Learning Representations. 2023.

Li J, Li D, Savarese S, Hoi S. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. Proc Mach Learn Res. 2023;202:19730-19742.

Liu H, Li C, Wu Q, Lee YJ. Visual instruction tuning. Adv Neural Inf Process Syst. 2023;36.

Dai W, Li J, Li D, Tiong AMH, Zhao J, Wang W, et al. InstructBLIP: towards general-purpose vision-language models with instruction tuning. Adv Neural Inf Process Syst. 2023;36.

Zhu D, Chen J, Shen X, Li X, Elhoseiny M. MiniGPT-4: enhancing vision-language understanding with advanced large language models. arXiv. 2023;2304.10592.

Huang S, Dong L, Wang W, Hao Y, Singhal S, Ma S, et al. Language is not all you need: aligning perception with language models. arXiv. 2023;2302.14045.

Chen X, Djolonga J, Padlewski P, Mustafa B, Changpinyo S, Wu J, et al. PaLI-X: on scaling up a multilingual vision and language model. arXiv. 2023;2305.18565.

Girdhar R, El-Nouby A, Liu Z, Singh M, Alwala KV, Joulin A, Misra I. ImageBind: one embedding space to bind them all. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 15180-15190.

Yang Z, Li L, Lin K, Wang J, Lin CC, Liu Z, Wang L. The dawn of LMMs: preliminary explorations with GPT-4V(ision). arXiv. 2023;2309.17421.

Li B, Zhang Y, Chen L, Wang J, Yang J, Liu Z. Otter: a multi-modal model with in-context instruction tuning. arXiv. 2023;2305.03726.

Ye Q, Xu H, Ye J, Yan M, Hu A, Liu H, et al. mPLUG-Owl: modularization empowers large language models with multimodality. arXiv. 2023;2304.14178.

Bai J, Bai S, Yang S, Wang S, Tan S, Wang P, et al. Qwen-VL: a versatile vision-language model for understanding, localization, text reading, and beyond. arXiv. 2023;2308.12966.

Published

2023-06-01