Multimodal Reasoning and Visual Agents: Native Vision-Language Pretraining, Document Intelligence, Long-Video Understanding, and Computer Interaction
Keywords:
Qwen2.5-VL, visual agents, Multimodal large language modelsAbstract
Multimodal large language models evolved rapidly from systems that primarily described individual images into increasingly general visual reasoning models capable of processing documents, diagrams, charts, long videos, multiple images, spatial information, and graphical user interfaces. This review examines the progression toward multimodal reasoning and visual agents through 2025. CLIP established scalable image-text alignment, while Flamingo demonstrated few-shot multimodal learning through the integration of pretrained visual and language components. BLIP-2 reduced multimodal training requirements by introducing a compact bridge between frozen image encoders and language models. LLaVA and InstructBLIP subsequently demonstrated the importance of multimodal instruction tuning. Models such as Qwen-VL and InternVL extended capabilities toward localization, multilingual understanding, and higher-resolution visual processing. By 2024, Qwen2-VL, LLaVA-OneVision, PaliGemma, and Molmo expanded dynamic-resolution, multi-image, and video capabilities. The 2025 generation further emphasized native multimodal pretraining and interaction with external environments. Qwen2.5-VL combined dynamic visual resolution, document parsing, object localization, temporal encoding, long-video comprehension, and visual-agent capabilities. InternVL3 introduced native multimodal pretraining rather than adapting only a pretrained text model, while Phi-4-Multimodal used modality-specific adapters to integrate vision and speech with a compact language backbone. Gemma 3 extended lightweight open models with image understanding, multilinguality, and long context. This review compares visual tokenization, dynamic resolution, connector architectures, multimodal pretraining, instruction tuning, temporal encoding, document reasoning, grounding, and interface interaction. Key limitations include hallucinated visual details, weak fine-grained spatial reasoning, OCR errors, temporal inconsistency, interface fragility, and limited evaluation of real-world agentic reliability.
References
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. Proc Mach Learn Res. 2021;139:8748-8763.
Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst. 2022;35:23716-23736.
Li J, Li D, Savarese S, Hoi S. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. Proc Mach Learn Res. 2023;202:19730-19742.
Liu H, Li C, Wu Q, Lee YJ. Visual instruction tuning. Adv Neural Inf Process Syst. 2023;36.
Dai W, Li J, Li D, Tiong AMH, Zhao J, Wang W, et al. InstructBLIP: towards general-purpose vision-language models with instruction tuning. Adv Neural Inf Process Syst. 2023;36.
Bai J, Bai S, Yang S, Wang S, Tan S, Wang P, et al. Qwen-VL: a versatile vision-language model for understanding, localization, text reading, and beyond. arXiv. 2023;2308.12966.
Chen X, Djolonga J, Padlewski P, Mustafa B, Changpinyo S, Wu J, et al. PaLI-X: on scaling up a multilingual vision and language model. arXiv. 2023;2305.18565.
Chen Z, Wu J, Wang W, Su W, Chen G, Xing S, et al. InternVL: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.
Wang P, Bai S, Tan S, Wang S, Fan Z, Bai J, et al. Qwen2-VL: enhancing vision-language model's perception of the world at any resolution. arXiv. 2024;2409.12191.
Li B, Zhang Y, Guo D, Zhang R, Li F, Zhang H, et al. LLaVA-OneVision: easy visual task transfer. arXiv. 2024;2408.03326.
Deitke M, Clark C, Lee S, Tripathi R, Yang Y, Park JW, et al. Molmo and PixMo: open weights and open data for state-of-the-art multimodal models. arXiv. 2024;2409.17146.
Beyer L, Steiner A, Pinto AS, Kolesnikov A, Wang X, Salz D, et al. PaliGemma: a versatile 3B VLM for transfer. arXiv. 2024;2407.07726.
Bai S, Chen K, Liu X, Wang J, Ge W, Song S, et al. Qwen2.5-VL technical report. arXiv. 2025;2502.13923.
Zhu J, Wang W, Chen Z, Liu Z, Ye S, Gu L, et al. InternVL3: exploring advanced training and test-time recipes for open-source multimodal models. arXiv. 2025;2504.10479.
Abouelenin A, Ashfaq A, Atkinson A, Awadalla H, Bach N, Bao J, et al. Phi-4-Mini technical report: compact yet powerful multimodal language models via mixture-of-LoRAs. arXiv. 2025;2503.01743.
Gemma Team. Gemma 3 technical report. arXiv. 2025;2503.19786.
Yue X, Ni Y, Zhang K, Zheng T, Liu R, Zhang G, et al. MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.
Lu P, Bansal H, Xia T, Liu J, Li C, Hajishirzi H, et al. MathVista: evaluating mathematical reasoning of foundation models in visual contexts. In: International Conference on Learning Representations. 2024.
Yu W, Yang Z, Li L, Wang J, Lin K, Liu Z, et al. MM-Vet: evaluating large multimodal models for integrated capabilities. Proc Mach Learn Res. 2024.
Li B, Wang R, Wang G, Ge Y, Ge Y, Shan Y. SEED-Bench: benchmarking multimodal large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.