Video-Capable Multimodal Large Language Models: Temporal Representation, Long-Video Memory, and Visual Instruction Tuning
Keywords:
temporal reasoning, long video, visual instruction tuning, MovieChat, TimeChatAbstract
Multimodal large language models initially focused primarily on static image-language interaction, but by 2024 increasing research attention had shifted toward video understanding. Video introduces temporal order, motion, event duration, repeated visual content, and substantially larger token requirements than individual images. This review examines the evolution of video-capable multimodal language models and long-video reasoning through 2024. Earlier approaches such as VideoBERT and contrastive video-language representation learning established foundations for joint visual-linguistic modeling. Large vision-language architectures including CLIP, Flamingo, BLIP-2, and LLaVA subsequently provided transferable visual-language components that could be adapted to video. Video-LLaMA extended language-model interaction to audio-visual inputs, while Video-ChatGPT introduced large-scale video instruction data and open-ended video conversation. Video-LLaVA unified image and video representation before projection into the language space. Chat-UniVi employed dynamic unified visual tokens, and LLaMA-VID compressed individual frames into context and content tokens to enable much longer videos. TimeChat introduced timestamp-aware features and temporal localization, whereas MovieChat used sparse memory mechanisms to support long-video understanding. MVBench provided a broad evaluation framework for temporal multimodal reasoning. The review compares frame sampling, visual token compression, temporal modeling, instruction tuning, long-video memory, visual-language projection, and multimodal evaluation. Limitations include temporal hallucination, inadequate event ordering, excessive visual-token counts, weak localization, benchmark contamination, and difficulty maintaining detailed representations across long videos.
References
Sun C, Myers A, Vondrick C, Murphy K, Schmid C. VideoBERT: a joint model for video and language representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. p. 7464-7473.
Miech A, Alayrac JB, Smaira L, Laptev I, Sivic J, Zisserman A. End-to-end learning of visual representations from uncurated instructional videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. p. 9879-9889.
Bain M, Nagrani A, Varol G, Zisserman A. Frozen in time: a joint video and image encoder for end-to-end retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 1728-1738.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. Proc Mach Learn Res. 2021;139:8748-8763.
Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst. 2022;35:23716-23736.
Li J, Li D, Savarese S, Hoi S. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. Proc Mach Learn Res. 2023;202:19730-19742.
Liu H, Li C, Wu Q, Lee YJ. Visual instruction tuning. Adv Neural Inf Process Syst. 2023;36.
Zhang H, Li X, Bing L. Video-LLaMA: an instruction-tuned audio-visual language model for video understanding. In: Proceedings of EMNLP 2023: System Demonstrations. 2023.
Maaz M, Rasheed H, Khan S, Khan FS. Video-ChatGPT: towards detailed video understanding via large vision and language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. p. 12585-12602.
Lin B, Ye Y, Zhu B, Cui J, Ning M, Jin P, Yuan L. Video-LLaVA: learning united visual representation by alignment before projection. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. p. 5971-5984.
Jin P, Takanobu R, Zhang W, Cao X, Yuan L. Chat-UniVi: unified visual representation empowers large language models with image and video understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024. p. 13700-13710.
Ren S, Yao L, Li S, Sun X, Hou L. TimeChat: a time-sensitive multimodal large language model for long video understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024. p. 14313-14323.
Song E, Chai W, Wang G, Zhang Y, Zhou H, Wu F, et al. MovieChat: from dense token to sparse memory for long video understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024. p. 18221-18232.
Li K, Wang Y, He Y, Li Y, Wang Y, Liu Y, et al. MVBench: a comprehensive multi-modal video understanding benchmark. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024. p. 22195-22206.
Li Y, Wang C, Jia J. LLaMA-VID: an image is worth 2 tokens in large language models. In: European Conference on Computer Vision. 2024.
Yan S, Zhu T, Wang Z, Cao Y, Zhang M, Ghosh S, et al. VideoCoCa: video-text modeling with zero-shot transfer from contrastive captioners. arXiv. 2022;2212.04979.
Wang Y, Li K, Li Y, He Y, Huang B, Zhao Z, et al. InternVideo: general video foundation models via generative and discriminative learning. arXiv. 2022;2212.03191.
Li Y, Wang Y, Wang Z, Wang Y, Jiang X, Yang K, et al. VideoChat: chat-centric video understanding. arXiv. 2023;2305.06355.
Tang Y, Bi J, Xu S, Song L, Liang S, Wang T, et al. Video understanding with large language models: a survey. arXiv. 2023;2312.17432.
Xu H, Ye Q, Yan M, Shi Y, Ye J, Xu Y, et al. mPLUG-2: a modularized multi-modal foundation model across text, image and video. Proc Mach Learn Res. 2023;202:38728-38748.