World Foundation Models and Large-Scale Generative Video: From Diffusion-Based Synthesis to Interactive Physical-AI Simulation

Authors

  • Oliver Jensen Author

Keywords:

Wan, physical AI, video generation, World foundation models

Abstract

Generative video modeling evolved rapidly from short text-conditioned video synthesis toward increasingly general models capable of representing temporal dynamics, physical interactions, camera motion, and interactive environments. This progression renewed interest in world models: learned representations that predict future states of an environment and can potentially support planning, robotics, simulation, and embodied artificial intelligence. This review examines the convergence of generative video and world modeling through 2025. Early world-model research demonstrated that compact learned dynamics could support control within imagined environments, while Dreamer-style systems learned behaviors directly through latent dynamics. The development of diffusion models and latent video representations subsequently improved high-fidelity visual synthesis. Make-A-Video transferred text-to-image knowledge into video generation without paired text-video data, while Stable Video Diffusion emphasized scalable dataset curation and latent video diffusion. Lumiere introduced a space-time architecture for coherent full-video generation, and VideoPoet treated video synthesis as multimodal autoregressive generation. Genie advanced the concept of generative interactive environments by learning latent actions from unlabeled video. CogVideoX and HunyuanVideo scaled diffusion-transformer architectures for longer and higher-quality videos. In 2025, Wan provided an open family of large-scale video generation models spanning multiple generation and editing tasks. The Cosmos World Foundation Model platform explicitly positioned generative world models as infrastructure for physical AI, combining video tokenization, pretrained world models, and post-training workflows. WorldSimBench subsequently proposed perceptual and embodied evaluation of video generation systems as world simulators. The review examines temporal coherence, action conditioning, visual tokenization, diffusion Transformers, latent dynamics, physical consistency, evaluation, and simulation-to-action transfer. Major limitations include incorrect physical dynamics, long-horizon drift, expensive data requirements, temporal hallucination, safety, and weak causal understanding despite photorealistic outputs.

 

References

Ha D, Schmidhuber J. World models. arXiv. 2018;1803.10122.

Hafner D, Lillicrap T, Ba J, Norouzi M. Dream to control: learning behaviors by latent imagination. In: International Conference on Learning Representations. 2020.

Hafner D, Lillicrap T, Norouzi M, Ba J. Mastering Atari with discrete world models. In: International Conference on Learning Representations. 2021.

Hafner D, Pasukonis J, Ba J, Lillicrap T. Mastering diverse domains through world models. arXiv. 2023;2301.04104.

Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Adv Neural Inf Process Syst. 2020;33:6840-6851.

Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. p. 10684-10695.

Singer U, Polyak A, Hayes T, Yin X, An J, Zhang S, et al. Make-A-Video: text-to-video generation without text-video data. In: International Conference on Learning Representations. 2023.

Blattmann A, Dockhorn T, Kulal S, Mendelevitch D, Kilian M, Lorenz D, et al. Stable Video Diffusion: scaling latent video diffusion models to large datasets. arXiv. 2023;2311.15127.

Peebles W, Xie S. Scalable diffusion models with Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. p. 4195-4205.

Bar-Tal O, Chefer H, Tov O, Herrmann C, Paiss R, Zada S, et al. Lumiere: a space-time diffusion model for video generation. In: Proceedings of SIGGRAPH Asia 2024 Conference Papers. 2024.

Kondratyuk D, Yu L, Gu X, Lezama J, Huang J, Schindler G, et al. VideoPoet: a large language model for zero-shot video generation. Proc Mach Learn Res. 2024;235:25105-25124.

Bruce J, Dennis MD, Edwards A, Parker-Holder J, Shi Y, Hughes E, et al. Genie: generative interactive environments. Proc Mach Learn Res. 2024;235:4603-4623.

Kong W, Tian Q, Zhang Z, Min R, Dai Z, Zhou J, et al. HunyuanVideo: a systematic framework for large video generative models. arXiv. 2024;2412.03603.

Yang Z, Teng J, Zheng W, Ding M, Huang S, Xu J, et al. CogVideoX: text-to-video diffusion models with an expert Transformer. In: International Conference on Learning Representations. 2025.

Polyak A, Zohar A, Brown A, Tjandra A, Sinha A, Lee A, et al. Movie Gen: a cast of media foundation models. arXiv. 2024;2410.13720.

Chen H, Zhang Y, Cun X, Xia M, Wang X, Weng C, Shan Y. VideoCrafter2: overcoming data limitations for high-quality video diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.

Ho J, Chan W, Saharia C, Whang J, Gao R, Gritsenko A, et al. Imagen Video: high definition video generation with diffusion models. arXiv. 2022;2210.02303.

Wan Team. Wan: open and advanced large-scale video generative models. arXiv. 2025;2503.20314.

NVIDIA. Cosmos world foundation model platform for physical AI. arXiv. 2025;2501.03575.

Qin Y, Shi Z, Yu J, Wang X, Zhou E, Li L, et al. WorldSimBench: towards video generation models as world simulators. Proc Mach Learn Res. 2025;267:50338-50362.

Published

2025-06-01