World Models and Vision-Language-Action Systems for Robotic Intelligence: Predictive Simulation, Policy Evaluation, and Whole-Body Control

Authors

  • Arman Salehi Author

Keywords:

VLA, vision-language-action, World models

Abstract

World models seek to represent how an environment evolves and thereby allow an intelligent system to predict the consequences of possible actions before executing them. Although latent world models have long been studied in reinforcement learning, advances in generative video, multimodal perception, and vision-language-action modeling have expanded the concept toward realistic physical simulation and robot control. This review examines world models and vision-language-action systems for robotics through September 2026. Early latent-dynamics methods demonstrated that agents could learn policies using imagined trajectories, while Dreamer-style systems progressively improved sample-efficient control through learned representations of environmental dynamics. Diffusion Policy introduced expressive diffusion-based action generation for visuomotor behavior. RT-1, RT-2, Open X-Embodiment, and Octo demonstrated increasingly general robotic policies trained across heterogeneous datasets and tasks. Generative video systems subsequently offered richer visual predictions of object and environment dynamics. By 2025–2026, this line of work increasingly converged with world modeling. PhysWorld explored robot learning from physical video-generation models. WorldGym used an action-conditioned video model as an environment in which robotic policies could be evaluated through simulated rollouts and vision-language-based rewards. Sequential World Models extended model-based reinforcement learning to multi-robot cooperation through autoregressive agent-wise dynamics. WholeBodyVLA integrated vision-language-action learning with latent actions and reinforcement-learning control for humanoid loco-manipulation. Research on video generation for robotics and world-action models increasingly formalized the transition from predicting future observations toward using predicted futures to determine executable actions. The review identifies physical fidelity, causal correctness, action grounding, sim-to-real transfer, long-horizon error accumulation, and safety as central challenges.

References

Ha D, Schmidhuber J. World models. arXiv. 2018;1803.10122.

Hafner D, Lillicrap T, Ba J, Norouzi M. Dream to control: learning behaviors by latent imagination. In: International Conference on Learning Representations. 2020.

Hafner D, Lillicrap T, Norouzi M, Ba J. Mastering Atari with discrete world models. In: International Conference on Learning Representations. 2021.

Hafner D, Pasukonis J, Ba J, Lillicrap T. Mastering diverse domains through world models. arXiv. 2023;2301.04104.

Chi C, Feng S, Du Y, Xu Z, Cousineau E, Burchfiel B, Song S. Diffusion Policy: visuomotor policy learning via action diffusion. Int J Robot Res. 2024.

Brohan A, Brown N, Carbajal J, Chebotar Y, Chen X, Choromanski K, et al. RT-1: robotics Transformer for real-world control at scale. arXiv. 2022;2212.06817.

Brohan A, Brown N, Carbajal J, Chebotar Y, Dabis J, Finn C, et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In: Conference on Robot Learning. 2023.

Open X-Embodiment Collaboration. Open X-Embodiment: robotic learning datasets and RT-X models. In: IEEE International Conference on Robotics and Automation. 2024.

Octo Model Team. Octo: an open-source generalist robot policy. arXiv. 2024;2405.12213.

Bruce J, Dennis MD, Edwards A, Parker-Holder J, Shi Y, Hughes E, et al. Genie: generative interactive environments. Proc Mach Learn Res. 2024;235:4603-4623.

NVIDIA. Cosmos world foundation model platform for physical AI. arXiv. 2025;2501.03575.

Mao J, He S, Wu HN, You Y, Sun S, Wang Z, et al. Robot learning from a physical world model. arXiv. 2025;2511.07416.

Mei Z, Yin T, Shorinwa O, Badithela A, Zheng Z, Bruno J, et al. Video generation models in robotics: applications, research challenges, future directions. arXiv. 2026;2601.07823.

Quevedo J, Sharma AK, Sun Y, Suryavanshi V, Liang P, Yang S. WorldGym: world model as an environment for policy evaluation. In: International Conference on Learning Representations. 2026.

Zhao Z, Guo H, Chen S, Xu K, Jiang B, Zhu Y, Zhao D. Empowering multi-robot cooperation via sequential world models. In: International Conference on Learning Representations. 2026.

Jiang H, Chen J, Bu Q, Chen L, Shi M, Zhang Y, et al. WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. In: International Conference on Learning Representations. 2026.

Zhang X, Zeng X, Zhang W. From world models to world action models: a concise tutorial for robotics. arXiv. 2026;2607.00836.

Peebles W, Xie S. Scalable diffusion models with Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. p. 4195-4205.

Blattmann A, Dockhorn T, Kulal S, Mendelevitch D, Kilian M, Lorenz D, et al. Stable Video Diffusion: scaling latent video diffusion models to large datasets. arXiv. 2023;2311.15127.

Qin Y, Shi Z, Yu J, Wang X, Zhou E, Li L, et al. WorldSimBench: towards video generation models as world simulators. Proc Mach Learn Res. 2025;267:50338-50362.

Published

2026-06-01