Multimodal Computer-Use Agents for Web and Desktop Environments: Cognitive Reasoning, Verification, Robustness, and Interactive Evaluation
Keywords:
multimodal agents, web agents, Computer-use agentsAbstract
Language-model agents increasingly operate through graphical interfaces rather than interacting only with text-based APIs. Computer-use agents must perceive screenshots, interpret textual and visual interface elements, reason over user goals, select actions, recover from interface changes, and determine whether an action has successfully achieved the intended outcome. This review examines multimodal computer-use agents and their evaluation through September 2026. Earlier web-navigation benchmarks such as WebShop, Mind2Web, WebArena, and WorkArena progressively increased environmental realism, while OSWorld extended interaction to complete operating-system environments. AgentBench provided broader interactive evaluation across several agent environments. Advances in multimodal large language models enabled agents to rely increasingly on pixels rather than structured accessibility trees or pre-extracted HTML. By 2025–2026, research shifted toward cognitive structure, robustness, verification, and safety. Web-CogReasoner introduced a framework separating factual, conceptual, and procedural knowledge for multimodal web-agent reasoning. ScienceBoard evaluated computer-using agents in realistic scientific workflows. EgoBench introduced interactive egocentric multimodal tasks requiring perception, tool invocation, multi-hop reasoning, and user interaction. AgentNoiseBench examined robustness under noisy user and tool conditions. Research on multimodal verifiers explored whether vision-language models could reliably judge outcomes across computer-use and robotic trajectories. BLIND-ACT demonstrated a tendency toward blind goal-directedness, in which computer-use agents continue pursuing goals despite ambiguity, infeasibility, or contextual reasons to stop. The review compares visual grounding, cognitive planning, action representation, tool invocation, multimodal verification, robustness testing, and safety. Reliable computer use requires not only stronger perception and reasoning but also explicit mechanisms for uncertainty, refusal, progress checking, and recovery.
References
Yao S, Chen H, Yang J, Narasimhan K. WebShop: towards scalable real-world web interaction with grounded language agents. Adv Neural Inf Process Syst. 2022;35:20744-20757.
Deng X, Gu Y, Zheng B, Chen S, Stevens S, Wang B, et al. Mind2Web: towards a generalist agent for the web. Adv Neural Inf Process Syst. 2023;36.
Zhou S, Xu FF, Zhu H, Zhou X, Lo R, Sridhar A, et al. WebArena: a realistic web environment for building autonomous agents. In: International Conference on Learning Representations. 2024.
Drouin A, Gasse M, Caccia M, Laradji IH, Del Verme M, Marty T, et al. WorkArena: how capable are web agents at solving common knowledge work tasks? arXiv. 2024;2403.07718.
Xie T, Zhang D, Chen J, Li X, Zhao S, Cao R, et al. OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environments. Adv Neural Inf Process Syst. 2024;37.
Liu X, Yu H, Zhang H, Xu Y, Lei X, Lai H, et al. AgentBench: evaluating LLMs as agents. In: International Conference on Learning Representations. 2024.
Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Hambro E, et al. Toolformer: language models can teach themselves to use tools. Adv Neural Inf Process Syst. 2023;36.
Qin Y, Liang S, Ye Y, Zhu K, Yan L, Lu Y, et al. ToolLLM: facilitating large language models to master real-world APIs. In: International Conference on Learning Representations. 2024.
Liu H, Li C, Wu Q, Lee YJ. Visual instruction tuning. Adv Neural Inf Process Syst. 2023;36.
Bai S, Chen K, Liu X, Wang J, Ge W, Song S, et al. Qwen2.5-VL technical report. arXiv. 2025;2502.13923.
Guo Y, Guo C, Sun A, He H, Yang X, Lu Y, et al. Web-CogReasoner: towards multimodal knowledge-induced cognitive reasoning for web agents. In: International Conference on Learning Representations. 2026.
Sun Q, Liu Z, Ma C, Ding Z, Xu F, Yin Z, et al. ScienceBoard: evaluating multimodal autonomous agents in realistic scientific workflows. In: International Conference on Learning Representations. 2026.
Liu Y, Niu T, Wang Z, Dai Z, Qing Y, Wang W, Liu J. EgoBench: an interactive egocentric multimodal benchmark for tool-using agents. arXiv. 2026;2605.27820.
Wang R, Chen Y, Wang Y, Wu C, Fang J, Cai X, et al. AgentNoiseBench: benchmarking robustness of tool-using LLM agents under noisy condition. arXiv. 2026;2602.11348.
Andrade M, Cha J, Ho B, Srihari V, Yadav K, Kira Z. Let’s think in two steps: mitigating agreement bias in MLLMs with self-grounded verification. In: International Conference on Learning Representations. 2026.
Shayegani E, Hines K, Dong Y, Abu-Ghazaleh NB, Lutz R, Whitehead S, et al. Just do it!? Computer-use agents exhibit blind goal-directedness. In: International Conference on Learning Representations. 2026.
Ma C, Zhang J, Zhu Z, Yang C, Yang Y, Jin Y, et al. AgentBoard: an analytical evaluation board of multi-turn LLM agents. In: ICLR Workshop on Large Language Model Agents. 2024.
Wu Q, Bansal G, Zhang J, Wu Y, Li B, Zhu E, et al. AutoGen: enabling next-generation LLM applications via multi-agent conversation. In: Conference on Language Modeling. 2024.
Wang G, Xie Y, Jiang Y, Mandlekar A, Xiao C, Zhu Y, et al. Voyager: an open-ended embodied agent with large language models. Trans Mach Learn Res. 2024.
Park JS, O'Brien JC, Cai CJ, Morris MR, Liang P, Bernstein MS. Generative agents: interactive simulacra of human behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023.